PLU-001 · Continuous AI Evaluation

Regulations are tightening. Auditors are arriving. And regulated mid-market teams can't replay a single AI decision from three years ago.

Continuous evaluationdrift ledgersigned recordstamper-evident evidence.

For regulated mid-market teams running models in healthcare, finance, or the public sector, the audit posture is no longer six-month-old. The EU AI Act and NIST AI RMF make years-later replay a baseline expectation — and your auditor opens a record, not a slide deck.

Plumbline AI turns every model decision into a continuous evaluation, written as a drift ledger of signed records. Each link in the chain is tamper-evident evidence — your auditor reconstructs a decision end to end, with no logs to piece together after the fact.

Start with one model, one evaluator, and one auditor-shaped question. We'll show you the replay before any contract is signed.

Not ready to register a model yet?
Drop your email and we'll let you know when a pilot slot opens. No marketing blast — only the next-step note.

The register-model flow is signed-in only — you'll be redirected to sign in if you haven't yet.

Frameworks
NIST · EU AI Act
Programs
HIPAA · SOC 2 · SOX
Records
Signed, replayable
Plumb line — the namesake instrument18130803μ = 0.946 ± 0.012DECISION · EVAL · SIGNED

§01 · Why now

Audits fail for one reason— the evidence was never captured.

NIST AI RMF and the EU AI Act are turning immutable evaluation, drift, and incident records into a baseline expectation. The governance stack a hyperscaler runs is overkill for most programs — and the spreadsheets most teams fall back on don't survive contact with an auditor.

Plumbline sits between those two: an evaluation substrate that mid-market and regulated teams can actually afford, with the records your auditor needs to replay a decision five months later.

§02 · The artifact

One ledger.
Tamper-evident end to end.

Every evaluation run is hashed, signed, and chained to the prior record. Tampering with one decision breaks the chain; the replay surfaces exactly where it broke.

Tamper-evident evaluation ledger#1,0370.951EVALsig:7f3c…ad4#1,0380.947Δ −0.004EVALsig:4a91…0be#1,0390.882Δ −0.065EVAL · INCIDENTsig:9b22…3f7#1,0400.913Δ +0.031EVAL · RESOLVEDsig:cc15…8e1CHAIN · sha-256 · KEY-ANCHORED · REPLICATED

Eval · Event · Chain-record #1,039 — sha-256 chain · signed · replicated

Each row is an evaluation event — a model or agent decision, scored, timestamped, and linked to its inputs. Drifts and incidents live in the same ledger; they don't get lost between tools.

  • Versioned. Replay the same decision against three different evaluation runs — without rebuilding the dataset.
  • Signed. Every record carries a key-anchored signature. Provenance is a query, not a forensic exercise.
  • Access-controlled. Retention tuned to HIPAA, SOC 2, and SOX windows — not out-of-the-box defaults.

§03 · Methodology

Four steps, one ledger.

A typical pipeline touches four Plumbline calls. Each one writes a record. The ledger doesn't depend on anyone remembering to run it.

  1. 01

    Capture

    The SDK tags every decision with a stable evaluation key — input, output, model, version.

  2. 02

    Evaluate

    Lightweight evaluators run inline or asynchronously. Scores attach to the same record, not a separate log.

  3. 03

    Alert & resolve

    Drift or incident triggers hit the workflow you already use. Resolution notes are signed back into the ledger.

  4. 04

    Replay

    Auditor opens a record; the chain walks them back to the input — every step provable, nothing reconstructed.

§04 · Compliance

Built for teams whose
auditor already has a name.

Optional compliance packages for healthcare, financial services, and public-sector buyers navigating AI-specific frameworks.

NIST AI RMF

US federal

Map, measure, manage — with evidence retrieval aligned to the risk profile.

EU AI Act

high-risk systems

Logging, post-market monitoring, and technical documentation that survives an inspection.

HIPAA

healthcare

PHI-aware retention and access scopes. BAA available on the enterprise tier.

SOC 2 · SOX

enterprise IT

Immutable change records for ITGC and model-risk management — auditor-friendly exports.

§05 · Pricing

Three tiers.
Sized to what you actually run.

Pricing tracks monitored models and monthly evaluation volume — not seats. Pick the tier that matches your footprint today; we'll re-size with you after the first 60 days.

The full billing surface — including the launch code and renewal math — lives on the pricing page. This section orients new visitors to the three shapes we offer.

Solo
Solo
For a single practitioner running one model in production.
$79 / mo
Monitored models
1 model
Monthly evaluations
Limited / month
Regulated Team
Most chosen
For auditors and the teams who brief them — role-aware seats and a signed trail.
$499 / mo
Monitored models
10 models
Monthly evaluations
Higher quota / month
Enterprise
Enterprise
Custom quotas, dedicated environment, and a DPA with named sub-processors.
Contact us
Monitored models
Unlimited
Monthly evaluations
Custom quotas

§06 · Trust

What an auditor actually checks.

Three bands of proof, each anchored to a control reference, a code path, or a tier term — not a hand-wave. We're in build-up: controls are mapped, signing lives in the codebase, and the export bundle is live. A formal SOC 2 Type II audit is on the roadmap, not a present claim.

Column 01 · Compliance coverage

Frameworks, mapped control by control.

Each framework below is grounded in a real catalog row — same artefact that populates the compliance page.

  • HIPAA

    Controls mapped

    45 CFR § 164

    • §164.312(b) — Audit controls (replayable signed records)
    • §164.308(a)(4) — Access control (per-model API keys)
    • §164.312(c)(1) — Integrity (SHA-256 hash chain)
  • SOX

    ITGC mapped

    §404 ITGC

    • §404 — Management assessment of AI systems
    • ITGC — Change + logical access
    • ITGC — Audit trail retention
  • SOC 2

    TSC mapped

    TSC 2017

    • CC6.1 — Logical access (per-model isolation)
    • CC7.1 / CC7.2 — Anomaly detection + monitoring
    • CC4.1 — Monitoring of controls

Column 02 · Signed records

A cryptographic audit trail, end to end.

The signed-record table isn't reconstruction — it's the evidence the auditor opens.

  • SHA-256 hash chain

    Every record is sha256(canonicalJson(record)). Tamper one row, break the chain — the replay finds it.

    sha256(eval) → sha256(eval·prev)

  • Key-anchored signatures

    Each capture carries a signature anchored to its registered model key. Provenance is a query.

    sig:7f3c…ad4

  • Replayable provenance

    Distinct (model, run) pairs in the same ledger. Auditor replays months later — no re-running inference.

    (model, runId) → chain

Column 03 · Security posture

Posture that holds up to a checklist.

What's in production today — verifiable in the code, the tier catalog, and the audit-bundle export.

  • Per-model API-key isolation

    SOC 2 CC6.1

    No model authenticates as another — the access boundary is the registered-model row.

  • Retention aligned to legal windows

    2,555-day default

    Team and Enterprise tiers default to the HIPAA / SOX-aligned 7-year window — not an out-of-the-box 30.

  • BAA on the enterprise tier

    Enterprise

    Healthcare buyers sign a BAA at the enterprise tier — on the order form, not ad hoc.

  • Tamper-evident audit bundle

    export_event

    One-click PDF export — eval records, drifts, alerts, retention, access — with a chain-of-custody seal.

§07 · The team

Audit trails were never optional in our prior lives.
So we built one for AI.

Two decades across healthcare-IT, Microsoft-channel enterprise sales, and regulated SaaS — applied to mid-market AI governance that teams and auditors can both live with.