☰ Contents

Metrics

F verified factP decided planC open challenge

Measurement mechanism

Metrics are computed by tools, not counted by agents. The platform's audit log has a pre-determined, parsable format (Platform PG-6); parser tools produce the numbers below and write dated summaries into this document's log section. Agent judgment is used only where tooling cannot measure (e.g. classifying a retraction), and each such metric names its judge. Why: reconstructing June 2026 agent activity took four mining agents over session logs and git — the cost of not having the log (evidence).

Baseline → target

Baselines are measured v1 facts (Baseline); targets are plans.

Metric v1 baseline (source) v2 target Basis Measured by
Human turns per routine support case median 3 (June 2026 sessions) 0 non-gate turns once a shape is proven; gate decisions only during training D111: median 3 → 0 non-gate turns after proof audit-log parser
Gate-decision minutes per case unmeasured ≤ 5 minutes per gate and ≤ 15 minutes per routine case; auto-confirmed gates counted separately (D70) D111: true median is ≈ 38 minutes today, so gate attention must fit inside 15 minutes audit-log parser
Human interventions per dev session median ~18 ≤ 4 gate decisions + ≤ 2 questions per case D111: ~18 interventions → at most six human stops audit-log parser
Recurrence of documented error families 4 families recurred after their rule existed (20.05.2026 audit) 0 Existing v1 failure family target retained; every recurrence means the rule did not fire gate-hit counter + judged review
Root-cause retraction cycles 3 known (SD-1627, SD-1636, SD-1651) 0 D111 keeps the v1 retraction count as the zero target judged from case records
Support throughput ~7 ticket folders/working day, ~5 per calendar day (Jul–Aug 2026) ≥ 7 closed cases per working day at unchanged human hours D111 preserves the measured working-day baseline and shifts the unit to closed cases audit-log parser
Dev acceleration 2.16× (UAT file) · 2.5× personal · CR schedule −24 % ≥ 2× acceleration sustained per case D111 sets the floor below the two measured acceleration examples audit-log parser + git
Vendor dependency on a customer-specific change (D79) correspondence with the vendor plus the full change procedure removed: brief → verified deploy inside hours or days, on the customer's own lineage D79/D111 measure the dependency as time and hand-offs, never money case records + git
People a case needs at once, and the desk at peak (D79) one operator per machine, two machines a case needs one controller; the desk absorbs a temporary need for 5–10 people at once through parallel cases, not headcount D79/D103 treat this as capacity, not staffing platform
Thin shapes (D115) the team's judgement named 7 domains held by one or two people; measured 03.09.2026 from the Internal Tag: 7 thin shapes (≤ 2 holders with ≥ 2 tickets) thin shapes 7 → 0; 35 of 35 shapes closed by ≥ 2 operators with the skill within two quarters D111/D115 replace the old single-holder wording with the measured thin-shape KPI Internal-Tag parser + skill use records
Precipitation rate (SG-3) practiced but unmeasured: 548 files written, and retrieval fails to fire (20.05.2026 audit) 100 % of closed cases carry a precipitation decision; recurring shapes without a skill after the 3rd occurrence = 0 D111 + D23/D25: every close records the precipitation model decision audit-log parser
Reuse hits unmeasured reuse recorded on ≥ 60 % of Support cases within two quarters D111: the catalogue already covers ≈ 2/3 of analysed tickets audit-log parser
Retrieval top-5 hit-rate unmeasured ≥ 50 % per domain before acceptance D111: baseline retrieval is 10 %, so acceptance needs a fivefold lift audit-log parser + held-out replay
Held approved writes 39 HDesk updates held 12 weeks, invisible 0 batches older than 5 business days without a recorded reason D111 converts invisibility into an aged queue with a reason requirement platform queue
Waiting cases without a recorded trigger and age (SG-15) 148 fixed tickets in Pending without a resolution date (archive of 02.09.2026) 0 — every waiting case names its trigger, its clock treatment and its age, and re-enters the route when the trigger fires D147: waiting is a route, not a status an agent sets platform queue
Configuration completeness half-built products happen (KI-052 class) unreachable products 0; every deployment-set decision recorded in the plan revision that set it D111: unreachable products 0 and deployment-set decisions must be attributable case records
Configuration elapsed time (CG-13) configuration-change tickets close in a median 9 days and master-data tickets in 11, over the 175 customer-closed tickets intake → preview ≤ 1 business day; intake → customer acceptance ≤ 5 business days median; unresolved[] ≤ 5 at H1 and 0 at close; ≤ 2 H1 correction rounds; QT rows 0 D111: CG-13 turns those closing medians into elapsed-time and quality targets case records
Configuration verification acceptance scripts exist, never executed by v1 skills green on target + smoke tests after production deploy; QT rows 0 D111: QT rows must stay 0 and target proof replaces hand-written scripts skill runners
Session resumability session death = disk archaeology resume with plan + evidence intact after kill PG-9: papers and ledger are state; the worker is disposable scripted test
Deployment discipline manual, out-of-band 100 % of applied changes carry scripts + test evidence + revert path D111/DG-9 make revert execution a measured deployment discipline audit-log parser
SLA visibility clocks watched by a human warning at 50 % of the window, escalation at 75 %, 0 breaches without a warning, response breaches ≤ 2 % per quarter D111 supplies the lead times and breach ceiling platform
Classification correction rate (SG-11) unmeasured ≤ 10 % after three months; ≤ 5 % steady D111 defines the training and steady-state thresholds audit-log parser
Re-routing rate unmeasured ≤ 5 % D111: route correction is separate from classification correction audit-log parser
Dedup rate and merged duplicates (PG-20) two channels carried one incident, unmeasured every duplicate attached, none investigated twice PG-20: the origin key makes duplicate work visible audit-log parser
Non-EU model calls (SG-9) n/a 0 D111/SG-9 keep the EU boundary absolute route records
Handovers Hd and CR hand-offs (SG-4, SG-11) unmeasured counted with their packets SG-4/SG-11 require the packet to carry the hand-off reason audit-log parser
Memory retractions (SG-14) unmeasured counted with affected cases reviewed SG-14: a wrong article creates ledger CORRECTION records and review tasks ledger CORRECTION records
Specification quality (CG-13) v1 does not execute configuration unresolved[] ≤ 5 at H1 and 0 at close; ≤ 2 H1 correction rounds; QT rows written = 0 D111 gives the numeric quality gates for CG-13 case records
Placement share, backport success, reverts (DG-9) unmeasured ≥ 80 % of changes entirely in seams 1–3; 100 % of seam-4 landings with a proposition; port MR ≤ 1 business day after PROD verification; port rework ≤ 10 %; reverts ≤ 1 per 20 deploys, executed ≤ 1 h D111: baseline is 5 of 8 case studies for seam placement ledger + git
Refusals before contact (PG-10) v1 has no enforcement every refusal recorded; writes without a grant = 0 D111/PG-10: grants, not content, are authority connector records
Auto-confirmed gates v1 has no enforcement corrections or rollbacks ≤ 1 % D111/D70: policy confirmation is safe only while corrections remain rare ledger + gate records
Parked hours per outage (PG-22) unmeasured resume ≤ 15 minutes after recovery D111: recovery is measured by time back to work, not outage duration alone audit-log parser
Verifier and auditor share unmeasured ≤ 25 % of case reasoning at steady state D111: checks must cost less than the errors they catch audit-log parser

Every target is a configuration default (D51) and is re-baselined by the ledger after the first ten cases of its kind (D68).

Spend envelope

P Spend is a control, not a metric (D79). Provider usage is budgeted, capped and escalated so that a case cannot run away; it is never reported as evidence that the platform works, and no figure here is expressed in currency. The measures of value are hours, bus factor, people involved, domains involved and elapsed time (Value and ROI).

Why the control is needed: v2 uses more provider capacity per case by construction — agents consult laterally, the verifier re-derives work the author already did, the write auditor re-derives the affected set, and eval sets re-run whenever a memory domain, skill or model route changes (CR-4). Unbudgeted, that is how a platform consumes more effort than the people it accelerates.

P Provider usage is a first-class control, not a monthly surprise:

Control
Per-case budget minutes — declared at case open from the case type's default and replaced by the predicted delivery time at classification, the larger of the two (D135); provider usage, expensive reads, sub-agent count and DB sessions are caps beside it (Agent Runtime § 11.5). Exceeding it escalates to the root agent, then to the human — it does not silently continue.
Per-profile budget a profile carries its own share or fixed ceiling in minutes (Agents § 6), so an expensive verifier cannot be spawned in a loop.
Eval-set batching re-runs are batched per domain, on a schedule and before publish — not on every memory write (Agent Framework § 6.5).
Route by profile cheap routes for mechanical stages (test runners, script emitters), expensive ones only where judgement is the work.

P Targets, measured from the ledger per module:

Control v1 v2 target
Budget breaches per period none enforceable — v1 has no budget object 0 silent continuations; every breach escalated and recorded
Provider usage against the declared budget unmeasured inside the budget the case type declares, or escalated before it is exceeded
Verifier and auditor share of a case unmeasured ≤ 25 % of case reasoning at steady state; a check that costs more than the errors it catches is re-tiered
Eval-set runs per module per week n/a bounded and reported; a set that costs more than the errors it catches is trimmed

P Verification overhead is the open number: an isolated verifier re-deriving a mechanism roughly doubles the reasoning cost of a case. Whether that is worth it is measured against the metric it exists for — root-cause retraction cycles, of which v1 had three known. Decide after the first ten real cases, not now.

P Initial estimations, to be further analysed (decision Vladimir, 02.09.2026 — D68; starting figures, not conclusions):

Quantity Initial estimate Basis
Mandatory-tier cost +100 % of the case's reasoning on verified cases a verifier re-derives the mechanism from the same evidence (Case Studies G-3: the verifier doubles a 30–40-minute ticket)
Sampled-tier cost at N = 5 +4–8 % spread across known-shape cases one full re-derivation per five runs
Retraction base rate 3 in 340 analyses (≈ 0.9 %), each ≈ a day of rework + credibility Baseline § 5; the analysis count from Data and Information Module § 0.2
Break-even the tier pays when prevented retractions × (rework + follow-up effort) > verifier usage — at the base rate, the mandatory tier belongs on customer-facing mechanisms and memory writes only, which is where it already sits derived

The first ten cases replace every figure here with ledger numbers; the note stays until they do.

First proofs

  1. Platform experiment — the C# session host (PG-1/PG-2) demonstrating: concurrent isolated sessions with sub-agents; plan confirmation upward; escalation to the human; resume after an answer with nothing lost; model route swap without session-record migration; tool interception with policy enforcement; audit records for every connector result. An SDK handshake alone proves nothing.
  2. One real configuration slice — a real source folder on an authorized target through S1, phase 1 and both branches of phase 2, closed by the joint proof and a full deployment set, deliberately including one initially missing element and one discovered mid-run, so the early-ask (H1) and per-stage iteration loops and MISSING handling are demonstrated. The 9951 run (Product Configurator § SRC) is the single-environment precedent; the v2 slice adds sessions, gates, audit, and deployment.
  3. Round-trip golden set — export ≥5 shape-diverse selling products with their real sources; re-derive each from sources alone in a clean session; grade with the isolated grader; re-run on knowledge/model changes. Specified and partially exercised in the PC workspace (facts).

Measurement log

(appended by tools — empty until the platform exists)