Metrics
Measurement mechanism
Metrics are computed by tools, not counted by agents. The platform's audit log has a pre-determined, parsable format (Platform PG-6); parser tools produce the numbers below and write dated summaries into this document's log section. Agent judgment is used only where tooling cannot measure (e.g. classifying a retraction), and each such metric names its judge. Why: reconstructing June 2026 agent activity took four mining agents over session logs and git — the cost of not having the log (evidence).
Baseline → target
Baselines are measured v1 facts (Baseline); targets are plans.
| Metric | v1 baseline (source) | v2 target | Basis | Measured by |
|---|---|---|---|---|
| Human turns per routine support case | median 3 (June 2026 sessions) | 0 non-gate turns once a shape is proven; gate decisions only during training | D111: median 3 → 0 non-gate turns after proof | audit-log parser |
| Gate-decision minutes per case | unmeasured | ≤ 5 minutes per gate and ≤ 15 minutes per routine case; auto-confirmed gates counted separately (D70) | D111: true median is ≈ 38 minutes today, so gate attention must fit inside 15 minutes | audit-log parser |
| Human interventions per dev session | median ~18 | ≤ 4 gate decisions + ≤ 2 questions per case | D111: ~18 interventions → at most six human stops | audit-log parser |
| Recurrence of documented error families | 4 families recurred after their rule existed (20.05.2026 audit) | 0 | Existing v1 failure family target retained; every recurrence means the rule did not fire | gate-hit counter + judged review |
| Root-cause retraction cycles | 3 known (SD-1627, SD-1636, SD-1651) | 0 | D111 keeps the v1 retraction count as the zero target | judged from case records |
| Support throughput | ~7 ticket folders/working day, ~5 per calendar day (Jul–Aug 2026) | ≥ 7 closed cases per working day at unchanged human hours | D111 preserves the measured working-day baseline and shifts the unit to closed cases | audit-log parser |
| Dev acceleration | 2.16× (UAT file) · 2.5× personal · CR schedule −24 % | ≥ 2× acceleration sustained per case | D111 sets the floor below the two measured acceleration examples | audit-log parser + git |
| Vendor dependency on a customer-specific change (D79) | correspondence with the vendor plus the full change procedure | removed: brief → verified deploy inside hours or days, on the customer's own lineage | D79/D111 measure the dependency as time and hand-offs, never money | case records + git |
| People a case needs at once, and the desk at peak (D79) | one operator per machine, two machines | a case needs one controller; the desk absorbs a temporary need for 5–10 people at once through parallel cases, not headcount | D79/D103 treat this as capacity, not staffing | platform |
| Thin shapes (D115) | the team's judgement named 7 domains held by one or two people; measured 03.09.2026 from the Internal Tag: 7 thin shapes (≤ 2 holders with ≥ 2 tickets) | thin shapes 7 → 0; 35 of 35 shapes closed by ≥ 2 operators with the skill within two quarters | D111/D115 replace the old single-holder wording with the measured thin-shape KPI | Internal-Tag parser + skill use records |
| Precipitation rate (SG-3) | practiced but unmeasured: 548 files written, and retrieval fails to fire (20.05.2026 audit) | 100 % of closed cases carry a precipitation decision; recurring shapes without a skill after the 3rd occurrence = 0 | D111 + D23/D25: every close records the precipitation model decision | audit-log parser |
| Reuse hits | unmeasured | reuse recorded on ≥ 60 % of Support cases within two quarters | D111: the catalogue already covers ≈ 2/3 of analysed tickets | audit-log parser |
| Retrieval top-5 hit-rate | unmeasured | ≥ 50 % per domain before acceptance | D111: baseline retrieval is 10 %, so acceptance needs a fivefold lift | audit-log parser + held-out replay |
| Held approved writes | 39 HDesk updates held 12 weeks, invisible | 0 batches older than 5 business days without a recorded reason | D111 converts invisibility into an aged queue with a reason requirement | platform queue |
| Waiting cases without a recorded trigger and age (SG-15) | 148 fixed tickets in Pending without a resolution date (archive of 02.09.2026) | 0 — every waiting case names its trigger, its clock treatment and its age, and re-enters the route when the trigger fires | D147: waiting is a route, not a status an agent sets | platform queue |
| Configuration completeness | half-built products happen (KI-052 class) | unreachable products 0; every deployment-set decision recorded in the plan revision that set it | D111: unreachable products 0 and deployment-set decisions must be attributable | case records |
| Configuration elapsed time (CG-13) | configuration-change tickets close in a median 9 days and master-data tickets in 11, over the 175 customer-closed tickets | intake → preview ≤ 1 business day; intake → customer acceptance ≤ 5 business days median; unresolved[] ≤ 5 at H1 and 0 at close; ≤ 2 H1 correction rounds; QT rows 0 |
D111: CG-13 turns those closing medians into elapsed-time and quality targets | case records |
| Configuration verification | acceptance scripts exist, never executed by v1 | skills green on target + smoke tests after production deploy; QT rows 0 | D111: QT rows must stay 0 and target proof replaces hand-written scripts | skill runners |
| Session resumability | session death = disk archaeology | resume with plan + evidence intact after kill | PG-9: papers and ledger are state; the worker is disposable | scripted test |
| Deployment discipline | manual, out-of-band | 100 % of applied changes carry scripts + test evidence + revert path | D111/DG-9 make revert execution a measured deployment discipline | audit-log parser |
| SLA visibility | clocks watched by a human | warning at 50 % of the window, escalation at 75 %, 0 breaches without a warning, response breaches ≤ 2 % per quarter | D111 supplies the lead times and breach ceiling | platform |
| Classification correction rate (SG-11) | unmeasured | ≤ 10 % after three months; ≤ 5 % steady | D111 defines the training and steady-state thresholds | audit-log parser |
| Re-routing rate | unmeasured | ≤ 5 % | D111: route correction is separate from classification correction | audit-log parser |
| Dedup rate and merged duplicates (PG-20) | two channels carried one incident, unmeasured | every duplicate attached, none investigated twice | PG-20: the origin key makes duplicate work visible | audit-log parser |
| Non-EU model calls (SG-9) | n/a | 0 | D111/SG-9 keep the EU boundary absolute | route records |
| Handovers Hd and CR hand-offs (SG-4, SG-11) | unmeasured | counted with their packets | SG-4/SG-11 require the packet to carry the hand-off reason | audit-log parser |
| Memory retractions (SG-14) | unmeasured | counted with affected cases reviewed | SG-14: a wrong article creates ledger CORRECTION records and review tasks | ledger CORRECTION records |
| Specification quality (CG-13) | v1 does not execute configuration | unresolved[] ≤ 5 at H1 and 0 at close; ≤ 2 H1 correction rounds; QT rows written = 0 |
D111 gives the numeric quality gates for CG-13 | case records |
| Placement share, backport success, reverts (DG-9) | unmeasured | ≥ 80 % of changes entirely in seams 1–3; 100 % of seam-4 landings with a proposition; port MR ≤ 1 business day after PROD verification; port rework ≤ 10 %; reverts ≤ 1 per 20 deploys, executed ≤ 1 h | D111: baseline is 5 of 8 case studies for seam placement | ledger + git |
| Refusals before contact (PG-10) | v1 has no enforcement | every refusal recorded; writes without a grant = 0 | D111/PG-10: grants, not content, are authority | connector records |
| Auto-confirmed gates | v1 has no enforcement | corrections or rollbacks ≤ 1 % | D111/D70: policy confirmation is safe only while corrections remain rare | ledger + gate records |
| Parked hours per outage (PG-22) | unmeasured | resume ≤ 15 minutes after recovery | D111: recovery is measured by time back to work, not outage duration alone | audit-log parser |
| Verifier and auditor share | unmeasured | ≤ 25 % of case reasoning at steady state | D111: checks must cost less than the errors they catch | audit-log parser |
Every target is a configuration default (D51) and is re-baselined by the ledger after the first ten cases of its kind (D68).
Spend envelope
P Spend is a control, not a metric (D79). Provider usage is budgeted, capped and escalated so that a case cannot run away; it is never reported as evidence that the platform works, and no figure here is expressed in currency. The measures of value are hours, bus factor, people involved, domains involved and elapsed time (Value and ROI).
Why the control is needed: v2 uses more provider capacity per case by construction — agents consult laterally, the verifier re-derives work the author already did, the write auditor re-derives the affected set, and eval sets re-run whenever a memory domain, skill or model route changes (CR-4). Unbudgeted, that is how a platform consumes more effort than the people it accelerates.
P Provider usage is a first-class control, not a monthly surprise:
| Control | |
|---|---|
| Per-case budget | minutes — declared at case open from the case type's default and replaced by the predicted delivery time at classification, the larger of the two (D135); provider usage, expensive reads, sub-agent count and DB sessions are caps beside it (Agent Runtime § 11.5). Exceeding it escalates to the root agent, then to the human — it does not silently continue. |
| Per-profile budget | a profile carries its own share or fixed ceiling in minutes (Agents § 6), so an expensive verifier cannot be spawned in a loop. |
| Eval-set batching | re-runs are batched per domain, on a schedule and before publish — not on every memory write (Agent Framework § 6.5). |
| Route by profile | cheap routes for mechanical stages (test runners, script emitters), expensive ones only where judgement is the work. |
P Targets, measured from the ledger per module:
| Control | v1 | v2 target |
|---|---|---|
| Budget breaches per period | none enforceable — v1 has no budget object | 0 silent continuations; every breach escalated and recorded |
| Provider usage against the declared budget | unmeasured | inside the budget the case type declares, or escalated before it is exceeded |
| Verifier and auditor share of a case | unmeasured | ≤ 25 % of case reasoning at steady state; a check that costs more than the errors it catches is re-tiered |
| Eval-set runs per module per week | n/a | bounded and reported; a set that costs more than the errors it catches is trimmed |
P Verification overhead is the open number: an isolated verifier re-deriving a mechanism roughly doubles the reasoning cost of a case. Whether that is worth it is measured against the metric it exists for — root-cause retraction cycles, of which v1 had three known. Decide after the first ten real cases, not now.
P Initial estimations, to be further analysed (decision Vladimir, 02.09.2026 — D68; starting figures, not conclusions):
| Quantity | Initial estimate | Basis |
|---|---|---|
| Mandatory-tier cost | +100 % of the case's reasoning on verified cases | a verifier re-derives the mechanism from the same evidence (Case Studies G-3: the verifier doubles a 30–40-minute ticket) |
| Sampled-tier cost at N = 5 | +4–8 % spread across known-shape cases | one full re-derivation per five runs |
| Retraction base rate | 3 in 340 analyses (≈ 0.9 %), each ≈ a day of rework + credibility | Baseline § 5; the analysis count from Data and Information Module § 0.2 |
| Break-even | the tier pays when prevented retractions × (rework + follow-up effort) > verifier usage — at the base rate, the mandatory tier belongs on customer-facing mechanisms and memory writes only, which is where it already sits |
derived |
The first ten cases replace every figure here with ledger numbers; the note stays until they do.
First proofs
- Platform experiment — the C# session host (PG-1/PG-2) demonstrating: concurrent isolated sessions with sub-agents; plan confirmation upward; escalation to the human; resume after an answer with nothing lost; model route swap without session-record migration; tool interception with policy enforcement; audit records for every connector result. An SDK handshake alone proves nothing.
- One real configuration slice — a real source folder on an authorized target through S1, phase 1 and both branches of phase 2, closed by the joint proof and a full deployment set, deliberately including one initially missing element and one discovered mid-run, so the early-ask (H1) and per-stage iteration loops and MISSING handling are demonstrated. The 9951 run (Product Configurator § SRC) is the single-environment precedent; the v2 slice adds sessions, gates, audit, and deployment.
- Round-trip golden set — export ≥5 shape-diverse selling products with their real sources; re-derive each from sources alone in a clean session; grade with the isolated grader; re-run on knowledge/model changes. Specified and partially exercised in the PC workspace (facts).
Measurement log
(appended by tools — empty until the platform exists)