Value and ROI
The unit rule
P Money is not a metric of this platform (decision Vladimir, 03.09.2026 — D79, restating and widening D52). No figure on this page, in Metrics or on any presentation of them is expressed in currency — not an hourly rate, not a cost per ticket, not a cost per saved day, and not the provider's bill. A currency figure is an operations line owned by whoever pays it; it is not evidence that the platform works.
The valid measures, per module:
| Module | What is measured |
|---|---|
| Support | hours — elapsed and human, per ticket shape; bus factor — how many people can work a knowledge domain, and how many domains have only one; people involved — how many humans a case needs at once and how many the desk needs at peak; knowledge domains involved per case |
| Configuration | hours and days — from intake to a preview, and from intake to a deployed, accepted product; people involved; domains involved |
| Development | time — from brief to a verified deploy; and vendor dependency removed — whether a customer-specific change still requires correspondence with a vendor and a long procedure, or does not |
P The peak-capacity measure the support desk actually needs (stated 03.09.2026): the desk must absorb a temporary need for five to ten people working at the same time — an incident day, a release week, a migration — without hiring for the peak and without the queue collapsing onto the one person who knows the shape. Parallel cases on one platform are how that is met; the metric is concurrent cases carried and the number of distinct humans involved while carrying them.
Purpose
Two audiences need the same number and cannot use the same page. The CEO wants half a page and yes/no answers — that is documented: the long form of July 2026 was rejected for exactly this (evidence). The development manager wants to see how the number is produced. This page is the second; § 6 is the first.
1. What a ticket costs today P
The salary-hours of the person who works it are the smallest part.
| Cost | What it is | Evidence / where the number comes from |
|---|---|---|
| Handling time | the planning baseline is 30 minutes of human time per ticket, on average, for non-agent work (decision Vladimir, 02.09.2026 — D69; roi.human_minutes_per_ticket, Whitelabel Catalogue § 5) — a non-development ticket takes 5 minutes to a couple of hours depending on the person and the ticket; an investigation runs 1–3 h; a multi-day case exists |
the mined effort distribution over the 331 tickets in the effort row: 40 trivial · 116 routine · 138 investigation · 37 multi-day, out of 485 unique tickets (archive of 02.09.2026); 340 ticket folders carry a written analysis (03.09.2026) (Data and Information Module § 0.2); the 30-min average sits with the measured median true time (§ 2b) |
| Context switching | there is no dedicated support team and no L2/L3 line: the same people configure products, develop and support, so every ticket is a switch out of development and back. The switch cost is about 20 % of the ticket's handling time (Vladimir, 02.09.2026 — D53; a configuration default, Whitelabel Catalogue § 5). The larger cost is not the minutes: it is the fogginess after the switch — a raised risk of introducing bugs and of slower or worse execution on the non-support work the person returns to. That cost is tracked as a count (defects and reworks on non-support work within the day after an interruption), not as hours | organisational fact, stated 02.09.2026; the CEO evidence records the same people on CRs and defects |
| Bus factor | measured thin shapes (≤ 2 holders with ≥ 2 tickets each) are 7, and 25 of 35 tag-bearing shapes have a lead share ≥ 50 % (D115); when the available person is away the ticket waits or takes hours instead of minutes | the bus-factor column of Data and Information Module § 0.2; v1 memory holds the articles, but articles are not executable |
| Rediscovery | knowledge that exists but does not fire costs the same hour again: cached knowledge is 10–60× cheaper and 30–300× faster than rediscovery | measurement; the 20.05.2026 audit (Baseline § 5) |
| Retractions | a confident wrong root cause costs a full cycle — HTML, Jira, memory — three known in v1 | Baseline § 5 |
| Waiting | approved work that waits invisibly: 39 HDesk updates held 12 weeks | backlog |
| Elapsed time on configuration requests | configuration-change tickets close in a median 9 days and master-data tickets in 11, against 1–2 for corrections — the customer waits a working fortnight for a config cell | mined resolution dates, 175 customer-closed tickets (Data and Information Module § 0.2) |
P Human effort is measured in hours, never monetised (decision Vladimir, 02.09.2026 — D52). The ledger computes per case: human hours + switch cost hours (default 20 % of the ticket's handling time — D53) + provider usage as a control (D117) + waiting days × SLA exposure, with the bus-factor exposure tracked as the thin-shape count and the fogginess risk as a count of defects introduced after interruptions (D53). No hourly rate appears anywhere in this wiki. For planning before the ledger exists, human time is taken as 30 minutes per ticket on average for non-agent work (D69) — replaced per case type by the stopwatch protocol (§ 5) and then by the ledger.
2. What v1 already delivered F
| Measure | Value |
|---|---|
| Development acceleration | 2.16× on the UAT Health file (106.75 → 49.5 developer-days over 46 rows); 2.5× personally (45 → 18 days); CR schedule −24 % (179 → 136.1 days), technical phases only |
| Support throughput | frozen baseline 31.08.2026: 509 ticket folders · 256 mail threads · 52 HDesk tickets; 340 ticket folders carry a written analysis (archive of 03.09.2026, Data and Information Module § 0.2); ~7 ticket folders per working day, ~5 per calendar day (Jul–Aug 2026) |
| Human effort per ticket | median 3 human turns per support session |
Sources: Baseline § 2–3, CEO evidence. These are the numbers a CEO can be given today; they are measured, not projected.
2b. Measured on the ticket archive — agent time against the human estimate F, adjusted P
Measured 02.09.2026 over 485 tickets — 185 with a reliable timing and 230 carrying a human estimate (Measurements § TT-2; script scripts/mine_tickets_timing.py). Two quantities per ticket:
- Agent time = elapsed from the moment AISA fetched the ticket (the JSON snapshot timestamp) to its last time-bearing AISA timeline stamp or, failing that, the first commit at or after the snapshot, plus 15 minutes of operator time (Vladimir's rule). Reliable for 185 tickets; the naive "last commit" estimator is unusable because bulk retrofits later touched every directory, and file modification times collapse onto three checkout dates.
- Human estimate = the role × hours table the analysis itself carries in its Financial card (230 tickets parsed, 194 cross-checked against amount ÷ rate). Per Vladimir, the true time of the highly skilled people who actually do this work is about 25 % of that estimate.
| Hours per ticket | n | median | p25 | p75 |
|---|---|---|---|---|
| Agent time, reliable subset — measured | 185 | 2.0 | 1.0 | 4.6 |
| Agent time, timeline-based only — measured | 111 | 1.5 | ||
| Agent active work windows (gaps > 4 h excluded), where measurable | 158 | 1.6 | 2.8 | |
| Agent time, adjusted −25 % (see the note) | 185 | 1.5 | 0.75 | 3.4 |
| Human estimate (analysis author's card) | 230 | 2.0 | 1.0 | 4.9 |
| Human true time (25 % of the estimate) — measured rule | 230 | 0.5 | ||
| Human true time, adjusted +25 % (see the note) | 230 | 0.63 |
The adjustment note P (decision Vladimir, 02.09.2026 — D54). A colleague's experience is that the agent's time is lower and the person's higher than the archive shows: the agent's elapsed includes operator waits and re-analysis days, and the 25 % rule is a flat guess. Until the ten-ticket stopwatch measurement replaces them (§ 5), the model uses the agent time reduced by 25 % and the person's time raised by 25 %, and every figure derived from them carries this note. Times vary by domain and by person — a cargo repair by the one colleague who knows the family is minutes, the same ticket for anyone else is hours — so the medians are the shape of the comparison, not a number for any one ticket. Both adjustments are configuration defaults (Whitelabel Catalogue § 5).
| Request type | Agent time, median h (measured → adjusted) | Human true, median h (measured → adjusted) | Ratio human / agent, adjusted |
|---|---|---|---|
| data correction | 1.8 → 1.35 | 0.5 → 0.63 | 0.46 |
| transfer / sync | 1.7 → 1.28 | 0.5 → 0.63 | 0.49 |
| incident / defect | 2.6 → 1.95 | 0.9 → 1.13 | 0.58 |
| access / account | 1.9 → 1.43 | 0.2 → 0.25 | 0.18 |
| configuration change | 2.4 → 1.80 | 0.4 → 0.50 | 0.28 |
| master data | 3.9 → 2.93 | 0.3 → 0.38 | 0.13 |
The honest reading. With the adjustment, a skilled person closes the median ticket in about 40 minutes of their own time; AISA v1 takes about 1.5 hours of elapsed time, of which the operator spends roughly 15–30 minutes (the 15-minute rule plus a median of three turns). So per ticket v1 saves the skilled person's own time but not the elapsed time — the person is busy for a quarter of the agent's run — while producing a written analysis, the Jira notes and the memory entry as by-products; and the comparison varies by domain and person (the note above). Where v1 wins is not the median ticket: it is the ticket the skilled person is not there for (the bus factor), the ticket nobody has seen before (rediscovery), and the ten tickets that can run at once. That is exactly the case v2 has to make with measured numbers, and it is why the catalogue's saving figures in Data and Information Module § 0.2 must be read against true human hours, not estimated ones: the mechanical saving is ~31 % of the figure shown there (the adjusted rule, § 2b), and the argument rests on bus factor, parallelism and retrieval.
Caveats that bound the claim F (timing report § Caveats): elapsed is not effort (an operator away for a day inflates the mean, not the median); commits are batched across two machines; the AISA timeline stamps are the agent's own approximate times (108 date-only stamps unused); the human estimate is the agent-written financial card whose scope varies (L1 diagnosis versus full CR development), not an independent human number. P The one measurement that would settle it is ten tickets timed by a skilled person with a stopwatch, alongside the same ten on the platform.
3. Where v2 adds value, and how each claim is measured P
| Lever | Mechanism | Estimate today | Replaced by (ledger metric) |
|---|---|---|---|
| Skills for the recurring shapes | 35 catalogued shapes; the mechanical part of a ticket runs as a governed skill | ~625 estimate-hours per six months over the mined archive (1 921 → 1 296 h), i.e. ~1 250 h/year on the analysis authors' estimates — about 390 h/year of true skilled time at the adjusted factor (31 % of the estimate, § 2b), varying by domain and person (Data and Information Module § 0.2) | human minutes per case by skill; reuse hits |
| Retrieval that fires | precedent lookup at triage, mechanics instead of memories, measured hit-rate | v1's 10–60× cached-vs-rediscovered ratio applied to the share of cases that have a precedent — unknown share today | triage hit-rate on held-out cases (Agents Memory § 7.2); time-to-first-relevant-item |
| No retraction cycles | isolated verifier re-derives the mechanism | 3 known cycles in v1; each a day of work plus credibility | retraction count judged from case records |
| Parallel cases, any operator | web platform, many viewers, one controller, durable cases | v1: one case per machine, two machines | concurrent cases; throughput per operator |
| Bus factor down | knowledge in skills and memory that any operator's agent executes | thin shapes: 7 today; 35 of 35 shapes must close by ≥ 2 operators with the skill (D115) | count of thin shapes and shapes closed by ≥ 2 operators with the skill |
| Switch cost down | the agent works the routine stretch; the human decides at gates, from the Decisions queue, when it suits | median 3 turns → gate decisions only | human turns per case; time between gate request and decision |
| Waiting made visible | held batches are cases with an age | 12 weeks invisible | held approved writes: count and age |
| Configuration executed, not described | the configuration module deploys as far as the user chooses | v1 stops at files; the PC sibling did 9951 end to end on dev | hours and days per configuration stage run, established on the first slice |
4. What v2 costs P
| Cost | Shape | Control |
|---|---|---|
| Build | two phases and two waves: T0 contracts are phase 1; T1 platform experiment + T2 configuration slice are wave 1; T3 is wave 2 (Delivery); the runtime is the largest unproven item (Agent Runtime) | no dates, named staff or headcount in this wiki (Non-Goals N-6) — the estimate is the development manager's |
| Run — provider usage | v2 uses more provider capacity per case by construction: lateral consults, the isolated verifier (~2× the reasoning of a case), eval re-runs | per-case and per-profile budgets; eval batching; cheap routes for mechanical stages (Metrics § Spend envelope) |
| Run — operations | one more compose unit on the QA host, one schema, Vault secrets, the six LT_USER_ROLES rows the platform adds (D134), two Bulstrad IT network decisions |
Software Architecture |
| Governance | prompts, skills and memory go through test → publish; someone holds the publisher role per stage | the role is a permission, not a headcount (Trust and Data § 3) |
5. How the ROI number is produced P
- Baseline, now, on v1 — before any platform exists: the mined effort distribution (Data and Information Module § 0.2), the v1 human-turn median, the thin-shape count, the triage hit-rate of grep over the current
memory/on 30 held-out tickets (Stage Precipitation § 0). These are one afternoon of work and they are the denominator. - First slice — the abacus-only configuration deployment and the first ten Support cases on the platform: the ledger yields human minutes, provider usage, gate latency, reuse hits per case (Metrics). Compare to the baseline per ticket shape, not in aggregate.
- Per skill — hours now → hours with skill, measured, replacing the inferred columns of the catalogue; the bus-factor count re-taken.
- The numbers — in hours:
human hours saved(handling + switch cost) per quarter,retractions avoided,waiting days avoided,defects after interruptions, and beside them provider usage as a control; no conversion of hours into money (D52/D117). Written into Metrics § Measurement log by the tools, never by hand.
P The protocol for the two missing inputs (Challenge Rounds § R2 § 10): ten consecutive real interruptions across at least three developers and both channels, four timestamps each written as they happen — T0 noticed · T1 ticket work starts · T2 ends · T3 back in flow (first meaningful edit or commit, cross-checked against git); switch cost = (T1−T0) + (T3−T2), handling = T2−T1, medians and p75 reported. The same ten tickets test the 25 % rule per type: (T2−T1) / human_est_h from the analysis card; the rule holds for a type when the median is inside 0.15–0.35, otherwise that type gets its own factor and the catalogue saving of Data and Information Module § 0.2 is recomputed; first on the three largest types (data correction, transfer, incident = 71 % of tickets). Owner: development manager; two calendar weeks. There is no hourly rate — hours are the unit (D52). C the ten rows exist.
6. The half page P
P The owner's purpose. Put the recurring configuration, development and support work currently carried by six people into software that the customer can use directly. The hours below support recovered capacity and reduced dependence on particular knowledge holders; they do not measure six complete workloads or promise six posts removed. Configuration preview time and accepted-delivery time are separate measurements; compare ticket closure with accepted delivery, not with a preview. This page explains the operating choice for the owners, not a purchase offer.
For the CEO. Everything here is either measured in v1 or named as the measurement that will replace the estimate.
Does the agent make implementation mistakes? Yes. Therefore tests, evals and an isolated verifier gate every change — the platform is built around that answer.
What is measured in v1: 2.16× on the UAT Health file; 2.5× personally; −24 % on the CR schedule, technical phases only; ~7 ticket folders per working day; one person leading 33 of 35 tag-bearing ticket shapes (D115). On support tickets v1 frees the skilled person's own time but not the calendar — about 1.5 hours of agent time against about 40 minutes of true human time, the operator busy for a quarter of it (adjusted figures, § 2b) — while writing the analysis and the notes; its value is on the tickets that person is not available for, has never seen, or cannot run ten at once. Times vary by domain and person.
What v2 changes, in three lines: 1. About half of the tickets in the effort row are trivial or routine (156 of 331); two thirds carry a database write in their analysis (220 of 333; the desk itself executed the write in 42 % of the 345 analysed cases and left 23 % as proposed SQL for Bulstrad IT or an operator — Typical cases § 0) and are the shapes the skills package — ~1 250 hours a year of mechanical work by today's inferred numbers, planning against 30 minutes of human time per ticket of non-agent work (D69), replaced by measured ones after the first month. 2. Thin ticket shapes (7 today) become executable by any operator's agent — the bus factor, which is the cost that does not show in hours (D115). 3. Product configuration is executed and deployed by the platform under human gates, instead of being described in documents — proven once on 9951, now with sessions, audit and deployment.
What it costs: a build in two phases and two waves, with no dates in this document; more provider usage per case, budgeted; one more service on the QA host and two network decisions from Bulstrad IT.
What we ask now: two calendar weeks of the ten-interruption measurement, and the first configuration slice on QA to replace every estimate on this page with a ledger number.