Post-mortem
Failure atlas
Every defect that actually happened here, what it cost, and the check that now holds it shut. Written by the suite, not by us.
SourceGenerated from this system's own regression suite
As ofregenerated on every verify run
LicenceOurs — published because a score with no error bar is an opinion
Size70 defects
Every defect that actually happened here
| The check that now holds it shut | What it cost when it happened | Guarded |
|---|---|---|
| R10 HOUSE carries no operator identity | every agent judged one reader | yes |
| R11 no operator dimension in any rubric | 14% of Track P scored the reader | yes |
| R12b verdict vocabulary consistent | three different vocabularies coexisted | yes |
| R12 no rejection-flavoured verdicts | BUILD/VALIDATE/WATCH/KILL | yes |
| R13 a low base rate does not crush differentiation | 8% prior capped everything at 14.7%, so great and terrible looked the same | yes |
| R14 every track runs BASERATE | the default track had no outside view at all | yes |
| R15 trailing-content JSON recovered | ARBITER_P failed 2 of 4 runs | yes |
| R16 truncation retried | DEFENDER + FINANCIER corrupted verdicts | yes |
| R17 completeness recorded | degraded runs looked identical to clean ones | yes |
| R18 stale-code detector exists | a 15-stage pipeline ran as if it were 17 | yes |
| R19 structured enum values are defined | gates guessed their own flag vocabulary | yes |
| R1 extractor ceilings above measured floor | 786 dead calls | yes |
| R20 gate count matches flag vocabulary | CONCERN-DARK mapped to no gate | yes |
| R21 scoring legends generated from live weights | duplicate L7, duplicate P7, sum 91 | yes |
| R22b rubric_for resolves every track | L and P silently fell back to V | yes |
| R22 snapshots record the rubric that ran | 5 L/P snapshots claimed Track V weights | yes |
| R23 QUANT sized above its observed output | double cost + run marked not comparable | yes |
| R24 time-to-revenue is reported, not scored | real businesses failed for being slow | yes |
| R25 rubric versions moved with the weights | w2.0 would have described two rubrics | yes |
| R26 no reader-constraint welded into a prompt | one profile scored into every business | yes |
| R27 every track pre-registers a demand test | L and P shipped verdicts with no falsifiable test for months | yes |
| R28 model output normalised at the intake boundary | 3 verdicts persisted with demand_test=null; 86 was a BUILD with no test at all | yes |
| R29 score variance measured against the live rubric | the calibration strip would have rendered an invented interval | yes |
| R2 arbiter ceiling sized to evidence load | $2.81 + lost verdict | yes |
| R30 the registry drives the harvester | five disabled sources were fetched nightly; then 1,979 registered sources could not be fetched at all | yes |
| R31 every enabled source category has an extractor | 1,979 discovered sources would all have used the generic prose lens at batch 1 | yes |
| R32 evidence-role sources are excluded from harvest | 735 licence/entity/lien registries would have been swept as if they were pain signal | yes |
| R33 no prompt narrows the product to one output shape | the procurement lens told itself to ignore construction, physical goods and site work | yes |
| R34 extractor strength is stored, not discarded | every signal ever extracted lost the model's own 1-5 confidence rating | yes |
| R35 a signal dropped at synthesis records the reason | 96 orphaned signals, and no way to tell a failure from a gap | yes |
| R36 every enabled source resolves to a fetch URL | the registry and the harvester described different worlds | yes |
| R37 StackExchange fetches question bodies | 183 demand sources registered against a filter that returns no question text | yes |
| R38 every console page renders for a signed-in operator | a page was shipped to the founder that 500'd on every load | yes |
| R39 no module defines the same name twice | _mask and the env writer were each defined twice; the wrong copy won silently | yes |
| R3b stale claims are reclaimed | candidate orphaned 208 min | yes |
| R3 queue claim is atomic | ~$4 + two verdicts for one candidate | yes |
| R40 the public surface renders and all ten templates compile | step 1 would have shipped with the same untested hole the console had | yes |
| R41 the surface reads only through the spine | a page that reaches into a candidate row makes core/fields.py decorative | yes |
| R42 every source carries the role its category implies | 97 licence registries were queued to be harvested as pain signal through a NULL | yes |
| R43 discovery.register assigns role at write time | role was set once by a migration and every later discovery arrived NULL | yes |
| R44 the evidence layer does not produce fees or lead times | a fabricated licence fee on a page someone acts on | yes |
| R45 truncated counts are labelled as floors | 6,000 reported as a licensee count when every registry had been page-capped | yes |
| R46 migration numbers are unique | two files were both numbered 012; a runner would apply one and skip the other | yes |
| R47 jurisdiction is derived from the portal, not asserted | 181 Canadian datasets were labelled US and every city portal lost its state | yes |
| R48 abbreviation matching cannot mis-file a portal | a substring match would have put Chicago, Cambridge and Calgary in California | yes |
| R49 no regression case id appears twice | a re-run patch script duplicated six cases and the suite reported a false 56/56 | yes |
| R4 budgets reset per candidate | silent stage skipping | yes |
| R50 a jurisdiction lookup reaches the jurisdictions inside it | the two real NY electrical registries were missed while mold contractors were returned | yes |
| R51 the licence matcher keeps real licences and drops fuzzy matches | Chiropractor was returned for an electrical lookup; Barbers was dropped for barber | yes |
| R52 an absent licence is reported as a finding | a real 'this trade is not licensed here' would have read as a broken lookup | yes |
| R53 observed_fee returns a distribution or nothing | a bare median would carry confidence nobody measured | yes |
| R53 observed_lead_time returns a distribution or nothing | a bare median would carry confidence nobody measured | yes |
| R54 every characteristics-bound field declares its provenance | a field pointed at a source that produced nothing and said nothing about it | yes |
| R55 an unmatched occupation returns unknown, not unlicensed | physicians were reported as an unlicensed occupation in all 50 states | yes |
| R56 demand density resolves across unrelated verticals | coverage across verticals was asserted in prose and never tested | yes |
| R57 the gate resolves to the licensed occupation, not the biggest headcount | physicians' offices returned a 6-state gate against a true 52 | yes |
| R58 labour supply and wage floor resolve from the industry | O4 and O6 had no producer and would have scored on nothing | yes |
| R59 pages showing licence data carry the CareerOneStop attribution | a licence-data page without attribution breaches the grant on its face | yes |
| R5 gates cannot kill | DealScope, TradePath, RootOptics all wrongly blocked | yes |
| R60 a generated page lists what it cannot tell you | an omitted gap reads as nothing-to-report | yes |
| R61 an unmatched licence title is labelled unmatched, never absent | California and Florida were shown as having no electrical gate at all | yes |
| R62 occupation stemming keeps its floor | electrolysis instructors were counted as an electrician licence gate | yes |
| R63 every harvestable source has a parser | 1,244 sources would have fed truncated JSON to a paid extractor, nightly | yes |
| R64 an unknown payload yields nothing, not a fake record | a 3,000-character JSON fragment was emitted as a signal | yes |
| R65 StackExchange parsing keeps the question body | 183 demand sources were fetched with bodies and parsed without them | yes |
| R6 field copy runs before branches | REMEDY build spec silently lost | yes |
| R7 no schema/FIELD_MAP drift | BRIEF thesis + CAPITAL_MAP routes lost | yes |
| R8 no unintended field collisions | assumptions/uncertain_about/note erased | yes |
| R9b numeric columns coerced | a string answer crashes a completed run | yes |
| R9 persist generates SQL from its mapping | $6.36 lost after all 22 stages ran | yes |