The record

What we predicted, and what happened

Predictions are frozen with a date and a falsifying threshold before the outcome is known. Nothing is edited afterwards, and the misses stay up.

InstrumentARBITER
Rubricw3.0-20260804
Interval±3.1 at 95%
Measuredn=10, sd=1.6, 2026-08-04
frozen before the outcome was known

Live pass marks

VentureActualMetricVisitorsResolvesLeftOutcome
GA4 Fix0 of 3customers102026-08-21-44dVOID
AD Sync Fix0 of 3customers42026-08-21-44dVOID
Fail2ban for journald0 of 4customers62026-08-21-44dVOID
DealScope0 of 3customers362026-08-30-35dVOID

Visitors is external traffic only — our own verification hits are flagged and excluded. 4 of these 4 are VOID: registered, never run, and contributing nothing to any calibration in either direction. Deadline voided by founder direction 2026-08-06. The test was never run: 18 external visitors across all four probes, no tagged link ever shared, and no outreach performed against the registered threshold. Recording FAIL would enter "demand was tested and absent" into the calibration record when nothing was tested. Void contributes nothing in either direction. 3 of these 4 have had fewer than 25 visitors, so a threshold missed on the deadline would record that distribution was never attempted, not that demand was absent. An unrun test is not a failed one and will not be filed as one. 3 carry no freeze record: the metric and threshold are set, but nothing captured what was committed to or when. (GA4 Fix, AD Sync Fix, Fail2ban for journald.)

written down before the answer existed

Frozen, not yet launched

#CandidateVerdictScoreFalsifying metricWindow
244On-Site Mobile Diesel Repair for West Texas Owner-OperatorsWATCH48Paid mobile repair jobs closed at west Texas truck stops/yards45d
243Calibrated Base-Rate Data for Business BuyersBUILD62Number of prepaid orders (cash collected, not signups or waitlist), plus the split between $149 and $399 acceptance, plus count of repeat or multi-dea45d
242Unbilled Change-Order Watchdog for SubcontractorsBUILD67(a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b)60d
94Deal Intelligence for Used Heavy-Duty Diesel TrucksVALIDATE59Stripe payments completed for paid VIN reports from strangers (free proof reports do not count)14d
31Skimmable Video Summaries for Linear RecordingsWATCH41Completed paid orders from strangers (no friends, no refunds, Fiverr or Stripe cleared)7d
32Supplier Recall Lookup for Injectable Kit AssemblersWATCH42Completed Stripe payments at $99 from firms with no prior relationship to the operator14d
89Pre-LOI Screening for SBA Business BuyersBUILD74Completed Stripe charges at $249 for the Pre-LOI Deal Screen, counting only strangers (no prior relationship), with refunds subtracted14d
90Publicly Scored Daily Volatility Ranges for TradersBUILD61Paid $19 two-week beta subscriptions collected via Stripe payment link14d
88Journald-to-Fail2ban Filter and Jail GeneratorWATCH40Completed Stripe charges of $15 from strangers (not refunded, log sample actually received)10d
87Azure AD Connect Password Sync Failure DiagnosticWATCH42(a) count of qualifying live PHS-failure threads found in 14 days, and (b) count of $19 Stripe payments collected14d
86GA4 Symptom-to-Root-Cause DiagnosticBUILD62Paid orders from strangers, plus post-delivery confirmations that the named cause was correct14d
25Tender Radar for School Transport Route BidsWATCH28Count of distinct operators with a live £15/mo Stripe subscription (card charged, not a verbal yes or a free-trial signup), AND count of route-level o21d
21Route Cost & Feasibility Checker for School Transport BiddersWATCH37Distinct UK operators who complete a £39 Stripe payment after hitting the third-calculation gate.30d
19Eligibility Check for NIH Grant MechanismsWATCH15Count of distinct institutions that pay $49 via Stripe for a single submission audit (not replies, not demo requests, not 'send me pricing').21d
15Route Costing Tool for School Transport Contract BiddersWATCH24Stripe charges completed at £49 from distinct operators who supply a real route sheet21d
2Line-Level Code Review Coverage AuditorWATCH30Number of qualified paid-intent conversations: a named engineering leader at a 30+ engineer org who (a) installed the Action on a private repo or rece30d

Each of these was recorded at the moment the verdict was issued. None has been run. They are published now so that what was predicted cannot be revised after the fact.

the gap in our own record

Verdicts carrying no test

Decided verdictsWith a usable frozen testWith none
361620

20 of 36 decided verdicts carry no usable pre-registered test. The cause is ours and it is documented: the arbitration agent nested the test object inside its dimensions block on 8 of 10 runs, and the pipeline read it from where it was expected to be rather than normalising it at the boundary, writing NULL. The payload is now normalised at a single intake point and a regression case holds it there. The affected verdicts need re-running, not patching — editing them would destroy the calibration record, which is the only thing here worth having.

the second gap in our own record

Evidence behind the verdicts

Verdicts issuedScored with no linked signalDemand signals heldNever processed
361917096

mention_count — whether a problem was raised once or fifty times — feeds three confidence functions at up to 20 points. It was a stored number that defaulted to 1 at creation and was updated only by a clustering step that never ran, so 19 of 36 verdicts were scored as though one person had raised the problem when no signal was linked at all. The database now derives it from what is actually attached. These verdicts are named rather than rescored: editing a frozen verdict destroys the calibration record it exists to be. 96 of 170 demand signals have never been processed into a candidate — not failed, never reached.

graded against outcomes, not opinions

Has the rubric been right?

63 predictions frozen · none resolved yet
The first four resolve 21 August 2026. Nothing is graded before then.

No prediction has resolved yet, so the rubric has never been graded. The first four resolve 21 and 30 August 2026. Until then every score this system publishes is uncalibrated against its own outcomes, and saying so is the difference between this and a confidence number with nothing behind it.

the loop, closed

What we have learned from being wrong

Predictions registeredStill runningVoidResolvedCan calibrate?
0000not yet

No prediction has resolved yet, so the rubric has never been graded. The first four resolve 21 and 30 August 2026. Until then every score this system publishes is uncalibrated against its own outcomes, and saying so is the difference between this and a confidence number with nothing behind it. Every verdict is written into a prediction ledger at the moment it is issued — the score, the confidence, the dimension scores as they stood, and the rubric version that produced them — and the outcome is written back automatically when a test resolves. Nothing in this category does this: fourteen tools were surveyed and none publishes a calibrated error bar, a pre-registered test, or its own failure history, let alone uses them to re-weight what it scores.

A frozen prediction is attributable to the rubric. D6 and D7 carry the loosest dimension-level agreement (sd 3.85 / 3.67) and the weights damp them; the total holds inside +/-3.1 at 95%.