What we predicted, and what happened
Predictions are frozen with a date and a falsifying threshold before the outcome is known. Nothing is edited afterwards, and the misses stay up.
Live pass marks
| Venture | Actual | Metric | Visitors | Resolves | Left | Outcome |
|---|---|---|---|---|---|---|
| GA4 Fix | 0 of 3 | customers | 10 | 2026-08-21 | -44d | VOID |
| AD Sync Fix | 0 of 3 | customers | 4 | 2026-08-21 | -44d | VOID |
| Fail2ban for journald | 0 of 4 | customers | 6 | 2026-08-21 | -44d | VOID |
| DealScope | 0 of 3 | customers | 36 | 2026-08-30 | -35d | VOID |
Visitors is external traffic only — our own verification hits are flagged and excluded. 4 of these 4 are VOID: registered, never run, and contributing nothing to any calibration in either direction. Deadline voided by founder direction 2026-08-06. The test was never run: 18 external visitors across all four probes, no tagged link ever shared, and no outreach performed against the registered threshold. Recording FAIL would enter "demand was tested and absent" into the calibration record when nothing was tested. Void contributes nothing in either direction. 3 of these 4 have had fewer than 25 visitors, so a threshold missed on the deadline would record that distribution was never attempted, not that demand was absent. An unrun test is not a failed one and will not be filed as one. 3 carry no freeze record: the metric and threshold are set, but nothing captured what was committed to or when. (GA4 Fix, AD Sync Fix, Fail2ban for journald.)
Frozen, not yet launched
| # | Candidate | Verdict | Score | Falsifying metric | Window |
|---|---|---|---|---|---|
| 244 | On-Site Mobile Diesel Repair for West Texas Owner-Operators | WATCH | 48 | Paid mobile repair jobs closed at west Texas truck stops/yards | 45d |
| 243 | Calibrated Base-Rate Data for Business Buyers | BUILD | 62 | Number of prepaid orders (cash collected, not signups or waitlist), plus the split between $149 and $399 acceptance, plus count of repeat or multi-dea | 45d |
| 242 | Unbilled Change-Order Watchdog for Subcontractors | BUILD | 67 | (a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b) | 60d |
| 94 | Deal Intelligence for Used Heavy-Duty Diesel Trucks | VALIDATE | 59 | Stripe payments completed for paid VIN reports from strangers (free proof reports do not count) | 14d |
| 31 | Skimmable Video Summaries for Linear Recordings | WATCH | 41 | Completed paid orders from strangers (no friends, no refunds, Fiverr or Stripe cleared) | 7d |
| 32 | Supplier Recall Lookup for Injectable Kit Assemblers | WATCH | 42 | Completed Stripe payments at $99 from firms with no prior relationship to the operator | 14d |
| 89 | Pre-LOI Screening for SBA Business Buyers | BUILD | 74 | Completed Stripe charges at $249 for the Pre-LOI Deal Screen, counting only strangers (no prior relationship), with refunds subtracted | 14d |
| 90 | Publicly Scored Daily Volatility Ranges for Traders | BUILD | 61 | Paid $19 two-week beta subscriptions collected via Stripe payment link | 14d |
| 88 | Journald-to-Fail2ban Filter and Jail Generator | WATCH | 40 | Completed Stripe charges of $15 from strangers (not refunded, log sample actually received) | 10d |
| 87 | Azure AD Connect Password Sync Failure Diagnostic | WATCH | 42 | (a) count of qualifying live PHS-failure threads found in 14 days, and (b) count of $19 Stripe payments collected | 14d |
| 86 | GA4 Symptom-to-Root-Cause Diagnostic | BUILD | 62 | Paid orders from strangers, plus post-delivery confirmations that the named cause was correct | 14d |
| 25 | Tender Radar for School Transport Route Bids | WATCH | 28 | Count of distinct operators with a live £15/mo Stripe subscription (card charged, not a verbal yes or a free-trial signup), AND count of route-level o | 21d |
| 21 | Route Cost & Feasibility Checker for School Transport Bidders | WATCH | 37 | Distinct UK operators who complete a £39 Stripe payment after hitting the third-calculation gate. | 30d |
| 19 | Eligibility Check for NIH Grant Mechanisms | WATCH | 15 | Count of distinct institutions that pay $49 via Stripe for a single submission audit (not replies, not demo requests, not 'send me pricing'). | 21d |
| 15 | Route Costing Tool for School Transport Contract Bidders | WATCH | 24 | Stripe charges completed at £49 from distinct operators who supply a real route sheet | 21d |
| 2 | Line-Level Code Review Coverage Auditor | WATCH | 30 | Number of qualified paid-intent conversations: a named engineering leader at a 30+ engineer org who (a) installed the Action on a private repo or rece | 30d |
Each of these was recorded at the moment the verdict was issued. None has been run. They are published now so that what was predicted cannot be revised after the fact.
Verdicts carrying no test
| Decided verdicts | With a usable frozen test | With none |
|---|---|---|
| 36 | 16 | 20 |
20 of 36 decided verdicts carry no usable pre-registered test. The cause is ours and it is documented: the arbitration agent nested the test object inside its dimensions block on 8 of 10 runs, and the pipeline read it from where it was expected to be rather than normalising it at the boundary, writing NULL. The payload is now normalised at a single intake point and a regression case holds it there. The affected verdicts need re-running, not patching — editing them would destroy the calibration record, which is the only thing here worth having.
Evidence behind the verdicts
| Verdicts issued | Scored with no linked signal | Demand signals held | Never processed |
|---|---|---|---|
| 36 | 19 | 170 | 96 |
mention_count — whether a problem was raised once or fifty times — feeds three confidence functions at up to 20 points. It was a stored number that defaulted to 1 at creation and was updated only by a clustering step that never ran, so 19 of 36 verdicts were scored as though one person had raised the problem when no signal was linked at all. The database now derives it from what is actually attached. These verdicts are named rather than rescored: editing a frozen verdict destroys the calibration record it exists to be. 96 of 170 demand signals have never been processed into a candidate — not failed, never reached.
Has the rubric been right?
No prediction has resolved yet, so the rubric has never been graded. The first four resolve 21 and 30 August 2026. Until then every score this system publishes is uncalibrated against its own outcomes, and saying so is the difference between this and a confidence number with nothing behind it.
What we have learned from being wrong
| Predictions registered | Still running | Void | Resolved | Can calibrate? |
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | not yet |
No prediction has resolved yet, so the rubric has never been graded. The first four resolve 21 and 30 August 2026. Until then every score this system publishes is uncalibrated against its own outcomes, and saying so is the difference between this and a confidence number with nothing behind it. Every verdict is written into a prediction ledger at the moment it is issued — the score, the confidence, the dimension scores as they stood, and the rubric version that produced them — and the outcome is written back automatically when a test resolves. Nothing in this category does this: fourteen tools were surveyed and none publishes a calibrated error bar, a pre-registered test, or its own failure history, let alone uses them to re-weight what it scores.
A frozen prediction is attributable to the rubric. D6 and D7 carry the loosest dimension-level agreement (sd 3.85 / 3.67) and the weights damp them; the total holds inside +/-3.1 at 95%.