Underwriting · V-track · 06 August 2026

Unbilled Change-Order Watchdog for Subcontractors

GAP: subs lose billable margin because change-order requests arrive informally (email, verbal, field notes) and get executed before being priced or signed, and by the time billing happens the paper trail is gone or disputed. WHO HAS IT: specialty trade subcontractors in the $5-50M revenue band, where change orders are common but back-office capacity to police every email thread is thin. WHAT IS SOLD: a monitoring age…

67 ± 3.1 BUILD
REWORK
PROVE
BUILD
rubric w3.0-20260804 · interval ±3.1 at 95% (n=10, sd=1.6, measured 2026-08-04)
Stages run26
Cost to produce$7.11
Wall clock34 min
Confidence55/100

What an analysis cost to produce belongs beside it. A reader deciding whether to trust a verdict is entitled to know whether it came from twenty-six stages or one, and nothing else in this category will tell them.

every charge, every rebuttal, every ruling

The case against it

ChargeRebuttalRuling
Signal lives outside email, so the product structurally can't see most of the leakCONCEDED — Conceded. The thesis itself flags this as an unstated assumption, and no evidence in the record establishes what share of change-order-triggering commupheld — Conceded with no counter-evidence; the share of change-order-triggering communication that touches email is unmeasured and is the load-bearing assumption of the entire product — this is the
Flagging without closing the loop doesn't recover the margin, so ROI attribution collapsesCONCEDED — Conceded as a design flaw in the current scope, though it is fixable rather than fatal. The value-based pricing story ('dollars of margin recovered') partial — Real as scoped, but the fix is known-in-kind and buildable (draft COR generation, e-signature chase, GC submission) and incumbents already ship the downstream half; this constrains pricing m
Mailbox-wide access request is a trust barrier the buyer persona won't clear quicklyPartially rebutted. Full OAuth mailbox access is not the only viable ingestion architecture: a dedicated forwarding address or shared CC/BCC alias (thdismissed — Forwarding/CC-alias ingestion is an established pattern in this exact category and avoids broad OAuth entirely; the residual issue is behaviour change in field/PM staff, which is an adoption
Precision is unproven at launch and trade-correspondence NLP is unusually noisyCONCEDED — Conceded outright. The data_moat section itself admits 'a new entrant starts with no verified precision/recall numbers and no proof of dollars actuallpartial — Precision genuinely unproven and alert fatigue is the likeliest churn mechanism — but the concierge design sidesteps it entirely for the first cohort by putting a human in the loop, which is
Incumbents can bolt on an email-ingestion layer faster than the entrant can build a data moatPartially rebutted. The incumbent_weakness section provides evidence of a structural, not just resource-based, disincentive: Clearstory's positioning dismissed — Banned reasoning territory: 'a funded competitor could copy this' is true of nearly every software wedge. The structural disincentive (Clearstory's whole pitch is moving work off email) is a
Escalating a flag to the GC creates relationship risk the sub may not want automatedPartially rebutted by the product's own design constraint: the autonomy spec treats GC-escalation as a 'needs_human' judgment call rather than an autopartial — The tool as specified does not auto-escalate, so the stated failure mode is not what's proposed — but the deeper tension (subs deliberately under-formalising paper trails to protect repeat w
The buyer may treat unbilled COs as an accepted cost of doing business, not an active problem to solvePartially rebutted. The Levelset channel evidence shows subs actively posting about this exact problem in their own words on public forums — unpromptepartial — Unprompted Levelset posting and $28M+ of venture funding into the category cut against full normalisation, but the conversion rate within the $5–50M band is genuinely unknown and is a legiti
Proposed channels don't scale past a handful of relationship-driven dealsCONCEDED — Conceded. The candidate's own channel notes admit the one scalable channel is already occupied: 'Buyers actively search this category, but Siteline, Rpartial — Conceded on SEO occupancy, but the ruling overreaches: ASA/CFMA chapter rolls, licence registries and the G2 category are enumerable lists addressable by outbound and paid placement — that i
Thin back-office capacity — the stated reason subs need the tool — is also the reason they can't operate itCONCEDED — Conceded. The buyer profile is explicitly defined by thin back-office capacity unable to police email manually, and the same team is required to triagupheld — Genuine design contradiction as specified. The only resolutions are high-confidence-only alerting or a done-for-you triage service (i.e. sell the concierge permanently, not as a stepping sto

A separate agent argued against this idea, a second answered, a third ruled. 5 of 9 charges were conceded rather than defended. Published in full because a score with the objections removed is a advertisement, and because the objections are usually more useful than the verdict.

dimension by dimension

How it scored

DimensionScoreReasoning
D170Unbilled/disputed change orders are a documented, recurring margin leak in trade subcontracting (24-day sign-to-submit lag, work performed before pricing), and subs post
D272Strong adjacent payment evidence: Knowify charges $99–$329/mo, Clearstory raised $28M+ on this exact workflow, eSUB $23.5M, Siteline at 10,000+ projects — subs demonstrab
D378Buyers are exceptionally addressable: state contractor licence registries, ASA and CFMA chapter rolls, the G2/Capterra 'Change Order Management' category with existing co
D430At v1 the wedge is an email classifier any funded incumbent can bolt on as a connector; the outcome-linked labeled corpus is a real but slow-accruing moat that does not e
D562The concierge version is genuinely cheap ($6.5k cash, 3 days to stand up, human-in-the-loop), but the automated product needs a precision-tuned classifier over noisy trad
D662Recurring per-firm subscription is a clean model with obvious annual contracts and expansion by project volume, but the stated value-based pricing ('dollars of margin rec
D758Tens of thousands of US specialty subs in the target band and a portable, non-English-market gap put a Clearstory-scale ($10–50M ARR) outcome inside the range, but only i
named, priced, and dated

Who already does this

CompetitorPricingFundingLaunchedOverlap
ClearstoryFree Basic tier; Standard/Professional add per-user + per-prSeed $5.3M, Series A $7M (Cloud Apps Capital),~2018 (per '6 years ago' rpartial
SitelineNot publicly listed; billing-software positioningNot disclosed in resultsPre-2020 (10,000+ projectspartial
KnowifyCore $99/mo (annual, 1 user); Advanced $329/mo (annual, up t$8.45M total, last raise $3M (2023)2012partial
eSUB CloudNot publicly listed$23.5M total, last raise $2M (2023)2008partial
GCPayNot publicly listedNot disclosedNot disclosedadjacent
Planyard14-day free trial, paid tiers not disclosed in resultsNot disclosedActive as of 2025adjacent
RhumbixNot publicly listedNot disclosedNot disclosedpartial
KynectionNot publicly listedNot disclosedNot disclosedadjacent
Mirage Metrics (AI agent layer for PM systems)Not disclosedNot disclosedNot disclosedpartial

Where the buyers actually are

ChannelWhy it reaches them
Levelset Payment Help community (payment-help/question threads)
CFMA local chapters + Connection Café
Long-tail organic search ("unbilled change order tracking," "change or
ASA (American Subcontractors Association) national/state chapters
G2/Capterra "Change Order Management Software" category
what stands in the way

Regulatory gates

GateFinding
G1Email monitoring and flagging is lawful. No prohibition on reading a customer's own email traffic with permission, analyzing it, and surfacing alerts. SaaS compliance (data handling, securit
G2Core operation is automated email monitoring and rule-based flagging. Requires no founder presence after deployment. Customer success, onboarding, and support are staffable. The agent itself
G3Existing paid competitors prove willingness to pay. Buildertrend, Knowify, Siteline, eSUB and others charge for change-order management modules. Subcontractors already budget for constructio
G4Email parsing, NLP-based change-order detection, and alert generation are all standard SaaS capabilities today. Requires no scientific breakthrough. LLMs can classify construction correspond
G5Subcontractors are reachable through multiple channels: construction software marketplaces (Appstore integrations with Buildertrend, Knowify), trade associations (AGC, specialty trade groups
written before the outcome is known

The pre-registered test

TermValue
days60
offerTwo-part paid pilot with 10 named specialty subs in the $5–50M revenue band, recruited from ASA and CFMA local chapter rosters plus direct outreach off state contractor licence registries. Part A (evidence): for a $250 p
price249
metric(a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b) Paid conversion: number of the 10 recruited firms that pay $249 up fr
channelASA and CFMA local chapter membership rosters and meetings (dues paid, in-person attendance), supplemented by direct outbound to state contractor licence registry listings in two metros
cost usd9500
falsifiesIf (a) misses, the assumption that email is where the change-order signal lives dies outright — the product cannot see the leak it claims to intercept and must be re-scoped to SMS/voice/field-ticket capture, which is a d
threshold(a) ≥50% email-traceability across the pooled audit AND ≥40% in at least 6 of the 10 individual firms; (b) ≥5 of 10 recruited firms pay the $249; (c) ≥3 of those payers authorise month two.

Recorded at the moment the verdict was issued and not editable afterwards. If this is launched, the result lands on the ledger whether it passes or fails.

and what moves it forward

Where this idea is

Phase nowAnalysed — Underwritten, with the argument against it on the record.
What you do hereRead the case against it first. An upheld charge you cannot answer is the verdict, whatever the score says.
To leave this phaseYou have read the upheld charges and decided the idea survives them.
Gate statusThis gate is a judgement, not a query. The system will not rule on it and will not pretend to — you decide, and the reason is recorded.
Next phaseValidating — A pre-registered test is live and running.

This gate is a judgement rather than a query, so the system states it and refuses to rule on it. Pretending software can decide whether a business "can take money from somebody who is not you" would make every gate on this site meaningless. Advancing an idea needs its link — the one handed back when it was submitted. Founder-owned ideas are advanced from the console. See the whole pipeline.

What to do in this phaseWhat it provesFrom which part of the analysis
Answer the upheld charge: Signal lives outside email, so the product structurally can't see most of the leakThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the partial charge: Flagging without closing the loop doesn't recover the margin, so ROI attribution collapsesThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the partial charge: Precision is unproven at launch and trade-correspondence NLP is unusually noisyThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the partial charge: Escalating a flag to the GC creates relationship risk the sub may not want automatedThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the partial charge: The buyer may treat unbilled COs as an accepted cost of doing business, not an active problem to solveThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the partial charge: Proposed channels don't scale past a handful of relationship-driven dealsThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration
Answer the upheld charge: Thin back-office capacity — the stated reason subs need the tool — is also the reason they can't operate itThe verdict survives its strongest objection, or it does not and you have learned that before spending.arbitration

Every step traces to a field this idea's own underwriting produced — not generic best practice, which is free everywhere. 0 of 7 complete. Mark them off in the console.

and what did not complete

How this was produced

MeasureValue
Cost to produce$7.11
Wall clock34 minutes

A verdict produced by 22 of 23 stages is not the same artefact as one produced by all of them, and which stages failed was recorded on every run and shown nowhere until now. If a stage that feeds a section died, the section came from somewhere else or nowhere — and you are entitled to know which is in front of you before you act on it.

The money

price pointNone
current spendNone
funding routepresale
revenue modelNone
churn monthly pctNone
cash to first dollar6500
marginal cost per unit38

What it costs to start

capital blockedFalse

What has to be built

data moatA growing labeled corpus linking specific email phrasing patterns (trade-specific, regional, per-customer) to confirmed outcomes - which flags turned into real signed/priced change orders and recovered dollars, versus which were noise. This outcome-linked dataset is what lets the classifier's precision improve over time and lets the vendor credibly price against margin recovered; a new entrant starts with no verified precision/recall numbers and no proof of dollars actually recovered per flag type, both of which take real customer usage over months to accumulate.
componentsEmail ingestion & mailbox sync (Gmail/O365 OAuth, polling, thread reconstruction): risk: med; units: 6; Change-order-signal classifier (distinguish CO request from routine correspondence): risk: high; units: 12; Field extraction (price presence, signature/authorization presence, dollar amount, project: risk: med; units: 8; Cross-reference layer against PM/accounting systems (Buildertrend/Knowify/QuickBooks/Proco: risk: high; units: 10; Alerting & notification (dashboard, email/SMS): risk: low; units: 4; Thread/state lifecycle tracking (dedupe, avoid re-flagging, status over time): risk: med; units: 6; Multi-tenant infra, mailbox security, data handling for sensitive contract content: risk: m
total units73
hardest unknownReliably telling a genuine unpriced/unsigned change-order request apart from ordinary field correspondence, at a false-positive rate low enough that an owner keeps checking the queue instead of muting it - and doing this without ground-truth access to what's already been priced or signed, since that fact lives in a separate PM/accounting system the email layer has to cross-reference, not in the email thread itself.
the verdict is not the end of the process

If you decide to do this

StepWhat it meansWhere it happens
1 · Read the case against it firstCharges the arbiter upheld are the ones to answer before committing. If an upheld charge is fatal for you, the verdict is not.on this page
2 · Commit the pre-registered testThe test is already written: (a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b) Paid conversion: number of the 10 recruited firms that pay $249 up front for the 30-day watch. (c) 30-day continuation: number of paying firms that authorise a second month at $249 or higher. at (a) ≥50% email-traceability across the pooled audit AND ≥40% in at least 6 of the 10 individual firms; (b) ≥5 of 10 recruited firms pay the $249; (c) ≥3 of those payers authorise month two.. Committing freezes it with a date, and it cannot be edited afterwards.promote it →
3 · Stand up the offerA landing page, a price, and an instrumented link. Nothing is proven until somebody who does not know you is asked to pay.ventures →
4 · Run distribution and let it resolveThe test resolves mechanically on its deadline: actual against threshold, no judgement. A test never distributed resolves VOID rather than FAIL — inaction is not evidence.automatic, daily
5 · The outcome grades this verdictWhatever happens is written back against this prediction and scored. That is what makes the next verdict better, and it is the only honest basis for ever claiming an accuracy.the ledger →

The evidence supports building it, and the objections below were answered rather than conceded. Steps 2 and 3 open the operator console, which lives under this same domain at /account and requires a log-in — the public record is readable by anyone, and committing a prediction against it is not. Step 5 happens automatically: this prediction is already frozen with its score, its confidence, and every dimension as it stood, waiting for an outcome to grade it against.