Unbilled Change-Order Watchdog for Subcontractors
GAP: subs lose billable margin because change-order requests arrive informally (email, verbal, field notes) and get executed before being priced or signed, and by the time billing happens the paper trail is gone or disputed. WHO HAS IT: specialty trade subcontractors in the $5-50M revenue band, where change orders are common but back-office capacity to police every email thread is thin. WHAT IS SOLD: a monitoring age…
What an analysis cost to produce belongs beside it. A reader deciding whether to trust a verdict is entitled to know whether it came from twenty-six stages or one, and nothing else in this category will tell them.
The case against it
| Charge | Rebuttal | Ruling |
|---|---|---|
| Signal lives outside email, so the product structurally can't see most of the leak | CONCEDED — Conceded. The thesis itself flags this as an unstated assumption, and no evidence in the record establishes what share of change-order-triggering comm | upheld — Conceded with no counter-evidence; the share of change-order-triggering communication that touches email is unmeasured and is the load-bearing assumption of the entire product — this is the |
| Flagging without closing the loop doesn't recover the margin, so ROI attribution collapses | CONCEDED — Conceded as a design flaw in the current scope, though it is fixable rather than fatal. The value-based pricing story ('dollars of margin recovered') | partial — Real as scoped, but the fix is known-in-kind and buildable (draft COR generation, e-signature chase, GC submission) and incumbents already ship the downstream half; this constrains pricing m |
| Mailbox-wide access request is a trust barrier the buyer persona won't clear quickly | Partially rebutted. Full OAuth mailbox access is not the only viable ingestion architecture: a dedicated forwarding address or shared CC/BCC alias (th | dismissed — Forwarding/CC-alias ingestion is an established pattern in this exact category and avoids broad OAuth entirely; the residual issue is behaviour change in field/PM staff, which is an adoption |
| Precision is unproven at launch and trade-correspondence NLP is unusually noisy | CONCEDED — Conceded outright. The data_moat section itself admits 'a new entrant starts with no verified precision/recall numbers and no proof of dollars actuall | partial — Precision genuinely unproven and alert fatigue is the likeliest churn mechanism — but the concierge design sidesteps it entirely for the first cohort by putting a human in the loop, which is |
| Incumbents can bolt on an email-ingestion layer faster than the entrant can build a data moat | Partially rebutted. The incumbent_weakness section provides evidence of a structural, not just resource-based, disincentive: Clearstory's positioning | dismissed — Banned reasoning territory: 'a funded competitor could copy this' is true of nearly every software wedge. The structural disincentive (Clearstory's whole pitch is moving work off email) is a |
| Escalating a flag to the GC creates relationship risk the sub may not want automated | Partially rebutted by the product's own design constraint: the autonomy spec treats GC-escalation as a 'needs_human' judgment call rather than an auto | partial — The tool as specified does not auto-escalate, so the stated failure mode is not what's proposed — but the deeper tension (subs deliberately under-formalising paper trails to protect repeat w |
| The buyer may treat unbilled COs as an accepted cost of doing business, not an active problem to solve | Partially rebutted. The Levelset channel evidence shows subs actively posting about this exact problem in their own words on public forums — unprompte | partial — Unprompted Levelset posting and $28M+ of venture funding into the category cut against full normalisation, but the conversion rate within the $5–50M band is genuinely unknown and is a legiti |
| Proposed channels don't scale past a handful of relationship-driven deals | CONCEDED — Conceded. The candidate's own channel notes admit the one scalable channel is already occupied: 'Buyers actively search this category, but Siteline, R | partial — Conceded on SEO occupancy, but the ruling overreaches: ASA/CFMA chapter rolls, licence registries and the G2 category are enumerable lists addressable by outbound and paid placement — that i |
| Thin back-office capacity — the stated reason subs need the tool — is also the reason they can't operate it | CONCEDED — Conceded. The buyer profile is explicitly defined by thin back-office capacity unable to police email manually, and the same team is required to triag | upheld — Genuine design contradiction as specified. The only resolutions are high-confidence-only alerting or a done-for-you triage service (i.e. sell the concierge permanently, not as a stepping sto |
A separate agent argued against this idea, a second answered, a third ruled. 5 of 9 charges were conceded rather than defended. Published in full because a score with the objections removed is a advertisement, and because the objections are usually more useful than the verdict.
How it scored
| Dimension | Score | Reasoning |
|---|---|---|
| D1 | 70 | Unbilled/disputed change orders are a documented, recurring margin leak in trade subcontracting (24-day sign-to-submit lag, work performed before pricing), and subs post |
| D2 | 72 | Strong adjacent payment evidence: Knowify charges $99–$329/mo, Clearstory raised $28M+ on this exact workflow, eSUB $23.5M, Siteline at 10,000+ projects — subs demonstrab |
| D3 | 78 | Buyers are exceptionally addressable: state contractor licence registries, ASA and CFMA chapter rolls, the G2/Capterra 'Change Order Management' category with existing co |
| D4 | 30 | At v1 the wedge is an email classifier any funded incumbent can bolt on as a connector; the outcome-linked labeled corpus is a real but slow-accruing moat that does not e |
| D5 | 62 | The concierge version is genuinely cheap ($6.5k cash, 3 days to stand up, human-in-the-loop), but the automated product needs a precision-tuned classifier over noisy trad |
| D6 | 62 | Recurring per-firm subscription is a clean model with obvious annual contracts and expansion by project volume, but the stated value-based pricing ('dollars of margin rec |
| D7 | 58 | Tens of thousands of US specialty subs in the target band and a portable, non-English-market gap put a Clearstory-scale ($10–50M ARR) outcome inside the range, but only i |
Who already does this
| Competitor | Pricing | Funding | Launched | Overlap |
|---|---|---|---|---|
| Clearstory | Free Basic tier; Standard/Professional add per-user + per-pr | Seed $5.3M, Series A $7M (Cloud Apps Capital), | ~2018 (per '6 years ago' r | partial |
| Siteline | Not publicly listed; billing-software positioning | Not disclosed in results | Pre-2020 (10,000+ projects | partial |
| Knowify | Core $99/mo (annual, 1 user); Advanced $329/mo (annual, up t | $8.45M total, last raise $3M (2023) | 2012 | partial |
| eSUB Cloud | Not publicly listed | $23.5M total, last raise $2M (2023) | 2008 | partial |
| GCPay | Not publicly listed | Not disclosed | Not disclosed | adjacent |
| Planyard | 14-day free trial, paid tiers not disclosed in results | Not disclosed | Active as of 2025 | adjacent |
| Rhumbix | Not publicly listed | Not disclosed | Not disclosed | partial |
| Kynection | Not publicly listed | Not disclosed | Not disclosed | adjacent |
| Mirage Metrics (AI agent layer for PM systems) | Not disclosed | Not disclosed | Not disclosed | partial |
Where the buyers actually are
| Channel | Why it reaches them |
|---|---|
| Levelset Payment Help community (payment-help/question threads) | |
| CFMA local chapters + Connection Café | |
| Long-tail organic search ("unbilled change order tracking," "change or | |
| ASA (American Subcontractors Association) national/state chapters | |
| G2/Capterra "Change Order Management Software" category |
Regulatory gates
| Gate | Finding |
|---|---|
| G1 | Email monitoring and flagging is lawful. No prohibition on reading a customer's own email traffic with permission, analyzing it, and surfacing alerts. SaaS compliance (data handling, securit |
| G2 | Core operation is automated email monitoring and rule-based flagging. Requires no founder presence after deployment. Customer success, onboarding, and support are staffable. The agent itself |
| G3 | Existing paid competitors prove willingness to pay. Buildertrend, Knowify, Siteline, eSUB and others charge for change-order management modules. Subcontractors already budget for constructio |
| G4 | Email parsing, NLP-based change-order detection, and alert generation are all standard SaaS capabilities today. Requires no scientific breakthrough. LLMs can classify construction correspond |
| G5 | Subcontractors are reachable through multiple channels: construction software marketplaces (Appstore integrations with Buildertrend, Knowify), trade associations (AGC, specialty trade groups |
The pre-registered test
| Term | Value |
|---|---|
| days | 60 |
| offer | Two-part paid pilot with 10 named specialty subs in the $5–50M revenue band, recruited from ASA and CFMA local chapter rosters plus direct outreach off state contractor licence registries. Part A (evidence): for a $250 p |
| price | 249 |
| metric | (a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b) Paid conversion: number of the 10 recruited firms that pay $249 up fr |
| channel | ASA and CFMA local chapter membership rosters and meetings (dues paid, in-person attendance), supplemented by direct outbound to state contractor licence registry listings in two metros |
| cost usd | 9500 |
| falsifies | If (a) misses, the assumption that email is where the change-order signal lives dies outright — the product cannot see the leak it claims to intercept and must be re-scoped to SMS/voice/field-ticket capture, which is a d |
| threshold | (a) ≥50% email-traceability across the pooled audit AND ≥40% in at least 6 of the 10 individual firms; (b) ≥5 of 10 recruited firms pay the $249; (c) ≥3 of those payers authorise month two. |
Recorded at the moment the verdict was issued and not editable afterwards. If this is launched, the result lands on the ledger whether it passes or fails.
Where this idea is
| Phase now | Analysed — Underwritten, with the argument against it on the record. |
| What you do here | Read the case against it first. An upheld charge you cannot answer is the verdict, whatever the score says. |
| To leave this phase | You have read the upheld charges and decided the idea survives them. |
| Gate status | This gate is a judgement, not a query. The system will not rule on it and will not pretend to — you decide, and the reason is recorded. |
| Next phase | Validating — A pre-registered test is live and running. |
This gate is a judgement rather than a query, so the system states it and refuses to rule on it. Pretending software can decide whether a business "can take money from somebody who is not you" would make every gate on this site meaningless. Advancing an idea needs its link — the one handed back when it was submitted. Founder-owned ideas are advanced from the console. See the whole pipeline.
| What to do in this phase | What it proves | From which part of the analysis |
|---|---|---|
| Answer the upheld charge: Signal lives outside email, so the product structurally can't see most of the leak | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: Flagging without closing the loop doesn't recover the margin, so ROI attribution collapses | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: Precision is unproven at launch and trade-correspondence NLP is unusually noisy | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: Escalating a flag to the GC creates relationship risk the sub may not want automated | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: The buyer may treat unbilled COs as an accepted cost of doing business, not an active problem to solve | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: Proposed channels don't scale past a handful of relationship-driven deals | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: Thin back-office capacity — the stated reason subs need the tool — is also the reason they can't operate it | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
Every step traces to a field this idea's own underwriting produced — not generic best practice, which is free everywhere. 0 of 7 complete. Mark them off in the console.
How this was produced
| Measure | Value |
|---|---|
| Cost to produce | $7.11 |
| Wall clock | 34 minutes |
A verdict produced by 22 of 23 stages is not the same artefact as one produced by all of them, and which stages failed was recorded on every run and shown nowhere until now. If a stage that feeds a section died, the section came from somewhere else or nowhere — and you are entitled to know which is in front of you before you act on it.
The money
| price point | None |
| current spend | None |
| funding route | presale |
| revenue model | None |
| churn monthly pct | None |
| cash to first dollar | 6500 |
| marginal cost per unit | 38 |
What it costs to start
| capital blocked | False |
What has to be built
| data moat | A growing labeled corpus linking specific email phrasing patterns (trade-specific, regional, per-customer) to confirmed outcomes - which flags turned into real signed/priced change orders and recovered dollars, versus which were noise. This outcome-linked dataset is what lets the classifier's precision improve over time and lets the vendor credibly price against margin recovered; a new entrant starts with no verified precision/recall numbers and no proof of dollars actually recovered per flag type, both of which take real customer usage over months to accumulate. |
| components | Email ingestion & mailbox sync (Gmail/O365 OAuth, polling, thread reconstruction): risk: med; units: 6; Change-order-signal classifier (distinguish CO request from routine correspondence): risk: high; units: 12; Field extraction (price presence, signature/authorization presence, dollar amount, project: risk: med; units: 8; Cross-reference layer against PM/accounting systems (Buildertrend/Knowify/QuickBooks/Proco: risk: high; units: 10; Alerting & notification (dashboard, email/SMS): risk: low; units: 4; Thread/state lifecycle tracking (dedupe, avoid re-flagging, status over time): risk: med; units: 6; Multi-tenant infra, mailbox security, data handling for sensitive contract content: risk: m |
| total units | 73 |
| hardest unknown | Reliably telling a genuine unpriced/unsigned change-order request apart from ordinary field correspondence, at a false-positive rate low enough that an owner keeps checking the queue instead of muting it - and doing this without ground-truth access to what's already been priced or signed, since that fact lives in a separate PM/accounting system the email layer has to cross-reference, not in the email thread itself. |
If you decide to do this
| Step | What it means | Where it happens |
|---|---|---|
| 1 · Read the case against it first | Charges the arbiter upheld are the ones to answer before committing. If an upheld charge is fatal for you, the verdict is not. | on this page |
| 2 · Commit the pre-registered test | The test is already written: (a) Email-traceability rate: share of executed change orders in the audited 90-day log that had a detectable email trace ≥3 days pre-mobilization. (b) Paid conversion: number of the 10 recruited firms that pay $249 up front for the 30-day watch. (c) 30-day continuation: number of paying firms that authorise a second month at $249 or higher. at (a) ≥50% email-traceability across the pooled audit AND ≥40% in at least 6 of the 10 individual firms; (b) ≥5 of 10 recruited firms pay the $249; (c) ≥3 of those payers authorise month two.. Committing freezes it with a date, and it cannot be edited afterwards. | promote it → |
| 3 · Stand up the offer | A landing page, a price, and an instrumented link. Nothing is proven until somebody who does not know you is asked to pay. | ventures → |
| 4 · Run distribution and let it resolve | The test resolves mechanically on its deadline: actual against threshold, no judgement. A test never distributed resolves VOID rather than FAIL — inaction is not evidence. | automatic, daily |
| 5 · The outcome grades this verdict | Whatever happens is written back against this prediction and scored. That is what makes the next verdict better, and it is the only honest basis for ever claiming an accuracy. | the ledger → |
The evidence supports building it, and the objections below were answered rather than conceded. Steps 2 and 3 open the operator console, which lives under this same domain at /account and requires a log-in — the public record is readable by anyone, and committing a prediction against it is not. Step 5 happens automatically: this prediction is already frozen with its score, its confidence, and every dimension as it stood, waiting for an outcome to grade it against.