Line-Level Code Review Coverage Auditor
Engineering leads and compliance-driven teams cannot tell which lines of code have actually been reviewed, by whom, or how many times — they only have git blame, which shows last-touch authorship, not review history. Today they either trust that a green PR approval means real scrutiny happened, or manually cross-reference diffs and comment threads, which collapses on large files, stacked PRs, and rubber-stamp approva…
What an analysis cost to produce belongs beside it. A reader deciding whether to trust a verdict is entitled to know whether it came from twenty-six stages or one, and nothing else in this category will tell them.
The case against it
| Charge | Rebuttal | Ruling |
|---|---|---|
| The thesis itself is built on a single weak signal (mention_count 1), insufficient to justify build time. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | upheld — One mention is noise, not demand; nothing in the record shows anyone asked for per-line review provenance. |
| Eight funded, established competitors already occupy the adjacent engineering-analytics/review-metrics space. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | partial — The granularity gap is real and specific, but a gap eight funded vendors chose not to fill is more likely low-value than overlooked — and any of them can ship it as a feature in a sprint. |
| There is no proprietary data moat; competitors can reconstruct the same signal from the same public GitHub/GitLab APIs. | CONCEDED — The defence conceded this. | upheld — Conceded, and correct: a prompt-free but still single-API-source product with no accumulating corpus. |
| No geographic arbitrage exists; incumbents already sell globally. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | upheld — Dev tooling is global from day one; the Japan observation is an absence of demand, not an absence of supply. |
| The buyer population figure is an unverified guess. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | upheld — 8,000 is a derived narrowing of a narrowing; the compliance-driven subset that would pay for review depth specifically is unmeasured. |
| A solo builder with no audience has no viable channel to reach compliance-driven engineering leads against funded incumbents. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | partial — Record has zero GTM, which is fatal as written — but dev tools are one of the few categories where a solo builder does have a real cold channel (GitHub Marketplace, open-source CLI, HN/Lobst |
| Technical feasibility: attributing review to lines across squash merges, rebases, and stacked PRs is nontrivial and unaddressed. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | upheld — This is the core claim and it is an approximation problem; a metric sold as audit evidence that is wrong 10% of the time is worse than no metric. |
| No validated willingness to pay for a line-level metric versus existing DORA/velocity metrics. | Rebuttal not recorded — the defence stage ran but its output was not persisted for this candidate. | upheld — No pilot, LOI, or conversation cited; the buyer already pays for a tool that reports 'reviewed' and their auditor already accepts it. |
A separate agent argued against this idea, a second answered, a third ruled. 1 of 8 charges were conceded rather than defended. This exchange is reconstructed from the arbiter's rulings. The prosecution and defence stages ran for this analysis but their output was not persisted, so each charge and each concession is as the ARBITER recorded it, not in the prosecutor's or the defender's own words. That is a defect in our storage and it is named here rather than presenting a partial record as a complete one. Published in full because a score with the objections removed is a advertisement, and because the objections are usually more useful than the verdict.
How it scored
| Dimension | Score | Reasoning |
|---|---|---|
| D1 | 28 | Real but latent: nobody is currently bleeding from not knowing per-line review counts — SOC2 auditors accept PR-approval logs today, so the pain is theoretical policy hyg |
| D2 | 30 | Adjacent budget clearly exists ($15-800/dev), but zero evidence anyone pays for review-depth specifically; the line item this replaces is already funded and satisfied. |
| D3 | 22 | No channel in the record at all; the only credible cold paths (GitHub Marketplace listing, free CLI/Action going viral) are low-conversion and crowded, and the compliance |
| D4 | 18 | Single public API, no accumulating dataset, no join across sources, no jurisdiction friction — a CI gate creates mild workflow stickiness but any incumbent already holdin |
| D5 | 35 | Plumbing is a weekend (webhooks, Postgres, Stripe) but the line-attribution engine under rebase/squash/rename is the whole product and is multi-week unbounded algorithmic |
| D6 | 48 | Per-seat recurring at $12/dev is clean and self-serve, but 6%/mo inferred churn on an unbundled single-metric point solution means the model leaks faster than a solo buil |
| D7 | 30 | 24 units × ~$500/mo ceiling assumes winning a compliance slice of a market where eight vendors sell the general case; realistic outcome is a sub-$5k MRR feature that gets |
| D8 | 62 | Genuinely in his competence — git internals, webhooks, Flask/Postgres — but he has no engineering-leadership audience and no enterprise sales motion, which is exactly wha |
Who already does this
| Competitor | Pricing | Funding | Launched | Overlap |
|---|---|---|---|---|
| Code Climate Velocity | not published; industry entry ~$15-30/dev/mo | acquired/operated by Code Climate (established | ~2018 (as GitPrime success | partial |
| Swarmia | not published, per-engineer seat pricing | ~$19M total ($8M seed/pre-seed 2021 + $11M Ser | 2020 | partial |
| LinearB | per-developer seat, not published; Gartner Magic Quadrant Le | VC-backed, multiple rounds (undisclosed in sea | 2019 | partial |
| Waydev | $449/year per active engineer | bootstrapped/VC, unclear | 2017 | partial |
| CodePulse | ~$15-25/developer/month | unclear, appears bootstrapped/early-stage | 2025 (2025 Engineering Ben | partial |
| Jellyfish | $500-800/developer/year (enterprise) | well-funded enterprise engineering-analytics p | 2017 | adjacent |
| GitPrime / Pluralsight Flow | enterprise, not published | acquired by Pluralsight (public co.) | 2015 (as GitPrime) | partial |
| Allstacks | enterprise contracts, contact sales | VC-backed value-stream analytics platform | ~2018 | adjacent |
The pre-registered test
| Term | Value |
|---|---|
| days | 30 |
| offer | Free, open-source GitHub Action + CLI: 'review-coverage' — runs on any repo, outputs a per-line review-depth heatmap (lines reviewed 0x / 1x / 2+ times, with reviewer names and staleness) plus an optional failing check w |
| price | 12 |
| metric | Number of qualified paid-intent conversations: a named engineering leader at a 30+ engineer org who (a) installed the Action on a private repo or received their own heatmap, AND (b) states in writing that PR-level review |
| channel | GitHub Marketplace listing for the Action + one Show HN + posts to r/ExperiencedDevs and Lobsters + 40 direct cold emails to VP Eng / Head of Platform at 30-300 engineer companies that publicly advertise SOC2 or ISO27001 |
| threshold | KILL unless ≥4 qualified paid-intent conversations AND ≥25 private-repo installs. PROCEED TO BUILD only if ≥2 of those 4 name a specific audit or policy failure caused by PR-level-only evidence AND at least 1 agrees to p |
Recorded at the moment the verdict was issued and not editable afterwards. If this is launched, the result lands on the ledger whether it passes or fails.
Where this idea is
| Phase now | Analysed — Underwritten, with the argument against it on the record. |
| What you do here | Read the case against it first. An upheld charge you cannot answer is the verdict, whatever the score says. |
| To leave this phase | You have read the upheld charges and decided the idea survives them. |
| Gate status | This gate is a judgement, not a query. The system will not rule on it and will not pretend to — you decide, and the reason is recorded. |
| Next phase | Validating — A pre-registered test is live and running. |
This gate is a judgement rather than a query, so the system states it and refuses to rule on it. Pretending software can decide whether a business "can take money from somebody who is not you" would make every gate on this site meaningless. Advancing an idea needs its link — the one handed back when it was submitted. Founder-owned ideas are advanced from the console. See the whole pipeline.
| What to do in this phase | What it proves | From which part of the analysis |
|---|---|---|
| Answer the upheld charge: The thesis itself is built on a single weak signal (mention_count 1), insufficient to justify build time. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: Eight funded, established competitors already occupy the adjacent engineering-analytics/review-metrics space. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: There is no proprietary data moat; competitors can reconstruct the same signal from the same public GitHub/GitLab APIs. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: No geographic arbitrage exists; incumbents already sell globally. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: The buyer population figure is an unverified guess. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the partial charge: A solo builder with no audience has no viable channel to reach compliance-driven engineering leads against funded incumbents. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: Technical feasibility: attributing review to lines across squash merges, rebases, and stacked PRs is nontrivial and unaddressed. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
| Answer the upheld charge: No validated willingness to pay for a line-level metric versus existing DORA/velocity metrics. | The verdict survives its strongest objection, or it does not and you have learned that before spending. | arbitration |
Every step traces to a field this idea's own underwriting produced — not generic best practice, which is free everywhere. 0 of 8 complete. Mark them off in the console.
How this was produced
| Measure | Value |
|---|---|
| Wall clock | 30 minutes |
A verdict produced by 22 of 23 stages is not the same artefact as one produced by all of them, and which stages failed was recorded on every run and shown nowhere until now. If a stage that feeds a section died, the section came from somewhere else or nowhere — and you are entitled to know which is in front of you before you act on it.
The money
| price point | anchor: Swarmia/CodePulse entry-tier per-developer pricing ($15-20/dev/mo and ~$4/dev/mo respectively); monthly: 12; rationale: This is a single-feature wedge, not a full DORA/velocity suite, so it cannot command Swarmia/LinearB full-platform pricing. But it targets a sharper, higher-stakes use case (audit evidence, not vanity dashboards) than CodePulse's cheapest tier, so pricing sits between the two: roughly $10-15/dev/month. At typical target org size (30-100 engineers) that's a ~$300-1,500/mo deal, comparable to what these orgs already allocate to one analytics tool line-item. |
| current spend | amount: ~$4-37 per developer per month depending on tier; enterprise flat contracts $100K+/yr; source: Competitor pricing found in search: <cite index="17-1">Leading platforms include Jellyfish ($100K+/yr for enterprises), LinearB (free tier + paid), CodePulse ($199/mo for 50 devs), and Swarmia ($15-20/dev/month).</cite> Additionally Waydev is priced at $449/year per active engineer per the candidate's own competitor data. No vendor was found charging specifically for per-line review-coverage as a standalone SKU — it does not exist as a line item today; teams currently pay for broader velocity/DORA suites and get review counts as a minor sub-feature, if at all.; on what: General engineering- |
| funding route | reinvest |
| revenue model | subscription |
| churn monthly pct | why: No churn data was found for this specific tool category (it doesn't exist as a standalone product yet). Inferred from general pattern: narrow, single-metric B2B dev-tool add-ons in an already-saturated category (8 direct/adjacent competitors identified, several enterprise-funded) face high substitution risk — a customer can get 'good enough' review visibility bundled into an existing LinearB/Swarmia/Jellyfish contract they already pay for, so a standalone tool is the first line item cut. 5-7% monthly churn is a reasonable inferred range for an unbundled point solution at this stage; treat as inferred, not observed.; value: 6 |
| cash to first dollar | 300 |
| marginal cost per unit | value: 1.5; components: Per active repo/org per month: git history + PR metadata ingestion (API calls to GitHub/GitLab/Bitbucket), incremental diff/blame computation to build per-line reviewer counts, storage of line-level history, and heatmap/CI-gate serving. On a single VPS this is dominated by CPU time for diff processing on large repos and storage growth, not by per-seat cost — marginal cost per additional developer seat within an already-onboarded repo is near zero (~$0.10-$0.50), while marginal cost per newly onboarded repo (initial full-history ingestion) is higher (~$2-5 one-time, amortized). This is an inferred engineering estimate from typical git-processing workloads, not a source |
What it costs to start
| capital blocked | False |
What has to be built
| data moat | Weak to none. All source data (git history, PR review threads) lives in GitHub/GitLab and is equally reconstructable by any competitor via the same public APIs — there's no proprietary corpus being accumulated. The only quasi-moat is calibration data if users flag mis-attributed lines (improves the heuristic engine over time), but that's a minor tuning edge, not a defensible moat, and every incumbent listed already has years of engineering-analytics data and distribution this candidate lacks. |
| components | GitHub/GitLab App install, OAuth, webhook receiver: risk: low; units: 2; Ingestion pipeline (commits, PR metadata, review comments, diff hunks): risk: med; units: 3; Line-provenance engine: map review comments through renames/squashes/force-pushes to curre: risk: high; units: 5; Per-line reviewer-count + staleness scoring, storage schema: risk: med; units: 2; Heatmap UI (file-level coverage overlay): risk: low; units: 3; CI gate / status-check integration (block merge on low coverage): risk: med; units: 2; Dashboard/reporting (repo and team rollups): risk: low; units: 2; Multi-tenant auth, org management, Stripe seat billing: risk: low; units: 2; Backfill + reconciliation cron, rate-limit/ba |
| total units | 24 |
| hardest unknown | Review comments are pinned to a diff hunk at PR time, but code keeps moving after — through rebases, squash merges, whitespace reformatting, renames, and stacked PRs. Reconstructing 'this exact current line was seen by this reviewer' months later is an approximation problem, not a lookup. Git blame already gets this wrong for last-touch authorship; doing it for review provenance is strictly harder and any errors directly undermine the product's core claim (an auditable metric). This is a multi-week algorithmic risk, not a plumbing task, and it's unvalidated against real repos with heavy rebase/squash workflows. |
If you decide to do this
| Step | What it means | Where it happens |
|---|---|---|
| 1 · Read the case against it first | Charges the arbiter upheld are the ones to answer before committing. If an upheld charge is fatal for you, the verdict is not. | on this page |
| 2 · Commit the pre-registered test | The test is already written: Number of qualified paid-intent conversations: a named engineering leader at a 30+ engineer org who (a) installed the Action on a private repo or received their own heatmap, AND (b) states in writing that PR-level review evidence has been questioned by an auditor or internal policy, AND (c) either books the call or explicitly asks for pricing/invoice terms. Secondary tracked metric: Action installs on private repos. at KILL unless ≥4 qualified paid-intent conversations AND ≥25 private-repo installs. PROCEED TO BUILD only if ≥2 of those 4 name a specific audit or policy failure caused by PR-level-only evidence AND at least 1 agrees to prepay 3 months at $12/dev before the hosted dashboard exists. 0-3 conversations or <25 installs = dead, do not revisit.. Committing freezes it with a date, and it cannot be edited afterwards. | promote it → |
| 3 · Stand up the offer | A landing page, a price, and an instrumented link. Nothing is proven until somebody who does not know you is asked to pay. | ventures → |
| 4 · Run distribution and let it resolve | The test resolves mechanically on its deadline: actual against threshold, no judgement. A test never distributed resolves VOID rather than FAIL — inaction is not evidence. | automatic, daily |
| 5 · The outcome grades this verdict | Whatever happens is written back against this prediction and scored. That is what makes the next verdict better, and it is the only honest basis for ever claiming an accuracy. | the ledger → |
Not now. Something specific would have to change first, and it is named in the ruling. Steps 2 and 3 open the operator console, which lives under this same domain at /account and requires a log-in — the public record is readable by anyone, and committing a prediction against it is not. Step 5 happens automatically: this prediction is already frozen with its score, its confidence, and every dimension as it stood, waiting for an outcome to grade it against.