Beta

Data integrity

The receipts only work if the rows are right.

Every claim on this site links to the row of primary-source data it came from. That citation is the user-facing promise — but it’s only worth what the underlying row is worth. So the rows have to be right. Keeping them right is the work this page describes.

Receipts Report ingests federal contracting, congressional votes, lobbying disclosures, campaign finance, stock-trade disclosures, earmarks, the Congressional Record, and the legislator register — eight upstream sources, none of them connected to the others, each of them imperfect. The integrity work happens at every layer: the ingest, the cross-reference, the publication, and the correction.

The covenant

What we owe the reader.

The platform makes three load-bearing promises. They’re the brand’s foundation; every integrity check we run exists to honor them.

  1. We cite everything.

    Every number, every chart, every hub field, every answer from the agent threads to a primary-source row with a permalink. The reader can always verify the receipt. If a claim can’t be linked back to its source, it doesn’t go on the page.

  2. We rank by the numbers, not by the party.

    The platform is party-neutral by design. Same data shape for every politician — Republican, Democrat, Independent. Ranking sorts by dollar amount and connection count, never by political flavor. Same question, same shape of answer.

  3. We say when we’re wrong.

    When we find a fact wrong, a source stale, or an entity mismatched, we fix it and say so. The brand’s failure posture is the same voice applied to a harder moment: transparent, factual, fix-forward. No defensiveness.

How we keep it right

Three disciplines, run continuously.

Data integrity isn’t a milestone we cross and move on from. Every ingest, every cross-reference, and every publication runs through layered checks before it reaches a page.

Continuous integrity checks

Nothing reaches a hub unchecked.

Each of the eight upstream sources has its own format, its own update cadence, and its own quirks. After every ingest, the data passes a battery of automated checks that look for shapes that shouldn’t exist — orphaned references, impossible values, gaps where there should be coverage. The checks run on a schedule and surface anomalies before they reach a hub.

Drift detection between sources

When two sources disagree, we hold both and look closer.

The federal data ecosystem is full of independent systems that record overlapping facts in incompatible ways. A senator’s name in the lobbying database doesn’t always match their name in the FEC. A contract recipient in USAspending may carry three different identifiers across three subsystems. Our cross-reference layer is where we catch the seams and reconcile them — and where we flag the ones we can’t reconcile yet.

Corrections as the work

Finding what’s wrong IS the job.

A correction isn’t a setback to manage; it’s the system working. Every error we catch — through automated checks, a reader email, or a manual review — gets traced to its source, fixed, and re-run through the checks that should have caught it. The system grows more rigorous each time. We treat the absence of corrections as a sign we’re not looking hard enough, not a sign of perfection.

What you can trust on every page

The citation IS the guarantee.

The integrity work behind the scenes shows up at the reader-facing layer in a few specific ways. None of them require taking our word.

  • Every claim is one click from its source. Dollar amounts, vote records, sponsor names, disclosure dates — each links to the primary-source row on the federal site that published it.

  • Every politician gets the same kind of page. Same fields, same charts, same shape of answer when you ask the agent about them. No editorial choice about who matters or how to frame them. The party-neutrality is visible in the structure.

  • Coverage windows are stated, not implied. Each entity surface labels what we have and from when — so the absence of a row doesn’t get read as the absence of a transaction.

  • The agent cites every claim inline. When you ask a question in plain English, the answer comes back with the rows it drew from — not a paraphrase, not a summary that loses the source. If the agent can’t cite a claim, it doesn’t make the claim.

The corollary

We publish daily. We correct daily.

The product is the ingest, the cross-reference, the hub, the agent, the citation. The integrity discipline is the substrate underneath all of it. There isn’t a finished state; there’s only the continuous practice of getting the rows right and re-running the checks that should have caught the ones we didn’t.

If you find a row that looks wrong, the citation chip on that row goes straight to the federal source it came from — start there. If the source is right and our row is wrong, the discrepancy is ours to fix.

The scoreboard

How we grade ourselves.

We hold an internal report card. Every category is scored against a written rubric and re-graded whenever the underlying surface changes.

As of 21 May 2026

78

Automated integrity checks

9

Cross-source reconcile surfaces

6

Anomaly detectors

6

Cross-reference review queues

  • A

    Finding bad data

    Three layers of automated checking, all running on schedule. The 78 sanity invariants catch shapes that shouldn’t exist — orphaned references, impossible values, gaps where coverage is required. The 9 cross-source reconcile surfaces compare every dataset against its upstream and flag the deltas. The 6 anomaly detectors watch for outliers in donor patterns, stock-trade amounts, person identity, and job timing. Each layer catches a different failure mode.

  • A

    Fixing bad data

    Every cross-source matcher — 6 of them, one per pairing — has a paired operator review queue. When automatic resolution falls below confidence, the candidate sits in the queue for human judgment rather than getting published with a wrong attribution. Re-runs of the matcher are idempotent against operator decisions — once a row is confirmed by a human, no automation overwrites it.

  • A

    Freshness

    Every ingester runs on a schedule and supports delta-mode refreshes that catch only what changed since the last sweep. A shared fetch wrapper handles transient upstream failures with exponential backoff; persistent failures land in a triage queue rather than disappearing. The cron lanes (daily delta · weekly sweep · monthly reconcile) cover every kind in the catalog.

  • A

    Forensics

    Every row write carries a provenance tag pointing at the job that wrote it. Every job records exit code, log tail, invocation id, and the parameters it ran with. The full schema — source datasets, cross-reference layer, audit trail — is recoverable from scratch via a verified bootstrap contract.

  • A−

    Corpus completeness

    Every surface we ship at launch is fully populated against its committed scope — campaign finance back to 1980, lobbying back to 1999, Senate roll-call votes back to 1989, the Congressional Record back to 1995. The minus reflects four explicit expansions queued for after launch: pre-115 bills, pre-2002 FEC election results, Senate annual personal financial disclosures, and pre-FY24 earmarks. They’re named in the gaps section below.

What we’re still building

The honest gaps.

The product is launching with deep coverage on the surfaces that drive most first-pass investigations — campaign finance back to 1980, lobbying disclosures back to 1999, Senate roll-call votes back to 1989, the Congressional Record back to 1995, every dollar of federal spending we can match to a recipient. But the data universe is enormous, and a few specific surfaces are still expanding. The honest list:

  • Bills + amendments before the 115th — congressional legislation is on file from 2017 forward (the 115th Congress onward). Earlier Congresses are queued for backfill.

  • House roll-call votes 1989–2012 — Senate roll-call votes go back to the 101st Congress (1989) and are complete. The House equivalent for that same window comes from a separate upstream and is mid-backfill; the 113th Congress (2013) forward is already on file.

  • FEC election results before 2002— contributions, PAC activity, and committee expenditures reach back to 1980 in full. The per-cycle election-outcome tables (who won, who lost) cover 2002–2022; pre-2002 outcomes require a separate parsing pass from PDF records and are queued.

  • Senate annual financial disclosures — stock-trade periodic transaction reports cover both chambers. The longer annual personal financial disclosure forms are on file for House members today; the Senate equivalent is queued.

  • Earmarks before FY24 — House Community Project Funding and Senate Congressionally Directed Spending requests are on file from FY24 through FY26. Pre-FY24 earmark cycles are queued for backfill.

  • Cross-reference entries in review— connecting a lobbyist or stock-trade row to the right person sometimes lands in a queue for manual confirmation. We’d rather hold a row in review than publish a misattribution. Pending entries don’t appear on hubs until they clear.

We’re explicit about all of this on the relevant entity surfaces — the coverage note on a hub tells you what window you’re looking at, and the absence of a row is never silent.