Methodology Glossary — Plain-Language Edition
Every technical term used in our methodology, explained so that no statistics degree is required. This document is the seed of the site's public methodology page. Companion to
docs/money-politics-roadmap.md(the design) anddocs/money-politics-panel-synthesis.md(the full formulas).
The data sources (the acronyms)
FEC — the Federal Election Commission, where campaign money is legally required to be reported. Who gave, to which committee, how much, and when.
Committee (campaign committee / PAC) — money never legally flows straight from a corporation to a politician. It flows into committees: a candidate's own campaign committee, a Political Action Committee (PAC), a party committee. When we trace funding, we are tracing these committee-to-committee flows — that's how the money actually moves.
Super PAC / independent expenditure — a committee that can raise unlimited money but may not coordinate with a candidate; it spends on behalf of candidates (ads, mailers) rather than giving to them. We track this money separately from direct contributions because it is legally and practically a different thing.
LDA filings — quarterly lobbying disclosures required by the Lobbying Disclosure Act. Companies report which issues and bills they lobbied on and their total quarterly lobbying spend. Important honesty note: the law requires a lump sum per quarter, not per bill — so we can say "Company X lobbied on this bill and spent $2M lobbying that quarter," but never "spent $2M on this bill."
CBO — the Congressional Budget Office, Congress's nonpartisan scorekeeper. Its "cost estimates" say what a bill would actually cost and who gains or loses money. We use CBO scores as ground truth to test claims against.
CRS — the Congressional Research Service, Congress's nonpartisan research arm. Plain-language expert reports on what bills do.
CREC (Congressional Record) — the official verbatim transcript of everything said on the House and Senate floor. Our primary source of on-the- record political statements: attributed, dated, and tied to specific bills.
ClaimReview / external fact-checks — a machine-readable standard used by fact-checking organizations (PolitiFact, FactCheck.org, AP, and others) to publish their verdicts. We import these as corroboration — a second opinion alongside our own ratings, so we are never the only voice.
How claims get rated
Claim vs. utterance — a claim is an assertion ("this bill cuts taxes for families"); an utterance is one instance of someone saying it (a speech, a press release, an ad). One claim may be uttered by forty politicians. For scoring, the claim counts once — but the number of speakers repeating it is shown separately, because coordinated repetition is its own story.
Fact / prediction / opinion — only factual claims (checkable against evidence) receive accuracy verdicts. Predictions ("this will create jobs") and opinions ("this is un-American") are labeled as such and never counted in accuracy scores. You can't fact-check a prophecy.
Graded verdicts — ratings aren't just true/false. A claim can be Mostly True, Half True, Mostly False, etc. Each grade maps to a published number so scores reflect degrees of accuracy instead of forcing everything into two buckets.
Sampling frame — the defined, complete pool of statements our ratings are drawn from — for example, "every floor statement about a bill that has a CBO cost estimate." Having a fixed pool means nobody hand-picks which statements are eligible.
Frame-random sampling (the audit series) — the computer randomly selects which statements from the frame get rated each quarter — the way a pollster dials random numbers instead of interviewing whoever calls in. The editor rates whatever the draw produces, including boring statements. This is our defense against the accusation that we cherry-picked which claims to check: we literally cannot, the dice choose.
Blind rating — while rating a statement, the rater does not see the speaker's funding profile, so knowing who funds a politician can't nudge the verdict. The verdict is sealed before the money context is attached.
Stakes-weighted (stratified) sampling — deception isn't spread evenly: a skilled operator is accurate about trivia and lies exactly where it counts. So, like a financial auditor who checks the $5M transfer more often than the $50 lunch, our random draws are deliberately weighted toward high-stakes statements — about bills the speaker voted on or sponsored, that reached a floor vote, involve large CBO dollar amounts, or touch their funders' interests. The selection is still made by published rules and dice; the dice are just aimed where lying pays.
High-stakes tier / routine tier / editorial tier — the three groups of rated claims and their three jobs. High-stakes (sampled): the accountability engine our metrics run on. Routine (sampled): the baseline. Editorial ("noticed" — claims we or fact-checkers chose to examine because they were prominent): the stories on the site. Editorial verdicts are displayed but never blended into the statistical scores, because chosen claims are, by definition, not a fair sample.
Strategic deception gap — a politician's routine-tier accuracy minus their high-stakes-tier accuracy, shown side by side (e.g., "91% accurate on routine statements; 38% on high-stakes ones"). It directly measures the pattern of being honest when it's free and dishonest when it counts — a pattern a single blended average would hide.
How the scores are computed
Base rate — the background average everyone is compared against. If the average politician's claims rate 30% inaccurate, then a politician at 32% is ordinary, not scandalous. Every score we publish is shown against its relevant base rate, never floating alone.
Stratum (party × chamber × topic) — comparisons are made within matched groups: House Republicans on energy are compared to House Republicans on energy. This prevents statistical illusions where a pattern that exists in a blended average disappears (or reverses) once you look inside the groups — and it prevents our scores from simply rediscovering "party A talks about topic X more than party B."
Small-sample correction (shrinkage) — when a politician has only a handful of rated claims, their score is statistically pulled toward their group's average rather than trusted at face value. Three rated claims shouldn't brand anyone a liar — or a saint. As more ratings accumulate, their own record takes over from the group average.
Matched expectation (permutation test) — to judge a funder's portfolio, we build thousands of imaginary "random portfolios" from comparable politicians (same party/chamber/topic mix) and ask: does this funder's actual portfolio score worse than those? The published number is the comparison — "portfolio score 0.41 vs. matched expectation 0.29" — not a naked score.
Uncertainty range (confidence interval) — every score is published as a range, not a point. "0.41 (range 0.33–0.49)" means the data supports a value somewhere in that band. If the band is wide, we say so; precision is never implied where it doesn't exist.
Capping (winsorization) — one mega-recipient can't dominate a funder's portfolio score: extremely large weights are capped so a score reflects the portfolio, not one relationship.
Recency decay (half-life) — old money matters less. Contribution weights fade with time (halving every two election cycles), so a score reflects current behavior, not ancient history.
Topic matching — the strongest signal isn't "this funder's politicians are inaccurate"; it's "inaccurate about the topics the funder profits from." We match claims to a funder's lobbying topics using the categories in their own lobbying filings.
Funding window (D-during) — deception scores count claims made during the period the funder was actually funding the politician — not before the relationship existed.
Reward signal (R_F) — the sharpest question we ask: after a politician makes inaccurate topic-matched claims, does this funder increase their funding next cycle? A funder that consistently rewards deception with more money exhibits a measurable pattern of conduct. Note what this is not: a claim about anyone's intent or knowledge — only an observed pattern.
Display gates (minimum data rules) — no score is published until there is enough data behind it: at least 30 rated claims, at least 5 distinct recipients, sufficiently spread across the portfolio. Below the gate, the page says "insufficient data" — we show nothing rather than something shaky.
Opacity metric — the percentage of an entity's funding that flows through untraceable vehicles (unitemized donations, dark-money nonprofits, fresh single-cycle super PACs). Published beside every score, so hiding money worsens the public picture instead of escaping it.
Funder network (entity resolution) — "one funder" often means a web of related committees and organizations. We group them using official records (FEC-registered connections, committee-to-committee transfers, corporate registries), and we publish which committees each network includes so the grouping itself can be checked.
The fairness safeguards
Funded ≠ false — the foundational rule. Money context sits beside a claim's verdict; it never determines it. A claim is judged on evidence alone. Funder scores describe statistical patterns in portfolios — never the truth of any individual claim, and never anyone's intent.
Pre-registration — the formulas, sampling rules, and thresholds are published before any scores are computed, with a permanent timestamped record. We cannot quietly tune the method after seeing who it flatters or condemns — and you can verify that.
Right of reply — any politician or funder we score can submit a response, which is displayed on their profile, unedited, alongside the score.
Coverage dashboard — a public, always-current accounting of who is getting rated, by party and topic, so any imbalance in our attention is measurable by anyone — including our critics.
Versioned corrections — verdicts and scores carry visible version histories. When we correct something, the change and its reason are public; nothing is silently edited.
Positive symmetry — clean records are published as prominently as bad ones. This is a lens, not a hit list; a funder whose portfolio shows no deception pattern gets that said, with the same rigor.
Replication — the rated-claim dataset is downloadable, so any researcher (or opponent) can recompute our numbers and check our work. Trust nothing; verify everything — including us.