Methodology
This page describes what was done, in order, using the project's README and scripts. Figures were generated 2026-10-01. All processing was run by scripts; the rules and thresholds below are applied identically to every member. Any step that involved human or assistant review is disclosed under “Validation steps and exceptions”.
1. Sources and retrieval
- House Clerk bulk index. The yearly index files
{YEAR}FD.zipfor 2008–2026 were downloaded on 2026-09-30 (the index lists one row per filing: filing type, DocID, date). Years before 2008 are not available from this source. - House filing PDFs. Downloaded 2026-09-30 (14:07–20:01 CDT) from the Clerk's public filing URLs: 7,173 PDFs in total, of which 5,530 are electronic/text PDFs (DocIDs beginning “1”, plus an initial pilot set) and 1,643 are scanned image PDFs (DocIDs beginning “8” or “9”). Electronic annual reports and amendments were fetched for the members who have at least two consecutive report years with an electronic filing, because only such members can form a year-over-year comparison. Scanned annual reports for report years 2013–2025 were fetched only for members who could gain rankability or extra years from validated OCR.
- Senate eFD. The electronic annual-report pages of Senators and Former Senators (calendar years 2013 onward) were retrieved 2026-09-30 (14:41–15:46 CDT): 1,489 report pages, accessed through the site's own click-through agreement flow. Six pages (all one senator, 2013–2018) returned no report content and were skipped. 441 paper (scanned) Senate reports in the index were not fetched or used.
- Roster. unitedstates/congress-legislators (current and historical files, CC0), retrieved 2026-09-30, for names, Bioguide IDs, terms, state and party.
- Portraits. unitedstates/images 450×550 official portraits, retrieved 2026-10-01, resized to 300 px wide; 744 portraits found, 3 not available.
- Pacing and politeness. Requests identify the crawler with a descriptive User-Agent. Pacing is at least 1 s between requests for the index files and at least 2 s for PDFs, Senate pages and portraits. On HTTP 403, 429 or 5xx responses the fetchers back off exponentially and retry without changing identity. No back-off was triggered in the bulk House run.
2. Coverage by year and chamber
One annual report is used per member per report year. The table shows the number of member-years used and the number of comparable year-pairs (member-years) that met the eligibility rules in section 6.
| Report year | House (electronic) | House (validated OCR / X-mark) | Senate (electronic) | Eligible year-pairs |
|---|---|---|---|---|
| 2012 | 0 | 24 | 0 | 0 |
| 2013 | 245 | 0 | 57 | 6 |
| 2014 | 290 | 1 | 72 | 219 |
| 2015 | 327 | 0 | 72 | 274 |
| 2016 | 323 | 0 | 78 | 272 |
| 2017 | 354 | 0 | 80 | 289 |
| 2018 | 318 | 0 | 90 | 282 |
| 2019 | 370 | 0 | 91 | 305 |
| 2020 | 344 | 1 | 90 | 314 |
| 2021 | 383 | 1 | 91 | 324 |
| 2022 | 347 | 1 | 92 | 305 |
| 2023 | 383 | 1 | 95 | 289 |
| 2024 | 352 | 1 | 93 | 324 |
| 2025 | 343 | 1 | 97 | 343 |
| Total | 4379 | 31 | 1098 | 3546 |
Total: 5,508 member-year reports (4,410 House, 1,098 Senate) for 747 members. Of 4,730 consecutive-year pairs with reports on both sides, 3,546 were eligible; the others were excluded because traced coverage was below 60% (685), the traced opening value was below $50,000 (495), the OCR continuity check failed (3) or Schedule B of the later year could not be decoded (1).
What is excluded
- Scanned House filings that did not pass the validation gates in section 8 (1,643 scanned PDFs were downloaded; 78 went through validation; 25 passed and are used; 97 scanned annual-report filings belonging to 9 members' ranked or listed years remain listed as “not used”; most others were never processed).
- Report years 2008–2011 (outside the OCR scope), and House report years before 2013 unless an OCR-validated 2012 filing exists (24 filings); electronic House reports begin with report year 2013.
- Paper Senate reports; periodic transaction reports (PTRs, which are image-only for the members examined); the Senate before report year 2013; filings by people who could not be matched to the roster (candidates, staff and others).
- Defined-benefit pensions (asset code DB) have no value range and are skipped. Items not reportable in disclosures (for example a personal residence) are absent from the data.
3. Parsing
- House electronic PDFs are read with pdfplumber word positions anchored on each page's column headers (Schedule A assets, Schedule B transactions, Schedule D liabilities). Senate eFD HTML is parsed from Part 3 (assets), Part 4a/4b (transactions) and Part 7 (liabilities). Senate owner labels map Self = blank, Spouse = SP, Joint = JT, Child = DC.
- Value ranges are converted to their midpoint (low + high) ÷ 2; “None” is 0. Open-ended top ranges (e.g. “Over $50,000,000”) are excluded from the percentage method and counted as lost coverage; in the net-worth method they are counted at their lower bound and flagged.
- Bracket completion (net worth only): if text extraction kept only “$X −” (the upper bound was lost at a line wrap), the upper bound is filled from the fixed statutory bracket that starts at X (a table lookup), and the row is flagged. In the percentage method such a row counts as an unparsed value and is skipped.
- Rows with a value that cannot be read are skipped and counted: 757 of 254,676 Schedule A rows (0.3%) across the 587 filings affected.
- Multiple rows with the same owner and normalised asset name are summed. The normalised name removes asset-type codes, punctuation and case.
- One filing per report year: the most recently filed among the originals and amendments whose Schedule A row count is at least 40% of the largest (guards against partial amendments); an electronic filing always takes precedence over an OCR-derived one for the same year.
4. Matching
- Filers to members. House Clerk index rows are matched to the roster by last name, state, filing date within the member's term window (with slack) and first-name token (nicknames and prefix fallbacks allowed), with documented tie-breaks (narrower term window, Jr/Sr suffix, most recent term start). The match run left 0 ambiguous rows and matched 1,146 members. Senate eFD filers are matched by name and Senate term; 1,489 filings map to 129 members.
- Positions across years. A position key is the owner code (self / SP / DC / JT) plus the normalised asset name. Positions are matched between year t−1 and year t on this key.
5. Flows and the gain formula (percentage method)
- A position present in both years is “held”. A position present only in year t is included only if Schedule B of year t reports a purchase of it (opening value 0). A position present only in year t−1 is included only if Schedule B of year t reports a sale (closing value 0). Any other appearance or disappearance (gift, inheritance, rename, items below reporting thresholds) is excluded as unexplained. Purchases and sales are midpoints of the Schedule B amount ranges; partial sales count as sales.
- Yearly estimate = (Vt − Vt−1 + incomet − purchasest + salest) ÷ Vt−1, over the traced positions, all at midpoints. Income is the Schedule A income-range midpoint; it is not added for rows whose income type includes “Capital gains”, to avoid double counting.
- Same-position band: the same formula with all ranges at their floors in both years versus all at their ceilings. This is a sensitivity check, not a strict bound. Strict envelope: worst and best case over every independent range, with the denominator fixed at the prior-year midpoint total; it is typically ±100% or wider and is the more honest measure of uncertainty.
6. Eligibility thresholds
- A year-pair counts only if the traced opening value is at least $50,000 and traced coverage (traced opening midpoint ÷ total bounded opening midpoint) is at least 60%.
- A member is ranked with at least 3 eligible year-pairs; exactly 2 is limited data (listed, not ranked); 0 or 1 is insufficient data.
- Counts: 747 members — 480 ranked, 112 limited data, 155 insufficient data; 3,546 eligible member-years.
- Ranking metric: the geometric mean of (1 + yearly estimate) − 1 over the member's eligible years, with each yearly estimate floored at −99% inside the geometric mean (a result at or below −100% is an artefact of range midpoints, not a total loss). The arithmetic mean is also reported. Members are sorted by the geometric mean, highest first.
7. Reliability flags
Year-level flags (shown in each member's table; any flag excludes that year from the “unflagged” top-10):
- high turnover — purchases plus sales exceed 100% of opening value;
- many new/exited positions — new plus exited positions exceed 25% of the number of held positions;
- single position drives the result — the largest mover exceeds 50% of the summed absolute movement of the 12 largest movers and the year's estimate exceeds ±50%;
- income-dominated — reported income exceeds 50% of opening value;
- |estimate| > 300% (likely range artefact);
- OCR-derived — the year uses a validated scanned filing.
Member-level flags: fewer than 5 eligible years (“few years”); trading-dominated years (at least max(2, one third) of eligible years carry the turnover or new/exited flag); average traced coverage below 75%; scanned filings not used; OCR-derived years; a year beyond ±100%; years driven by income, one position or a range artefact; and any unparsed values. 138 of the 480 ranked members carry no member-level flag.
8. Scanned filings: OCR and X-mark decoding
A scanned filing enters the analysis only if it passes every gate below; otherwise the member-year stays “insufficient data – scanned” and the report lists the failed gate. Conflicts are dropped, never guessed.
- Page preparation (rotation, deskew, grid-line removal), then two engines: RapidOCR for layout and text and Tesseract 5 for re-reading each value cell, owner cell and transaction-type glyph. A value is kept only if both engines agree (or Tesseract cannot read the cell and RapidOCR gives an exact legal bracket); values are snapped to the statutory brackets.
| Gate | Rule |
|---|---|
| G1 value coverage | ≥ 90% of Schedule A table rows have a legal value bracket |
| G2 engine agreement | where both engines read a value cell, ≥ 90% agree; conflicts ≤ 5% of rows; trusted-snap rate ≥ 92% |
| G3 page completeness | ≤ 10% of Schedule A/B pages yield no rows; ≥ 1 Schedule A page |
| G4 income coverage | income amount read for ≥ 90% of kept rows |
| G5 continuity | against adjacent report years, for ≥ 5 name-matched positions, ≥ 87.5% within ±1 value bracket (the electronic-vs-electronic baseline is 97.5%, 17,020 of 17,453; threshold = baseline − 10 points); re-checked per year-pair |
| G6 internal consistency | income-type vs. amount agree on ≥ 75% of checkable rows (informational); Schedule B dropped unless ≥ 90% of its rows have a readable P/S/E type |
Result: 78 scanned filings went through the validation step and 25 passed (24 members: 24 filings for report year 2012 and 1 for 2014); they contribute 7 eligible year-pairs (6 in 2013, 1 in 2025 — the latter via the X-mark filings below). Failure reasons recorded (a filing can fail more than one gate): G4 income coverage 34, G3 page completeness 16, G2 engine agreement 14, checkbox (X-mark) layout 10, G1 1, G5 1, other 6. Only a fraction of the scanned queue was processed in the time available.
X-mark (checkbox-column) forms. On these forms a value is encoded by which column holds an “X”. A decoder was written and run for one member's filings (Rep. Ro Khanna, report years 2020–2025); it was not run for other members' X-mark filings. Per cell, an ink-density decoder must see exactly one mark and an independent centroid decoder must pick the same column, otherwise the cell is dropped. Filing-level gates: decoder agreement ≥ 98%, decoder-conflict rows ≤ 2%, dropped value rows ≤ 10%, ≥ 95% of Schedule A grid pages parsed, income read on ≥ 90% of kept rows; Schedule B is used only if ≥ 90% of its grid pages parse and ≥ 90% of its rows have one readable type. Observed: decoder agreement 99.5–100%; all six filings were accepted; Schedule B was usable for 2021, 2022, 2023 and 2025 but not for 2020 or 2024. Of that member's year-pairs only 2024→2025 passed all gates, so the member is “insufficient data” and not ranked. Details are in that member's report.
Accuracy figures (ground-truth sample)
| Field | Correct / total | Accuracy |
|---|---|---|
| Row recall | 174/178 | 97.8% |
| Asset name | 165/167 | 98.8% |
| Schedule A value bracket | 145/147 | 98.6% |
| Schedule A income bracket | 144/147 | 98.0% |
| Schedule A owner code | 141/143 | 98.6% |
| Schedule B amount | 29/31 | 93.5% |
| Schedule B date | 29/29 | 100% |
| Schedule B owner | 29/31 | 93.5% |
| Schedule B type (P/S/E) | 27/31 | 87.1% |
Sample: 15 pages from 12 distinct filings plus one 2019 page of the X-mark member's typed form. The sample is small (treat accuracies as about ±3 points). Misses: 4 rows not found, 2 “None” income values read blank, 4 transaction-type misreads (such rows are dropped by gate G6 when frequent).
9. Estimated net worth and dollar-gain rankings
- Net worth for an annual report = Σ Schedule A asset values − Σ Schedule D (House) / Part 7 (Senate) liabilities. Mid uses range midpoints; Low = assets at range floors − liabilities at range ceilings; High = assets at ceilings − liabilities at floors. Low and High are outer bounds of the reported ranges, not likely outcomes. The personal residence is not reportable as an asset (a mortgage on it can appear as a liability, lowering the estimate). Retirement/federal accounts, excepted or blind trusts, small cash balances, vehicles, household goods and some spouse/dependent items are absent or partial.
- Liabilities are extracted from the filing text (House: Schedule D section via
pdftotext -layout; Senate: Part 7 Amount column). The number of liability rows was cross-checked against an independent count of amount tokens in each filing's Schedule D: 6,909 of 6,909 text-readable filings agree exactly (110 image-only filings could not be checked). - Only electronic filings are used for net worth. Validated-OCR and X-mark years have no parsed Schedule D, so net worth is shown as “not available” for them.
- A year is not computable if the liabilities section cannot be located, any liability amount is unreadable, the filing lists no Schedule A assets, or more than 5% of Schedule A rows have an unreadable value.
- Dollar gain = latest computable year's net worth − earliest computable year's net worth, in nominal dollars (no inflation adjustment): Mid = latest Mid − earliest Mid; Low = latest Low − earliest High; High = latest High − earliest Low. The start is labelled “before entering office” only if the report year is not later than the year of the member's first term; otherwise the earliest available report is used. The change includes salary, gifts, inheritances, business income and revaluation, not only investment results.
- Counts: 727 members have a computable estimate; 20 are listed as not available with the reason.
- Recommended list (550 members) excludes any member with: a span under 3 years (137 members); a year inside the span that is not computable (32); open-ended (lower-bound) asset or liability values or unreadable values at either endpoint (16); a scanned-only filing inside the span that could not be used; or an OCR-derived endpoint. A member can fall into more than one category; 177 members are not in the recommended list. The full list shows all 727 with flags.
10. Top-10 lists
“Top 10 single-year” lists the ten highest yearly estimates among all eligible member-years. The recommended version (“unflagged”) first removes any member-year carrying a year-level flag from section 7 (including OCR-derived), then takes the top ten.
11. Known limitations
- Ranges are wide, so midpoint estimates are noisy; a holding can change range with no real change in value, or the reverse. A range crossing on a few very large holdings can drive a year's result.
- Trading-heavy filers show large values because purchases and sales are themselves ranges and value and flow offsets do not cancel exactly. Account-structure changes between years can hide positions. Blind trusts, residences, some pensions and items under $1,000 are absent.
- The geometric-mean ranking rewards or penalises range jumps on a few large accounts; members with only 3 eligible years rank on thin evidence.
- Scanned filings are mostly not used, so many members have little or no coverage; several earlier report years are missing. Parsing is automated; apart from the sample in section 8 it has not been hand-verified against the PDFs.
- Party labels reflect the member's most recent term in the roster; party history is shown in each report. Figures are estimates, not investment returns, and imply nothing about wrongdoing by anyone.
12. Automation, editorial input, and validation steps
Every number on this site is produced by the scripts listed in the project README, from the sources in section 1. The same thresholds, formulas and flags are applied to every member; no individual member's values, thresholds or ranking position were adjusted by hand, and no editorial content was added to the reports. The following human- or assistant-involved steps are disclosed in full:
- Choice of rules. The ranking threshold (≥ 3 eligible years, changed from an initial 2) and the −99% floor were set by the project owner and the analyst before the full run and then applied uniformly.
- Scope decisions. Which scanned filings were OCR-processed was set by a priority list of members who could gain rankability (199 members); the X-mark decoder was built and run for a single member (section 8), as a case study. These choices affect coverage, not the rules applied to the included data.
- Ground-truth sample. The 15 pages in section 8 were transcribed by the analyst (an AI assistant) from rendered page images and used to measure OCR accuracy; 10 of them were also used while tuning OCR settings and 2 were added afterwards as held-out checks (28 of 28 rows correct for value and income). This sample was used only for accuracy measurement and tuning, never to alter any member's extracted values.
- Visual spot-checks. For the X-mark member, about 12 rendered crops across 2020–2025 (value, income, type and Schedule B amount) were compared with the decoder output: no disagreement found. These checks were read through an image-description tool that was occasionally inconsistent between views, so they are a sanity check, not a transcription; no year is labelled manually verified.
- Name matching. A sample of unmatched “Hon.” index rows was reviewed to improve the generic matching rules; the rules, not individual overrides, were changed.
- Publication. The project owner approved publishing these outputs.
Scripts, thresholds and counts on this page match the pipeline run of 2026-10-01; if the pipeline is re-run, figures may change.