Methodology

This page describes what was done, in order, using the project's README and scripts. Figures were generated 2026-10-01. All processing was run by scripts; the rules and thresholds below are applied identically to every member. Any step that involved human or assistant review is disclosed under “Validation steps and exceptions”.

1. Sources and retrieval

2. Coverage by year and chamber

One annual report is used per member per report year. The table shows the number of member-years used and the number of comparable year-pairs (member-years) that met the eligibility rules in section 6.

Report yearHouse (electronic)House (validated OCR / X-mark)Senate (electronic)Eligible year-pairs
201202400
20132450576
2014290172219
2015327072274
2016323078272
2017354080289
2018318090282
2019370091305
2020344190314
2021383191324
2022347192305
2023383195289
2024352193324
2025343197343
Total43793110983546

Total: 5,508 member-year reports (4,410 House, 1,098 Senate) for 747 members. Of 4,730 consecutive-year pairs with reports on both sides, 3,546 were eligible; the others were excluded because traced coverage was below 60% (685), the traced opening value was below $50,000 (495), the OCR continuity check failed (3) or Schedule B of the later year could not be decoded (1).

What is excluded

3. Parsing

4. Matching

5. Flows and the gain formula (percentage method)

6. Eligibility thresholds

7. Reliability flags

Year-level flags (shown in each member's table; any flag excludes that year from the “unflagged” top-10):

Member-level flags: fewer than 5 eligible years (“few years”); trading-dominated years (at least max(2, one third) of eligible years carry the turnover or new/exited flag); average traced coverage below 75%; scanned filings not used; OCR-derived years; a year beyond ±100%; years driven by income, one position or a range artefact; and any unparsed values. 138 of the 480 ranked members carry no member-level flag.

8. Scanned filings: OCR and X-mark decoding

A scanned filing enters the analysis only if it passes every gate below; otherwise the member-year stays “insufficient data – scanned” and the report lists the failed gate. Conflicts are dropped, never guessed.

  1. Page preparation (rotation, deskew, grid-line removal), then two engines: RapidOCR for layout and text and Tesseract 5 for re-reading each value cell, owner cell and transaction-type glyph. A value is kept only if both engines agree (or Tesseract cannot read the cell and RapidOCR gives an exact legal bracket); values are snapped to the statutory brackets.
GateRule
G1 value coverage≥ 90% of Schedule A table rows have a legal value bracket
G2 engine agreementwhere both engines read a value cell, ≥ 90% agree; conflicts ≤ 5% of rows; trusted-snap rate ≥ 92%
G3 page completeness≤ 10% of Schedule A/B pages yield no rows; ≥ 1 Schedule A page
G4 income coverageincome amount read for ≥ 90% of kept rows
G5 continuityagainst adjacent report years, for ≥ 5 name-matched positions, ≥ 87.5% within ±1 value bracket (the electronic-vs-electronic baseline is 97.5%, 17,020 of 17,453; threshold = baseline − 10 points); re-checked per year-pair
G6 internal consistencyincome-type vs. amount agree on ≥ 75% of checkable rows (informational); Schedule B dropped unless ≥ 90% of its rows have a readable P/S/E type

Result: 78 scanned filings went through the validation step and 25 passed (24 members: 24 filings for report year 2012 and 1 for 2014); they contribute 7 eligible year-pairs (6 in 2013, 1 in 2025 — the latter via the X-mark filings below). Failure reasons recorded (a filing can fail more than one gate): G4 income coverage 34, G3 page completeness 16, G2 engine agreement 14, checkbox (X-mark) layout 10, G1 1, G5 1, other 6. Only a fraction of the scanned queue was processed in the time available.

X-mark (checkbox-column) forms. On these forms a value is encoded by which column holds an “X”. A decoder was written and run for one member's filings (Rep. Ro Khanna, report years 2020–2025); it was not run for other members' X-mark filings. Per cell, an ink-density decoder must see exactly one mark and an independent centroid decoder must pick the same column, otherwise the cell is dropped. Filing-level gates: decoder agreement ≥ 98%, decoder-conflict rows ≤ 2%, dropped value rows ≤ 10%, ≥ 95% of Schedule A grid pages parsed, income read on ≥ 90% of kept rows; Schedule B is used only if ≥ 90% of its grid pages parse and ≥ 90% of its rows have one readable type. Observed: decoder agreement 99.5–100%; all six filings were accepted; Schedule B was usable for 2021, 2022, 2023 and 2025 but not for 2020 or 2024. Of that member's year-pairs only 2024→2025 passed all gates, so the member is “insufficient data” and not ranked. Details are in that member's report.

Accuracy figures (ground-truth sample)

FieldCorrect / totalAccuracy
Row recall174/17897.8%
Asset name165/16798.8%
Schedule A value bracket145/14798.6%
Schedule A income bracket144/14798.0%
Schedule A owner code141/14398.6%
Schedule B amount29/3193.5%
Schedule B date29/29100%
Schedule B owner29/3193.5%
Schedule B type (P/S/E)27/3187.1%

Sample: 15 pages from 12 distinct filings plus one 2019 page of the X-mark member's typed form. The sample is small (treat accuracies as about ±3 points). Misses: 4 rows not found, 2 “None” income values read blank, 4 transaction-type misreads (such rows are dropped by gate G6 when frequent).

9. Estimated net worth and dollar-gain rankings

10. Top-10 lists

“Top 10 single-year” lists the ten highest yearly estimates among all eligible member-years. The recommended version (“unflagged”) first removes any member-year carrying a year-level flag from section 7 (including OCR-derived), then takes the top ten.

11. Known limitations

12. Automation, editorial input, and validation steps

Every number on this site is produced by the scripts listed in the project README, from the sources in section 1. The same thresholds, formulas and flags are applied to every member; no individual member's values, thresholds or ranking position were adjusted by hand, and no editorial content was added to the reports. The following human- or assistant-involved steps are disclosed in full:

Scripts, thresholds and counts on this page match the pipeline run of 2026-10-01; if the pipeline is re-run, figures may change.