senate_expenditures
Secretary of the Senate expenditure tables
One row per ruled row of one ruled table on one page of a Report of the Secretary of the Senate, with the cells exactly as the print states them and the roles its own header band names. Keyed on package, file, page, table, row and the page text's digest — the package id alone is not a key, because each package publishes Part I (-1.pdf) and Part II (-2.pdf), which repeat the A- summary pages before diverging. All columns are stored as VARCHAR.
Coverage. Sampled: a page window inside each file. Secretary-of-the-Senate expenditure reports reprinted as GovInfo CDOC packages from the 2024-01-01 issue floor, read under a per-run package cap. Every count over this table is a floor on the volume: each report runs to well over a thousand pages and only the first 80 of each file are read, so pages_capped is true on every row. In Part I (-1.pdf) those pages carry the A- summary-of-transactions section and the start of the B- detailed statement; in Part II (-2.pdf) they carry the A- summary again and then its own B- pages from the middle of the statement; the C- compensation-of-members and D- mail-allocation sections lie far past the read window and are not reached (measured 2026-09-20, receipt rollups-pdf-families-2026-09-20/). (measured 2026-09-28)
Data quality. Most rows carry one cell, not a row of cells, and cells_ruled is how you tell. Of the ruled body rows across the pages spicy-docs measured when it wrote the contract (release 0.23.0), those in an appropriation_summary grid carry all nine cells and every other one carries exactly one cell of ten or of seven: the print rules a whole payee block as a single cell. That is why the contract has no payee_name, document_number, date_posted or singular amount column — those would be a parse of one cell's newlines presented as separate facts. WHERE cells_ruled selects the rows whose cells are separate facts; everything else is in cells_json exactly as printed. cells_json is the row; every other content column is a promotion out of it. office, funding_year, appropriation_title and section_heading are read from the page text and are NULL on a continuation page, which is 56 of the 139 measured table pages — forward-fill them in printed_page order. printed_page, not page, is what the report's own contents index by. grid_kind is NULL where the first row states none of the three known labels, so an unrecognised grid is visible rather than mislabelled. An amount must carry a decimal point to be parsed into amounts_json: without that rule a bare integer read stacked fiscal years and account numbers as money, and every amount-shaped line measured carries a .dd tail while the lines that do not are all account numbers or fiscal years. The check that makes the extraction falsifiable is the print's own Totals row, which agreed to the cent with the sum of its account rows over seven money columns on both measured sections.
What each package's read window holds. Part I's pages reach the A- summary and the leadership and officers' B- pages; Part II's reach a short alphabetical run of senators' offices (in GPO-CDOC-119sdoc6: Justice, Kaine, Kelly and Kennedy), so a Senate-wide or other senator's figure is not here; SELECT DISTINCT office for the package says which offices are. The A- summary is by appropriation account only, with no object-class (travel and the like) rollup: object classes are in each office's ORGANIZATION TOTALS block. One report period's spending is spread over one block per funding year (FY2024, FY2025 and FY2026 in the October 2025 to March 2026 report), so sum the blocks' period amounts. A zero prints as $.00.
A payroll cell can interleave two printed lines letter by letter ("D SEIR NE AC TT OO RR" for DIRECTOR and SENATOR, on GPO-CDOC-119sdoc5-2 page B-1470 and 119sdoc6-2 page B-1262), so the cells hold the print exactly for the amounts, which stayed intact, and not always for the text.
- Parquet file:
senate_expenditures.parquet - MCP
query_sqlsupport: Configured; requires an available artifact. - Publication status: Not established by this schema page or its measurement date.
- Row count: Not stated here; the MCP
describe_tablereply gives the live count underpublication.
| Column | Type | Description |
|---|---|---|
package_id |
VARCHAR |
The GovInfo package id the page belongs to, GPO-CDOC-119sdoc3. Not unique on its own: one package publishes the whole report and each of its parts as separate PDFs, so file_name is in the identity beside it. |
file_name |
VARCHAR |
Which PDF of that package this page is in, GPO-CDOC-119sdoc3-1.pdf. Part of the identity because the package's two files, Part I (-1.pdf) and Part II (-2.pdf) as its MODS titles them, open with the same printed pages -- measured byte-identical extracted text over the first 60 pages of GPO-CDOC-119sdoc3 -- before Part II goes on to its own half of the B- statement, and keying on the package alone would collide them. |
page |
VARCHAR |
The one-based PDF page the table was found on, as the extraction states it. |
table_ordinal |
VARCHAR |
Which table on that page, zero-based, in the order the extraction returned them. Measured 139 of 139 pages carry exactly one, so this is 0 throughout the sample and exists because nothing in the print guarantees it. |
row_ordinal |
VARCHAR |
Which ruled row of that table, zero-based, top to bottom. |
text_sha256 |
VARCHAR |
Digest of the page text this row's context was read from, and part of the identity. A page, table and row ordinal mean nothing without it: a re-extraction that finds one more ruled band moves every ordinal after it, and keying on the digest is what stops two extractions of one page from colliding on one identity and silently merging. It pins the text, not the table finder, so a change in the table finder alone is not visible here. |
printed_page |
VARCHAR |
The print's own page label, the last non-empty line of the page text: A-7, B-1243, x. Present on 139 of 139 measured table pages. This is the locator the volume's table of contents indexes by; the PDF page number is not. |
section_heading |
VARCHAR |
The section heading the page states, one of SUMMARY OF TRANSACTIONS BY APPROPRIATIONS or DETAILED AND SUMMARY STATEMENT OF EXPENDITURES. NULL on a continuation page, which repeats neither. |
grid_kind |
VARCHAR |
Which of the print's three ruled grids this table is, decided by the label its own first row states: appropriation_summary, organization_detail or payee_detail. NULL where the first row states none of the three, so an unrecognised grid is visible rather than mislabelled. |
row_kind |
VARCHAR |
What the print makes this row: header where at least two of its cells are header labels, total where it is the grid's own Totals row or carries ORGANIZATION TOTALS, entry otherwise. |
cells_ruled |
VARCHAR |
Whether the print rules this row into more than one cell. true on every appropriation_summary body row (nine cells of nine) and false on every organization_detail and payee_detail entry row (one of ten, one of seven), so WHERE cells_ruled selects exactly the rows whose cells are separate facts rather than one block of text. |
office |
VARCHAR |
The office or account the page names between its section heading and its Funding Year line, SENATOR TIM KAINE; a label the print wraps over two lines is joined with one space (CHAIRMAN MAJORITY POLICY COMMITTEE (D), printed with (D) on its own line), so the Majority and Minority Conference chairs are not both COMMITTEE (D). NULL on a continuation page, which the print leaves unlabelled; a consumer forward-fills in printed_page order. |
funding_year |
VARCHAR |
The funding year the page's own Funding Year line states, and the first of them where it states a span, so a single-year block and a multi-year appropriation are comparable on this column. NULL where the page states no funding-year line, which is every continuation page. |
funding_year_end |
VARCHAR |
The last year of a multi-year appropriation, which the print spells Funding Year 2021-2023 and its own contents spell FY 21/23. NULL where the print states one year, so a span and a single year are told apart rather than folded together. Five of the 83 measured office pages are spans -- the Chaplain's, printed B-48 to B-55. |
appropriation_title |
VARCHAR |
The appropriation the page names under its funding year, joined to one line; the module docstring quotes one in full. NULL where the page states none. |
period_start |
VARCHAR |
The first day of the period the table's own column headers state, as ISO 2024-10-01. Read from FUNDS AVAILABLE AS OF October 1, 2024 or from the first date of a PERIOD OF header, the only two spellings measured; NULL where no header states one, which is every payee_detail table. |
period_end |
VARCHAR |
The last day of that period, as ISO 2025-03-31, from UNEXPENDED BALANCE AS OF March 31, 2025 or from the same header's second date. NULL on the same terms as period_start. |
column_headers_json |
VARCHAR |
One entry per column: the header label governing this row at that index, whitespace collapsed and exactly as the print spells it, or null where the band names none. The band is the nearest run of header rows above this one, which is what makes the role correct in an organization_detail table -- it carries two bands, and column 0 goes from the office block to DOCUMENT NO. partway down. |
cells_json |
VARCHAR |
One entry per column: the cell exactly as the print states it, newlines and all, or null where the extraction found no cell region there. This is the row; every other content column is a promotion out of it and none replaces it. |
account_title |
VARCHAR |
The cell whose governing header is APPROPRIATION TITLE, which carries the account name with its fiscal years stacked under it, or Totals on the section's own total row. Filled only where cells_ruled, so it is never a blob cell read as a title. |
account_number |
VARCHAR |
The cell whose governing header is NO., the Senate's four-digit appropriation account number (0100, 4046). Filled on the same terms as account_title; the print states it empty on the Totals row. |
amounts_json |
VARCHAR |
One entry per column: null where the cell states no amount, otherwise every line of that cell that parses, in printed order, as objects carrying the exact printed line and its decimal. A money cell stacks one amount per fiscal year, which is why this is a list and not a scalar, and a line that does not parse is not here and stands unchanged in cells_json. |
amount_count |
VARCHAR |
How many amounts this row states across every column; 0 where it states none. |
table_row_count |
VARCHAR |
How many ruled rows the table this row belongs to has. |
table_column_count |
VARCHAR |
How many ruled columns it has. Measured 7, 9 and 10 across the sample, one value per grid kind. |
bbox_x0 |
VARCHAR |
Left edge of the table's bounding box, in normalized displayed-page coordinates from a top-left origin. |
bbox_y0 |
VARCHAR |
Top edge of that box, on the same scale. |
bbox_x1 |
VARCHAR |
Right edge of that box, on the same scale. |
bbox_y1 |
VARCHAR |
Bottom edge of that box, on the same scale. |
page_count |
VARCHAR |
How many pages the whole PDF has, as the extraction states it: 1,335 and 1,264 on the two measured volumes. |
pages_read |
VARCHAR |
How many pages the extraction actually read, which a bounded read makes smaller than page_count. |
pages_capped |
VARCHAR |
Whether the read stopped short of the document, so every count taken over these rows is a floor. |
body_rendition |
VARCHAR |
Which rendition the cells were read from; pdf for this family, which offers no other. |
body_derivation |
VARCHAR |
How that rendition became cells and text, as the caller states it: pdf-extraction-lines for the NativeText read these rows were measured on. |
extraction_rule_version |
VARCHAR |
This module's rule version, so a re-extraction under a corrected rule is attributable; the merge prefers the larger value. Zero-padded decimal, because the column is compared as a string. |