budget_volumes
President's budget volumes
One row per published volume of the President's budget, with what its print adds to its own MODS index. Thirteen measured package parts are accepted; only volumes offering a supported package-root body yield rows. All columns are stored as VARCHAR.
Coverage. Sampled. BUDGET packages from the 2023-01-01 issue floor whose root PDF follows the provider's thirteen-part grammar, read under a per-run package cap. TAB, DB, CLIMATE and LRB have no root PDF; no granule route is attempted. (measured 2026-09-28)
Data quality. The bill counts are congress-blind and are not join keys. A BUDGET package states no Congress anywhere — not in its id, not in its summary — so the print can only spell HR7806 where the MODS states 119-hr-7806. distinct_bills and distinct_bills_beyond_index_congress_blind therefore reduce both sides to {bill_type}-{number} before comparing, which is why the column carries congress_blind in its name; neither can join congress_bills.bill_id, which needs the Congress this family never states. For the same reason the per-row stated_by_index on a bill_number citation of a budget volume is NULL, not false: the comparison was not possible, and answering false would report a comparison nobody made. Measured across the eight volumes spicy-docs sampled at full page depth, 6 of the 8 distinct printed bills are print-only. The family's real yield is laws: 504 of 518 distinct public laws the prints name are absent from their own MODS at full page depth, against 43 of 69 at a 60-page cap. fiscal_year is not the issue year — BUDGET-2026-MSR was issued 2025-09-05 — and is stated twice, by the package id and by the MODS, which spicy-docs asserts agree. associated_bills_json is published whole because it is the only statement of this edge that carries a Congress.
Thirteen measured parts are accepted by spicy-docs 0.24.0. The seven additions are OBJCLASS, TAB, DB, CLIMATE, LRB, CROSSCUT and DOD. On one measured volume per part, TAB, DB and CLIMATE offer PDF only at granule stems; LRB offers an XLS and no supported body. A GovInfoFormatNotOfferedError at the package root is counted and logged as the publisher's answer, not as a failed or missing volume. No granule route is attempted: shape_budget_volume takes package summary/MODS and does not support a granule identity. These observations remain eligible on the next run and produce no invented document row.
- Parquet file:
budget_volumes.parquet - MCP
query_sqlsupport: Configured; requires an available artifact. - Publication status: Not established by this schema page or its measurement date.
- Row count: Not stated here; the MCP
describe_tablereply gives the live count underpublication.
| Column | Type | Description |
|---|---|---|
package_id |
VARCHAR |
The GovInfo package id, which is this row's identity. |
fiscal_year |
VARCHAR |
The fiscal year this volume is the budget for, as the MODS states it in field name="Fiscal Year", falling back to the year the package id itself carries where the MODS states none. Not the year it was issued: BUDGET-2026-MSR was issued 2025-09-05. |
part |
VARCHAR |
Which part of the budget this volume is, in the publisher's own id spelling (APP for the Appendix, MSR for the Mid-Session Review). The sealed vocabulary is in sources/govinfo/bodies.py's package-id grammar, so a part no measurement has seen is a refusal rather than a row. |
title |
VARCHAR |
The volume's title as the keyed summary states it. |
date_issued |
VARCHAR |
The date the package was issued, as the summary states it. |
last_modified |
VARCHAR |
When the publisher last modified the package; the merge prefers the larger value. |
stated_page_count |
VARCHAR |
How many pages the volume has, as the keyed summary's own pages field states it -- never re-derived from the bytes, so a capped read reports how far it got beside the document's own extent rather than publishing its own count as the volume's. |
associated_bill_count |
VARCHAR |
How many bills the MODS names; every one is in associated_bills_json. |
associated_bills_json |
VARCHAR |
Every root-level <bill> the MODS names, as a JSON array of objects carrying the bill's natural key and the publisher's own context marker, in document order. Published whole because it is the only statement of this edge that carries a Congress: a budget volume's print names a bill without one, so a printed key can be compared with these only congress-blind and can join congress_bills not at all. |
associated_law_count |
VARCHAR |
How many laws the MODS names; every one is in associated_laws_json. |
associated_laws_json |
VARCHAR |
Every root-level <law> the MODS names, as a JSON array of joined laws identities, in document order. This is what distinct_laws_beyond_index is measured against. |
associated_usc_section_count |
VARCHAR |
How many U.S. Code sections the MODS names; every one is in associated_usc_sections_json. |
associated_usc_sections_json |
VARCHAR |
Every <USCode> section the MODS names, as a JSON array of objects carrying the {title}-{section} key and the publisher's own subsection detail, in document order. A chapter-only or appendix-only block contributes nothing: neither is a section and neither has a hosted key. |
associated_cfr_part_count |
VARCHAR |
How many CFR parts the MODS names; every one is in associated_cfr_parts_json. |
associated_cfr_parts_json |
VARCHAR |
Every <cfr> part the MODS names, as a JSON array of {title}-{part} keys, in document order. Budget volumes are the only sampled collection whose MODS states one at all. |
associated_statute_count |
VARCHAR |
How many Statutes at Large pages the MODS names; every one is in associated_statutes_json. |
associated_statutes_json |
VARCHAR |
Every <statuteAtLarge> page the MODS names, as a JSON array of {volume}-{pages} keys, in document order. |
distinct_bills |
VARCHAR |
How many distinct bills the print names, by the shared citation rules. A floor bounded by pages_read, and not a join key: a budget volume states no Congress, so the rule leaves each one as the congress-free HR7806 and congress_bills.bill_id cannot be built from it. The MODS's own bill list, which does carry a Congress, is in associated_bills_json. |
distinct_bills_beyond_index_congress_blind |
VARCHAR |
How many of those the MODS does not state with the Congress dropped from both sides -- the only comparison this family supports, and the one the re-check made when it reported 6 print-only of 8 distinct across the eight sampled volumes at full page depth. Named for what it is: a count, never a key, and weaker than every other *_beyond_index column here, because two measures numbered alike in different Congresses are one value. The per-row document_citations.stated_by_index stays NULL for bills, because the strict comparison a row would have to claim still cannot be made. |
distinct_laws |
VARCHAR |
How many distinct public laws the print names, by the shared citation rules. A floor, not a total: it counts what pages_read reached. |
distinct_laws_beyond_index |
VARCHAR |
How many of those the MODS does not already state. This is the family's headline: 504 of 518 across the eight sampled volumes at full page depth, against 43 of 69 at the rollup's 60-page cap. The per-row document_citations.stated_by_index carries the same comparison exactly; this column is its summary for one volume. |
distinct_usc_sections |
VARCHAR |
How many distinct U.S. Code sections the print names; a floor bounded by pages_read. |
distinct_usc_sections_beyond_index |
VARCHAR |
How many of those the MODS does not already state: 97 of 922 across the eight sampled volumes at full page depth. The MODS states far more than the print here, which is the do-not-recreate rule working in the other direction. |
distinct_cfr_parts |
VARCHAR |
How many distinct CFR parts the print names; a floor bounded by pages_read. |
distinct_cfr_parts_beyond_index |
VARCHAR |
How many of those the MODS does not already state: 13 of 18 across the eight sampled volumes at full page depth. |
distinct_statutes |
VARCHAR |
How many distinct Statutes at Large pages the print names; a floor bounded by pages_read. |
distinct_statutes_beyond_index |
VARCHAR |
How many of those the MODS does not already state: 1 of 81 across the eight sampled volumes at full page depth, which is why this kind carries no contract of its own. |
citation_rows |
VARCHAR |
How many citation rows this volume produced, across every kind. |
pages_read |
VARCHAR |
How many pages the extraction actually read, which a capped read makes smaller than stated_page_count. |
pages_capped |
VARCHAR |
Whether the read stopped short of the volume, so every count above is a floor. NULL where the read states no page split or the publisher states no numeric extent. |
body_rendition |
VARCHAR |
Which rendition the text was derived from. pdf here because the acquirer is asked for sources.govinfo.bodies.PRINT_BODY_PREFERENCE -- the sealed order with PDF first, for the families whose contracts publish a page -- and not because PDF is all a budget volume offers. A package-root format refusal on this family is the publisher's answer, not a missing volume: 4 of the 13 measured parts state no PDF at the package root (one id per part, measured 2026-09-20). TAB, DB and CLIMATE state theirs inside a constituent record at a granule stem (pdf/BUDGET-2027-TAB-1.pdf), which acquire_granule reaches and the package locator does not derive; LRB states one XLS and no body rendition at all. A published/BUDGET walk that meets GovInfoFormatNotOfferedError on one of those has found that shape. |
body_derivation |
VARCHAR |
How that rendition became text. |
text_sha256 |
VARCHAR |
Digest of the normalized text the citation spans index into. |
rule_set_version |
VARCHAR |
Digest over every citation rule's name, version, pattern and rejects, so these counts name the rules that produced them. |