committee_reports
Committee reports
One row per published part of a captured GovInfo committee report package, keyed (package_id, part_id). bill_id is the measure the report accompanies, read from the package's own MODS metadata. All columns are stored as VARCHAR.
Coverage. Sampled. Selected committee report packages from GovInfo, a mixed-Congress selection that grows as packages are added; not complete historical or constituent-granule coverage. committee_report_reads records each package read and its outcome. (measured 2026-09-28)
Data quality. Only the MODS PRIMARY bill fills bill_id. The thirteen CBO columns come from spicy-docs read_cbo_estimate on the same body already acquired. The cover recital gates publication of a letter: a CBO heading alone is not an estimate. report_states_estimate=false records an absent recital; NULL means the rule has not run. A positive recital with no resolved letter span remains distinct from an extracted letter. recital_bill_id comes from the printed recital and is retained beside the MODS bill_id. bill_id is the part's own PRIMARY bill: a multi-part package's root states none, so each part of CRPT-119hrpt494 carries the H.R. 3495 its constituent record names, where its row read before parts were keyed has NULL; its separate recital_bill_id is a text interpretation. Every part a record states is read, all or none: a package whose parts cannot all be read is refused in committee_report_reads and keeps its prior rows. Parts, by decision 29 (docs/research/fork-delivery-decisions-2026-09-22.md): a report published in one part is one row whose part_id is its package id and whose part_number is NULL; a report published as parts is one row per part, part_id the part's granule id (CRPT-119hrpt455-pt2) and part_number 1 to N, except that a Part 1 the publisher left unsuffixed (CRPT-119hrpt494) has part_id equal to the package id and part_number 1. A package whose one part is spelled -pt1 (CRPT-119hrpt811; spicy-docs docs/tables.md names the others it measured) is (package-pt1, 1), the same shape as Part 1 of two, so a reader counts a package's rows to tell a lone part from Part 1 of several. title, date_issued and last_modified are the package's and repeat on every part row. A row published before part_id existed is kept as its package's one part (part_id the package id, part_number NULL) until the package is read again, which replaces its rows as a set; the committee_report_reads rule version re-reads every report once for this, which is what re-keys the CRPT-119hrpt811 row published 2026-09-24 as (CRPT-119hrpt811-pt1, 1).
- Parquet file:
committee_reports.parquet - MCP
query_sqlsupport: Configured; requires an available artifact. - Publication status: Not established by this schema page or its measurement date.
- Row count: Not stated here; the MCP
describe_tablereply gives the live count underpublication.
| Column | Type | Description |
|---|---|---|
package_id |
VARCHAR |
The GovInfo package id: with part_id, this row's identity. Summary facts (title, date_issued, last_modified) are the package's and repeat on every part row. |
collection |
VARCHAR |
The GovInfo collection the package belongs to (CRPT, CHRG). |
congress |
VARCHAR |
The numbered Congress, parsed from the package id's own grammar. |
report_type |
VARCHAR |
The report's document-type code (hrpt, srpt, erpt). |
report_number |
VARCHAR |
The report's number within its Congress and type. |
chamber |
VARCHAR |
The chamber, from the package id's document-type code rather than from its first letter. |
title |
VARCHAR |
The package title as the keyed summary states it. |
date_issued |
VARCHAR |
The date the package was issued. |
last_modified |
VARCHAR |
When the publisher last modified the package; the merge prefers the larger value. |
bill_id |
VARCHAR |
The bill this package concerns, where a linkage exists; nullable by design (C11). |
format |
VARCHAR |
Which rendition was read (htm, xml, txt, pdf). |
media_type |
VARCHAR |
The response media type, proved against the format before the body was accepted. |
requested_url |
VARCHAR |
The URL the body fetch asked for. |
resolved_url |
VARCHAR |
The URL the body actually came from. |
byte_size |
VARCHAR |
Length in bytes of the captured body. |
sha256 |
VARCHAR |
Digest of the captured body bytes. |
observed_at |
VARCHAR |
When the body was captured. |
page_count |
VARCHAR |
How many pages the extraction read, where a page-based extractor ran. |
text_sha256 |
VARCHAR |
Digest of the extracted text, so a re-extraction that changed nothing is visible as such. |
estimate_rule |
VARCHAR |
Which rule read this report's cost-estimate statement. Named even though there is one today, for the reason document_citations names its own: a second reader over another rendition would have to say which one produced the span. |
estimate_rule_version |
VARCHAR |
That rule's version, a digest over all patterns, flags, rejects, heading thresholds and rule revision, so a re-read under a corrected rule is attributable. |
report_states_estimate |
VARCHAR |
Whether the report's own cover carries the statutory recital Including cost estimate of the Congressional Budget Office, printed in brackets. This is the gate and a heading is never one: three retained reports print a CBO heading over a section saying the estimate was not received, and one prints the estimate under a heading no pattern set had. false is requested-empty with estimate_absence_reason attached, never absence: 3 of 13 reported bills whose index named an estimate had none in the report. |
recital_bill_id |
VARCHAR |
The measure the cover states this report accompanies, keyed the way congress_bills.bill_id is, with the Congress taken from this package's own identity because the cover states none. The print's own answer to the report-to-bill join, which is what cbo_cost_estimates.bill_id joins against; bill_id beside it stays whatever index record the caller read. |
estimate_heading |
VARCHAR |
The section heading the located span opens on, as the committee spelled it, collapsed to one line where GPO wrapped it across two. Committee-specific prose no index states. |
estimate_heading_rule |
VARCHAR |
Which entry of the heading vocabulary matched. That vocabulary is a floor: it only ever locates a span the recital already declared, and one retained report uses a spelling the routes measurement's five patterns missed. |
letter_span_start |
VARCHAR |
Character offset of the reprinted letter in the extracted text text_sha256 digests, counted from zero. The letter is not carried in the row, so the span is how it is re-read. |
letter_span_end |
VARCHAR |
Offset just past the letter, so text[letter_span_start:letter_span_end] is it. |
letter_text_sha256 |
VARCHAR |
Digest of exactly those characters, so a consumer can prove the span it re-read is the span that was measured. A re-extraction that moved one character moves every offset after it, which is why text_sha256 says which text these offsets index into. |
letter_end_rule |
VARCHAR |
How the end was found: director_attribution on the letter's own close, or next_heading_in_series where the report prints a summary table with no attribution at all. A declared letter whose end could not be found publishes a NULL span rather than a guessed boundary. |
letter_signatory |
VARCHAR |
The Director as the letter names them; NULL where the attribution states no name, which the one table-shaped estimate does. |
estimate_absence_reason |
VARCHAR |
The publisher's own paragraph saying why the estimate is not here, verbatim, where the cover declares none. Carried whole rather than as a span, unlike the letter: it is a paragraph, and a row that holds the words needs no offset to read them. |
estimate_absence_rule |
VARCHAR |
Which reason pattern matched: not_available (the estimate had not arrived when the report was filed) or not_received (the committee asked and CBO had not answered). NULL where the report gives no reason, which is not the same as having none to give. |
part_id |
VARCHAR |
The part of the report this row is: the publisher's own granule id for it, which is also the file stem its body was read at (CRPT-119hrpt455-pt2). The package id itself where the part's stem is the package's: a report published in one part, and a Part 1 the publisher left unsuffixed (CRPT-119hrpt494). Never NULL, because it is half the identity; a package's part rows are replaced as a set. |
part_number |
VARCHAR |
The part's number as the publisher states it (its partNumber, or the -pt{N} its id carries, which must agree). NULL on a report published in one part at the package id, whose record numbers none; a report the publisher issued as a lone -pt1 is ({package}-pt1, 1), so a reader counts rows per package to tell it from Part 1 of several. |
body_completeness |
VARCHAR |
What the read text states about itself (sources.govinfo.bodies.publisher_body_status): publisher_placeholder where it is the publisher's own notice that the text is only in the PDF, so text_sha256 digests that notice, not the document; pdf_extracted for text extracted from a PDF, including one read in place of a placeholder; not_flagged otherwise. No value asserts the text is complete. NULL on a row not re-read since the column was added. |
text_derivation |
VARCHAR |
The derivation that produced the text text_sha256 digests (markup-reader, text-rendition-cleanup or pdf-extraction-gpo-normalized), from the body sha256 digests; NULL on a row not re-read since the column was added. |