Skip to content

bill_committee_actions

Committee-print bill actions

One row per action phrase a committee print states about one bill it names in the same sentence: the print's own phrasing, the sealed stage and BILLSTATUS action codes it maps to, and how reliable the pairing is. Joins document_citations through document_key, text_sha256 and mention_span_start = span_start. All columns are stored as VARCHAR.

Coverage. Derived: the committee actions printed in the accepted House activity reports (house_activity_reports), not all committee actions. (measured 2026-09-28)

Data quality. These are rule-derived interpretations with measured errors. attachment_confidence is single where the sentence names one bill and multi where it names several. Hand-checked over 60 mentions by spicy-docs: a single row is both the right kind of action and the right bill 83.3 percent of the time (30 of 36), a multi row 50 percent (2 of 4). WHERE attachment_confidence = 'single' selects the class with higher measured precision, at an 8 percent volume cost in that sample (4,089 of 4,456 rows were single). It does not qualify an individual row as correct. The vendored column description below still says "hosted-quality"; that wording is superseded by these limits and awaits adoption of the corrected source package. multi rows are kept rather than dropped because sentence_start, sentence_end and matched_text make them readable, and bills_in_sentence is published raw so a consumer can set its own threshold. Precision is not recall: recall is 59.6 percent, re-weighted by stratum over 148 checked contexts. The print writes "the bill" after naming it once, sets an en-bloc disposition as a sentence about "the measures", and states a committee consideration date in a ruled table's column header — so a consumer counting hearings from these rows is counting a floor. sealed_stage is NULL on 1,670 of 4,456 rows and that is a finding rather than a gap: a rung is never invented, so passed the House stays NULL rather than widening the sealed passed house matcher. spicy-docs 0.51.0 reads these phrases through its new stage rules on the next print-citations build: over the 12,695 live rows of 2026-09-29, 2,058 referred rows move from other_chamber to committee, and 476 rows staged law by a phrase that merely contains "public law" become NULL, since the became-law matcher must now start the phrase (5,245 NULL become 5,721; spicy-docs' own sealed_stage over each row's matched_text). billstatus_action_code is NULL on 280 rows whose phrasing has no known code at all (vetoed, not_considered, declined_markup, favorably_forwarded, included_in), and a further 639 rows are coded through codes the publisher's own retained guide never lists — 13 of the 35 distinct codes in the retained responses appear nowhere in its action-code table, which the guide itself says is representational rather than authoritative. stated_date is extracted and published but was never scored: nothing establishes that the date the sentence states belongs to this action. On a 20-bill probe the publisher had no counterpart at all to 10 of the 15 subcommittee hearings these prints state, while all 8 markups were stated and coded — which is what this table is for.

  • Parquet file: bill_committee_actions.parquet
  • MCP query_sql support: Configured; requires an available artifact.
  • Publication status: Not established by this schema page or its measurement date.
  • Row count: Not stated here; the MCP describe_table reply gives the live count under publication.
Column Type Description
document_key VARCHAR The stating document's own natural key in its family's spelling; the GovInfo packageId for a package. Spelled exactly as document_citations.document_key, so the two tables join.
document_kind VARCHAR Which family the stating document belongs to, so one table serves all of them.
bill_id VARCHAR congress_bills.bill_id for the measure this action is attached to, in the citation rule's own spelling. The Congress is the one the document says it covers (house_activity_reports.covered_congress), unless the print states the bill's own -- set right after it (H.R. 6752, 115th Cong.) or in a Congress subheading over it -- because a print mostly writes H.R. 1093 with no Congress beside it; a bill the document states no Congress for has no bill_id and attaches no action. house_activity_reports.bills_congress_mismatch counts the bills the MODS keys under another Congress.
print_phrasing VARCHAR What the print wrote, from the sealed additions-only vocabulary in interpretation.bill_actions.PRINT_ACTION_RULES -- held_hearing, ordered_reported, favorably_forwarded, declined_markup and 21 others. It names the phrasing, never the legislative event: the event is sealed_stage, and a phrasing with no rung keeps its own name rather than being given an invented code. Additions-only because this value is published: a renamed key rewrites rows that are already out.
span_start VARCHAR Character offset of the action phrase in the document's normalized text, counted from zero, and part of the identity. The phrase rather than the bill: one sentence states signed by the President and became Public Law No: 118-83 about the same bill, which is two events' worth of evidence and two rows, and only the phrase offset tells them apart.
span_end VARCHAR Character offset just past the phrase, so text[span_start:span_end] is matched_text.
matched_text VARCHAR The exact characters the phrase rule matched, so a wrong row is readable without the document in hand. The two measured classification failures both show up here: a Pub. L. cite naming the law a bill amends, and reported to Congress inside a bill's own subject matter.
mention_span_start VARCHAR Character offset of the bill designator this phrase was attached to. Equal to a document_citations.span_start for the same document and digest, which is how a consumer walks from an action to the citation row that found the bill.
mention_span_end VARCHAR Character offset just past that designator.
sentence_start VARCHAR Character offset the sentence this was read in begins at, so a consumer can re-read the whole clause and judge the row. The sentence is the unit because the entry is not one: this family writes a bill's long title as its own sentence and the disposition as the fragment after it, and every committee sets an entry differently.
sentence_end VARCHAR Character offset just past that sentence.
bills_in_sentence VARCHAR How many distinct bills that sentence names. This is the reliability predicate, published as the raw count so a consumer can set its own threshold rather than trusting a label: attachment_confidence is derived from it and nothing else.
attachment_confidence VARCHAR single where the sentence names one bill and multi where it names several. Measured joint precision -- right kind and right bill -- is 83.3% for single (30 of 36 hand-checked) and 50% for multi (2 of 4). 4,089 of 4,456 rows are single, so filtering to single excludes the remaining 8% of the volume but does not certify an individual finding. Both classes retain source evidence for verification; an acceptance threshold must be stated separately.
sealed_stage VARCHAR The rung of interpretation.bill_stage's ladder this phrase resolves to, or NULL. Derived by running the matched phrase through infer_stage_from_text, never asserted beside it. NULL on 1,670 of 4,456 rows and that is a finding, not a gap: the print's register is not BILLSTATUS's. passed the House is the sharpest case -- the sealed matcher is passed house, one word away -- and it is recorded NULL rather than closed by widening the sealed matcher list.
sealed_stage_matcher VARCHAR The exact matcher string inside that rung's rule that fired, so the mapping is auditable rather than trusted; NULL wherever sealed_stage is.
billstatus_action_code VARCHAR The <actionCode> values the publisher uses for this phrasing in this row's chamber, joined; NULL where no code is known for it. Both kinds are not in the retained user guide. A House committee hearing is coded H21000 and a markup H15000-B, H15001 or H22000, and none of the four appears in the guide's section 3 -- which states in its own first paragraph that it is representational and that no authoritative list exists. They were read off the publisher's responses instead, and interpretation.bill_actions.GuideCode.source records which of the two a code came from, because only a guide-listed one can be checked against a committed fixture. NULL on favorably_forwarded, declined_markup, not_considered, included_in and vetoed -- 280 rows -- where no code is known from either.
chamber VARCHAR Which chamber's vocabulary the codes were chosen from, read off the row and not off the document: a House committee's report states the Senate passed H.R. 2365 and reports a Senate committee's action on a Senate bill, and both belong to the Senate side while sitting in a House print. The phrase decides it where the phrase names a chamber, and the measure's own type otherwise.
stated_date VARCHAR The first date the sentence states, ISO, or NULL. Unscored: the dates are extracted and published, and whether this one belongs to this action rather than to a neighbouring clause was never hand-checked. stated_date_count says when the sentence stated more than one, which is when the attribution is least safe.
stated_date_count VARCHAR How many dates that sentence states. Above one, stated_date is the first of several and a consumer wanting certainty re-reads the sentence at sentence_start.
evidence_page VARCHAR The one-based page the bill designator sits on, counted in the rendition's own page sequence (for a PDF, the page's position in the file, not the folio printed on it), where the rendition states page boundaries; NULL where it states none.
rule_name VARCHAR Which rule in interpretation/bill_actions.py fired. Equal to print_phrasing today, and a separate column for the reason document_citations.rule_name is one: a phrasing can gain a second pattern and the row must then say which one read it.
rule_version VARCHAR The phrasing vocabulary's own version, moved when a phrasing is appended. Zero-padded decimal, because this column is compared as a string.
rule_set_version VARCHAR Digest over every phrasing's name and pattern, so these rows name the rules that produced them even when someone forgets to move rule_version; an equality token, never a freshness ordering. A successfully corrected generation supersedes its prior regardless of digest spelling.
citation_rule_version VARCHAR The version of the bill_number citation rule that found the designator this row is attached to. Two rule sets produced this row and a re-extraction can move either, so both are published.
body_rendition VARCHAR Which rendition the text was derived from (pdf, htm, xml, txt).
body_derivation VARCHAR How that rendition became text (pdf-extraction-gpo-normalized for a print), which is what the offsets are offsets into.
text_sha256 VARCHAR Digest of the normalized text every offset in this row indexes into, and part of the identity. A span means nothing without it, and keying on it is what stops two extractions of one document from colliding on one identity and silently merging. Append-only per digest, exactly as document_citations is.