document_citations
Document citation link table
One row per occurrence of one cited key in one document's text: the key, the exact characters that named it, and the character span it was read at. Citation kinds run from bill numbers and public laws to RINs, agency dockets and GAO product ids (SELECT DISTINCT cite_kind lists them). One shared link table over every document family, written by exactly one rollup. All columns are stored as VARCHAR.
Coverage. Derived, and selective: not a list of the texts the service holds. Every House activity report (govinfo_package) and budget volume (budget_volume) the print-citations family holds is read, and house_activity_reports and budget_volumes state each read (pages_read, citation_rows; citation_rows 0 is a read that found none). The held-field kinds (bill_section, report_section, lobbying_activity, comment_inline, communication_authority, communication_report_nature, communication_record_entry, court_opinion_derived_pdf) are read only for the fields an operator selects, at most 100 a run, so a held text of those kinds has rows here only once it has been selected. A document with no rows may be held and unread: look a committee report up in committee_reports and report_sections, a bill's text in bill_sections. Each row's document_kind is one the resolve_document_citations tool accepts; the tool refuses any other kind, and a kind with no held rows yet resolves to no occurrences. A budget root the reader refuses yields no citations. Every kind is written by the print-citations family. (measured 2026-10-03)
Data quality. The span is in the identity because the span is the yield. Keying on document, kind and target alone would collapse CRPT-118hrpt968's repeated mentions of one bill into one row; the aggregate is derivable from this table with GROUP BY document_key, cite_kind, target_key and the reverse is not. Different text digests retain their histories. A source revision producing a new text_sha256 adds findings beside earlier text versions. A successful rule correction for the same document and text digest replaces that digest's findings, including an empty result; it can retire obsolete rows. A failed read preserves previous findings. Size grows with retained text revisions and their citations. Filter to the digest the document row states (house_activity_reports.text_sha256 or budget_volumes.text_sha256) to read exactly one extraction. Prints are revised rarely, so this is a property to know before building on the table; pruning earlier text digests would require a separate retention decision. stated_by_index is three-valued and the third value matters: true/false where a keyed index record was read and could state a key of this shape, and NULL where no comparison was possible — either no index record was supplied or the MODS vocabulary has no element of that shape, which is the case for every bill citation of a budget volume. So the consumer predicate for "what the print adds" is WHERE stated_by_index IS NOT TRUE, never = false. target_resolved false means the cite was found and its key is unsettled, not that the cite is wrong. target_rule says how a key was reached and separates lookups from inferences: exact and roster_prefix are lookups against a supplied vocabulary, while name_prefix and sibling_prefix are inferences from the printed text that a consumer may decline. evidence_page is NULL where the rendition states no page boundary, and a match straddling a page break is attributed to the page it began on.
A document with no rows here is not shown to be unheld. Each held-field kind has rows only for the documents selected so far (on 2026-10-03 the report_section rows covered one of the report packages report_sections holds); count distinct document_key by document_kind for the live figure. GAO product ids join gao_reports on report_number = target_key, GAO's own spelling (report_id is the same id in lower case); the few that do not resolve are …SU products, which gao_reports does not hold, and misprints such as GAO-21-105520 (GAO-23-105520 is held).
- Parquet file:
document_citations.parquet - MCP
query_sqlsupport: Configured; requires an available artifact. - Publication status: Not established by this schema page or its measurement date.
- Row count: Not stated here; the MCP
describe_tablereply gives the live count underpublication.
| Column | Type | Description |
|---|---|---|
document_key |
VARCHAR |
The citing document's own natural key in its family's spelling: the GovInfo packageId for a package, and one value per family as the table takes each one. |
document_kind |
VARCHAR |
Which family the citing document belongs to, so one table serves all of them. |
cite_kind |
VARCHAR |
What kind of thing is cited, by the rule that read it: bill_number, public_law, statutes_at_large, usc_section, cfr_section, federal_register_cite, rin, gao_product_id, crs_report_id, docket_number, case_docket_number, us_reports_cite or committee_name. |
target_key |
VARCHAR |
The cited key in the hosted target's own spelling, which is part of this row's identity: congress_bills.bill_id for a bill, the joined laws identity for a law, a committee's system_code. Where nothing settled it, the rule's canonical form of the printed text stands instead and target_resolved says so, because an unsettled key is still evidence and dropping it would lose what only the print holds. |
span_start |
VARCHAR |
Character offset of the match in the document's normalized text, counted from zero. Part of the identity: one document naming one bill forty times is forty rows, and the offset is what tells them apart. |
span_end |
VARCHAR |
Character offset just past the match, so text[span_start:span_end] is matched_text. |
evidence_page |
VARCHAR |
The one-based page the match starts on, counted in the rendition's own page sequence (for a PDF, the page's position in the file, not the folio printed on it: BUDGET-2027-PER's PDF page 83 prints "Governmental Receipts 75"), where the rendition states page boundaries; NULL where it states none. A match straddling a page break is attributed to the page it began on. |
matched_text |
VARCHAR |
The exact characters the rule matched, so a false positive is readable from the row without the document in hand. |
target_table |
VARCHAR |
The namespace the target key belongs to, as the rule names it (congress_bills, gao products (product id as GAO prints it)): what the key identifies, not whether a host holds a table of it. |
target_resolved |
VARCHAR |
Whether target_key is the hosted target's own spelling. false means the rule found the cite but nothing settled the key: a bill with no stated Congress, or a committee name no supplied roster reaches. A shape check, not an existence check: a resolved key can name nothing the target table holds; join the target table to learn whether it is held. |
target_rule |
VARCHAR |
How the key was reached. Bill routes distinguish inline_congress, congress_subheading, document_fallback, unstated and document_fallback_refused after the bill_number: prefix. A document fallback is context supplied by the caller, not proof of historical applicability. Other kinds use the rule name except committee_name, where it is the resolution route: exact and roster_prefix are lookups in the roster vocabulary, while name_prefix and sibling_prefix are inferences from the printed text of this one document. A consumer wanting only roster lookups filters on this column rather than on target_resolved. |
stated_by_index |
VARCHAR |
Whether a keyed index record for this document already states this target key -- the package MODS's own <bill>, <law> and <USCode> section elements and its authoring committee. NULL where no index record was supplied. This is the owner's do-not-recreate rule, carried per row: where it is true the row's value is the evidence span, not the key, and WHERE stated_by_index IS NOT TRUE selects the kinds that are genuinely new. |
rule_name |
VARCHAR |
Which rule in interpretation/citations.py fired. Equal to cite_kind today, and a separate column because one kind can gain a second rule -- interpretation/bill_signals.py already runs two patterns for one bill number -- and the row must then say which one read it. |
rule_version |
VARCHAR |
That rule's version, so a re-extraction under a corrected rule is attributable the way prompt_version is; the merge prefers the larger value. Zero-padded decimal, because this column is compared as a string. |
body_rendition |
VARCHAR |
Which rendition the text was derived from (pdf, htm, xml, txt). |
body_derivation |
VARCHAR |
How that rendition became text (pdf-extraction-gpo-normalized for a print), which is what the offsets are offsets into. |
text_sha256 |
VARCHAR |
Digest of the normalized text the offsets index into, and part of the identity. A span means nothing without it -- a re-extraction that moved one character moves every offset after it -- and keying on it is what stops two extractions of the same document from colliding on one identity and silently merging. The table is append-only per digest: a superseded extraction's rows are not retired, and a consumer filters to the digest house_activity_reports.text_sha256 states for that document. |