Skip to content

Publish complete table generations

Rollups now retain a complete local generation and, when upload is enabled, publish it through publication.json. Readers that understand this index resolve its immutable table paths once per operation. Existing bare Parquet URLs remain legacy, unversioned data; this change does not rewrite or certify them.

What is checked

A rollup must return every declared output, including a real zero-row Parquet file for successful empty results. Admission checks the exact member set, expected schemas where declared, every data page through a bounded full-column read, row counts, and every byte digest. It uses the installed Rulespec artifact library for membership, canonical identity, manifests and verification.

The resulting artifact records the host implementation digest, installed package versions, the captured publication index, any unchanged siblings carried forward by a partial writer, and its parents. A parent is each declared input the build read: the SHA-256 and size of its bytes (with its family and generation when a managed family published it, checked against the captured index), or for an input read in place over HTTP (remote_inputs) its storage ETag and size, which must not change while the rollup builds. The non-table docket_search.json.gz has no generation, so only the refresh receipt binds its parent. It does not establish source completeness, common publisher timestamps, correct interpretation, or model qualification. The captured index identifies managed table inputs, not all original source requests.

When R2_PUBLIC_URL is configured, each run uses a new retained directory under output/.builds/. Old local priors therefore cannot bypass the captured index. Offline runs without a public URL can still use explicitly retained local inputs; they produce local candidates and cannot publish. Completed artifacts live under output/generations/<artifact-digest>/.

Publication and failure

All shrink guards run before any upload. Objects are created conditionally under generations/<family>/<artifact-digest>/. The publisher rereads and verifies all remote bytes before conditionally replacing the small publication index. A concurrent change refuses the update; rebuild from a fresh index before retrying. A crash before the pointer update can leave unreferenced immutable objects but keeps the previous published family intact. If the connection fails after submitting the pointer update, reread the index: the complete new family may already be current even though the caller saw an error.

A family cannot silently change its table membership or take another family's table. Those changes need an explicit migration. Successful empty tables still face the existing shrink guard; an intentional destructive replacement requires the existing explicit override after reviewing the result.

The Congress.gov archive walker is a partial writer of bill-family. Its bill update preserves the existing wider columns using the existing merge. It copies all sibling tables byte-for-byte from the captured complete family and records which generation supplied them. These siblings are carried forward, not reprocessed. A public cold start must first produce a complete bill-family generation; a lone bill file cannot bootstrap a complete family. An offline cold-start partial candidate is marked local-partial; the generic publisher refuses it too. Merely retaining a valid Parquet artifact does not promote that partial result to a complete family.

Regulatory base tables

The ETL rewrites the bare dockets.parquet and documents.parquet after every sweep batch and primes each run from those bare objects (r2.download_working_copy), never from a managed family. Once a sweep completes, the regulatory refresh publishes each as its own family, dockets and documents (run-rollup-dockets, run-rollup-documents), from the same working copy, refusing a null or repeated identity. It does this before the comments mirror captures base versions, so every dependent reads that sweep's snapshot, and check_refresh_inputs.py holds both family pins along with the bare objects' ETags. Priming from a family instead would drop the batches of a partly completed sweep whose keys the manifest had already retired. Bare-URL readers, such as notebooks and the browser, keep reading the working copies, which change batch by batch as before.

Before the families publish, fill-docket-gaps (pipelines/docket_gaps.py) asks the Regulations.gov API once for each docket that documents or the comments index name and the mirror never captured, and merges what it serves into the dockets working copy through the ETL's own extraction. It records every other answer in the bare docket_gap_outcomes.parquet and does not ask again for 30 days: 404 means the publisher does not publish that docket, and 400 "Invalid ID" means the id is outside its grammar. A legacy -RULEMAKING id whose base id is served is recorded as an alias of it. If a request gets no answer, the job fails, but the families still publish whatever it served.

Readers

  • Rollup reads share one captured index. Managed download failures and digest mismatches abort; only an index 404 permits legacy resolution.
  • MCP captures one index per cached connection, creates managed views at immutable URLs, checks their schemas, and refuses a connection if a managed member is missing. list_sources distinguishes actual available tables from declarations; describe_table labels managed generations versus legacy data. Query responses include the publication pin of each table the query names. The remote query reader relies on immutability; it does not rehash whole tables for each query.
  • The rulemaking dataset publishes under its own pointer, materialized/rulemaking/latest.json. Each MCP connection build reads it once beside the index (publication.load_rulemaking_snapshot): as the pipeline reads its prior generation, the pointer and its manifest must be one readable format of this dataset naming one snapshot, and each public artifact must sit under that snapshot's prefix. An invalid pointer refuses the connection, as an invalid index does. Views read those immutable keys, so a moved pointer is seen when the cached connection rebuilds (SPICY_REGS_CONNECTION_TTL, default 300 seconds), and a member that cannot be read refuses the connection. The manifest states no columns, so describe_table compares each view with the dictionary. Status reports rulemaking_snapshot with the snapshot id, and the ledger's snapshot_… pins compare with it. Local MCP does not read the pointer.
  • CLI download accepts rollup names. When the requested set includes managed data, it downloads the requested set into one batch, verifies every managed member, records the index and selected keys, then switches current only on success. Local reads resolve this link once per command. Requested legacy members remain explicitly unversioned even when downloaded in that batch.
  • Local MCP accepts that download root, its current link, or a specific batch directory through SPICY_REGS_DATA_DIR. A connection selects the batch once, rehashes managed members and checks their schemas. It exposes only the selected tables and reports their pins as managed_download; loose Parquet directories remain local_unversioned. File-change checks around tool statements refuse ordinary replacement or mutation after verification. A new current target takes effect when the cached connection rebuilds. These checks do not make writable local storage immutable or establish the source's completeness.
  • Dictionary remote schema discovery captures the same index and rulemaking pointer (publication.published_urls), as does scripts/check_table_joins.py, and holds each live column's type to its declaration as well as its name. Declared schema pages alone continue to make no claim of production availability.

Auditing a published generation

scripts/audit_generation.py re-reads one family's (or one table's) current generation anonymously and writes a JSON report. It never writes to storage.

uv run --frozen python scripts/audit_generation.py --family laws --prior 43130abc \
  --env-file ../spicy-docs/.env --output report.json

The report keeps its sections separate: publication (every member byte admitted; index, manifest and root agree), schema (names and types against the dictionary), identity (duplicate and NULL keys counted before any row is paired), conservation against a prior pin from the generation's captured chain (both directions, and multisets where identity cannot pair rows; a table whose build journals a rows-retired event, as table_merge.retired_rows computes it, must remove exactly the identities it names, and a removal it does not name, one it names that the output still holds, or one the prior never held is a failure), evidence (admission, binding, credential and body-shape scans) and state (an empty table is state, not a failure). Every report names source qualification and deployment as not assessed, and limits lists what the run could not see. Exit status 1 means at least one fail finding; 2 means the audit could not run. check_ledger_pins.py compares pins only; this tool audits what a pin holds.

Rollout limits

The offline object-store tests exercise conditional creation, interrupted uploads, concurrent index changes, byte corruption and stale inputs. They are not a live R2 deployment rehearsal. Validate these operations in a disposable bucket before enabling the new writer in production. Partitioned comments, Iceberg publication and the browser's non-table docket_search.json.gz object are outside this table-generation path; base regulations.gov dockets and documents joined it as families (above). Historical rebuilds, semantic qualification and public adoption remain separate work. No publication index is created remotely by the test suite.