Publish complete table generations
Rollups now retain a complete local generation and, when upload is enabled,
publish it through publication.json. Readers that understand this index resolve
its immutable table paths once per operation. Existing bare Parquet URLs remain
legacy, unversioned data; this change does not rewrite or certify them.
What is checked
A rollup must return every declared output, including a real zero-row Parquet file for successful empty results. Admission checks the exact member set, expected schemas where declared, every data page through a bounded full-column read, row counts, and every byte digest. It uses the installed Rulespec artifact library for membership, canonical identity, manifests and verification.
The resulting artifact records the host implementation digest, installed package
versions, the captured publication index, any unchanged siblings carried
forward by a partial writer, and its parents. A parent is each declared input
the build read: the SHA-256 and size of its bytes (with its family and
generation when a managed family published it, checked against the captured
index), or for an input read in place over HTTP (remote_inputs) its storage
ETag and size, which must not change while the rollup builds. The non-table
docket_search.json.gz has no generation, so only the refresh receipt binds
its parent. It does not establish source completeness, common
publisher timestamps, correct interpretation, or model qualification. The
captured index identifies managed table inputs, not all original source requests.
When R2_PUBLIC_URL is configured, each run uses a new retained directory under
output/.builds/. Old local priors therefore cannot bypass the captured index.
Offline runs without a public URL can still use explicitly retained local inputs;
they produce local candidates and cannot publish. Completed artifacts live under
output/generations/<artifact-digest>/.
Publication and failure
All shrink guards run before any upload. Objects are created conditionally under
generations/<family>/<artifact-digest>/. The publisher rereads and verifies all
remote bytes before conditionally replacing the small publication index. A
concurrent change refuses the update; rebuild from a fresh index before retrying.
A crash before the pointer update can leave unreferenced immutable objects but
keeps the previous published family intact.
If the connection fails after submitting the pointer update, reread the index:
the complete new family may already be current even though the caller saw an error.
A family cannot silently change its table membership or take another family's table. Those changes need an explicit migration. Successful empty tables still face the existing shrink guard; an intentional destructive replacement requires the existing explicit override after reviewing the result.
The Congress.gov archive walker is a partial writer of bill-family. Its bill
update preserves the existing wider columns using the existing merge. It copies
all sibling tables byte-for-byte from the captured complete family and records
which generation supplied them. These siblings are carried forward, not
reprocessed. A public cold start must first produce a complete bill-family
generation; a lone bill file cannot bootstrap a complete family.
An offline cold-start partial candidate is marked local-partial; the generic
publisher refuses it too. Merely retaining a valid Parquet artifact does not
promote that partial result to a complete family.
Regulatory base tables
The ETL rewrites the bare dockets.parquet and documents.parquet after every
sweep batch and primes each run from those bare objects
(r2.download_working_copy), never from a managed family. Once a sweep
completes, the regulatory refresh publishes each as its own family,
dockets and documents (run-rollup-dockets, run-rollup-documents), from
the same working copy, refusing a null or repeated identity. It does this
before the comments mirror captures base versions, so every dependent reads
that sweep's snapshot, and check_refresh_inputs.py holds both family pins
along with the bare objects' ETags. Priming from a family instead would drop
the batches of a partly completed sweep whose keys the manifest had already
retired. Bare-URL readers, such as notebooks and the browser, keep reading
the working copies, which change batch by batch as before.
Before the families publish, fill-docket-gaps (pipelines/docket_gaps.py)
asks the Regulations.gov API once for each docket that documents or the comments
index name and the mirror never captured, and merges what it serves into the
dockets working copy through the ETL's own extraction. It records every other
answer in the bare docket_gap_outcomes.parquet and does not ask again for 30
days: 404 means the publisher does not publish that docket, and 400 "Invalid
ID" means the id is outside its grammar. A legacy -RULEMAKING id whose base
id is served is recorded as an alias of it. If a request gets no answer, the
job fails, but the families still publish whatever it served.
Readers
- Rollup reads share one captured index. Managed download failures and digest mismatches abort; only an index 404 permits legacy resolution.
- MCP captures one index per cached connection, creates managed views at immutable
URLs, checks their schemas, and refuses a connection if a managed member is
missing.
list_sourcesdistinguishes actual available tables from declarations;describe_tablelabels managed generations versus legacy data. Query responses include the publication pin of each table the query names. The remote query reader relies on immutability; it does not rehash whole tables for each query. - The rulemaking dataset publishes under its own pointer,
materialized/rulemaking/latest.json. Each MCP connection build reads it once beside the index (publication.load_rulemaking_snapshot): as the pipeline reads its prior generation, the pointer and its manifest must be one readable format of this dataset naming one snapshot, and each public artifact must sit under that snapshot's prefix. An invalid pointer refuses the connection, as an invalid index does. Views read those immutable keys, so a moved pointer is seen when the cached connection rebuilds (SPICY_REGS_CONNECTION_TTL, default 300 seconds), and a member that cannot be read refuses the connection. The manifest states no columns, sodescribe_tablecompares each view with the dictionary. Status reportsrulemaking_snapshotwith the snapshot id, and the ledger'ssnapshot_…pins compare with it. Local MCP does not read the pointer. - CLI download accepts rollup names. When the requested set includes managed data,
it downloads the requested set into one batch, verifies every managed member,
records the index and selected keys, then switches
currentonly on success. Local reads resolve this link once per command. Requested legacy members remain explicitly unversioned even when downloaded in that batch. - Local MCP accepts that download root, its
currentlink, or a specific batch directory throughSPICY_REGS_DATA_DIR. A connection selects the batch once, rehashes managed members and checks their schemas. It exposes only the selected tables and reports their pins asmanaged_download; loose Parquet directories remainlocal_unversioned. File-change checks around tool statements refuse ordinary replacement or mutation after verification. A newcurrenttarget takes effect when the cached connection rebuilds. These checks do not make writable local storage immutable or establish the source's completeness. - Dictionary remote schema discovery captures the same index and rulemaking
pointer (
publication.published_urls), as doesscripts/check_table_joins.py, and holds each live column's type to its declaration as well as its name. Declared schema pages alone continue to make no claim of production availability.
Auditing a published generation
scripts/audit_generation.py re-reads one family's (or one table's) current
generation anonymously and writes a JSON report. It never writes to storage.
uv run --frozen python scripts/audit_generation.py --family laws --prior 43130abc \
--env-file ../spicy-docs/.env --output report.json
The report keeps its sections separate: publication (every member byte admitted;
index, manifest and root agree), schema (names and types against the dictionary),
identity (duplicate and NULL keys counted before any row is paired), conservation
against a prior pin from the generation's captured chain (both directions, and
multisets where identity cannot pair rows; a table whose build journals a
rows-retired event, as table_merge.retired_rows computes it, must remove
exactly the identities it names, and a removal it does not name, one it names that
the output still holds, or one the prior never held is a failure), evidence (admission, binding,
credential and body-shape scans) and state (an empty table is state, not a
failure). Every report names source qualification and deployment as not assessed,
and limits lists what the run could not see. Exit status 1 means at least one
fail finding; 2 means the audit could not run. check_ledger_pins.py compares
pins only; this tool audits what a pin holds.
Rollout limits
The offline object-store tests exercise conditional creation, interrupted
uploads, concurrent index changes, byte corruption and stale inputs. They are
not a live R2 deployment rehearsal. Validate these operations in a disposable
bucket before enabling the new writer in production. Partitioned comments,
Iceberg publication and the browser's non-table docket_search.json.gz object
are outside this table-generation path; base regulations.gov dockets and
documents joined it as families (above). Historical rebuilds, semantic
qualification and public adoption remain separate work. No publication index is
created remotely by the test suite.