Skip to content

Fork execution log — September 21–24, 2026

This log holds the dated narrative moved out of the fork output ledger on 2026-09-23: the September 21–22 scope gaps and status bullets, with their "(Superseded …)" markers, and the September 23 continuation. The text moved verbatim; only one pointer changed, "(rows above)" to "(ledger rows)". The ledger keeps one row per output with its qualified pin, plus the open items, and points here for the measurements and receipts behind each row. A later entry supersedes an earlier one where it says so.

The reviewed test exclusions and catch-up retry record the current regulatory recovery. The subsequent source and successful-run audits record report qualification, remaining source defects and the complete daily rollup-log review. Earlier snapshots below remain dated history.

Scope and evidence gaps

  • FEC covers the pinned catalog, retained 2024/2026 committee traversal and 649 explicitly selected collections. Broader individual records, committee history and correction streams remain T05.
  • Full docket/document repairs and docket search are published. All 670 selected source releases replayed; newer and unrelated prior observations survive. This is metadata repair: all document body-text, extraction-status and PDF-extraction-evidence values remain null; acquiring and extracting bodies remains open. The full 23,889,661-row comments object is now retained locally, but three raw witnesses still demonstrate omitted attachments/names/organizations; the current extraction-evidence column is missing. The separate 112,861-group index totals 23,888,128 comments, 1,533 below the independently pinned full object; a common generation is unproved. Comments remain unadmitted. Full reconciliation now identifies 1,555 native NULL-date rows omitted by the old partition rule, offset by 22 surplus counts in three older index groups. All 1,555 originals were acquired and match the parent. Standard Hive null partitions now retain those rows; the local full index conserves every parent row. Wider source-field repairs and the complete dated partition tree remain open.
  • Nominations (2,204) and treaties (2) are published with public-byte and MCP verification at their declared 119th Congress scopes. Record issues have their detail on every row (detail_missing: 0); the other T08 families retain their source/schema/field gaps.
  • Bill-body retry/budget handling is implemented and manually audited across four retained-fixture runs. Independent review found and repaired missing-version/diff-child checkpoints and stale child rows after successful correction. Exact successful scopes now replace their children while failed, capped and unrelated scopes survive. Full 18-table coverage, older backfills and model retry/access remain open. All five report-family members are rebuilt and published for exactly 118 selected packages; one hearing is a fresh source observation with changed bytes.
  • CFR failure paths refuse partial output. Its title 14, volume 4 correction preserves 318,063 unrelated rows and verifies 1,444 repaired rows; source ancestry and other packages remain open. (The September 23 parsing survey found that 1,434 of those 1,444 nulled parts were right; see the CFR bullet in the continuation.) CRS/FCC/GAO/USAspending refusal fixes do not qualify their full populations. CourtListener now uses the source-owned strict reader. A full keyless attempt stopped at HTTP 429 after 34 retained pages/680 records; an exact failure replay preserved prior output. Full scope, usable rate/access budgets and ordinary successful-source retention remain open. The retained Agenda edition now passes full native-field qualification; FR still needs it.
  • (Superseded 2026-09-23: bounded SAM and lobbying initial loads are published; only their schedules wait on the push.) SAM remained withdrawn and disabled pending a qualified initial load; lobbying remained paused.
  • Derived tables need verified parent pins. Monthly volume and discovery are published from the exact repaired parent with accounting/as_of metadata. Discovery's reviewed UTC correction is published and passes independent full-parent counting plus both MCP read paths. Explicit source offsets survive; offset-free dates use UTC, with the policy stored in Parquet metadata. Lifecycles needs a defined pairing/unknown-docket policy: 19 prior null-docket groups collapsed unrelated proposals and 647 groups were lost when the earliest final predated the earliest proposal. The ledger adds the hidden public-comments dependency of organization links and the complete-family publication prerequisite of the narrow bill writer.

The continuation's independent review reports and exact raw/output receipts are linked from the machine ledger. See discovery-utc/, bill-resume-fix-audit/, full-comments/reconciliation/ and courtlistener-refusal/ under the execution receipt directory. These repairs do not mark the unfinished populations complete.

Latest continuation

The 05:30 UTC status snapshot records 20 generated-and-verified outputs, two verified bounded corrections, 29 published outputs awaiting source/parent audit, 15 local candidates and 16 other unfinished outputs. These are 82 distinct output keys, not percentages of work or historical coverage. The fork serves 51 outputs: 48 in the current managed publication index and three base objects with retained public-byte audits. The current index matches the newly reviewed member, report and vote generations.

The 17:13 UTC recheck against the live index supersedes those counts. The index now serves 26 families with 57 table objects, plus the three base objects. Six former local candidates were published without qualification receipts: gao_reports (25 rows), committee_meetings (2,753), house_communications (4,964), the four print-citations tables, record_issues (363) and senate_expenditures (2,623). The later qualification audit (18:10 UTC, scheduled-published-qualification/) now qualifies all eleven objects: public-byte verification passes; fresh source-owned captures reproduce every published cell that still exists in the current source (treaties 48/48, record issues 6,171/6,171, nominations 28,652/28,652, GAO 200/200, house communications 158,368/158,368, committee meetings 74,314/74,331); the two PDF rollups replay byte-identically through the source-owned pipeline against the fork's own prior tables; and actual stdio MCP reads pass in both download and direct-R2 modes for all eleven tables. Recorded drift: 15 house-communication identities the publisher withdrew after the run, four meetings carrying additive publisher updates, plus growth rows the publisher added since. The re-published nominations (2,204) and treaties (2) pins are thereby re-qualified. CourtListener bulk acquisition is complete: all 46 selected objects passed full-file SHA-256 and source ETag verification (67.3 GB), zero pending, acceptance audit passed. The private native docket cache finished at 06:06 UTC with a PASS verdict — 71,677,647 rows, all IDs unique, semantic digest matching the expected value. The later court-dockets delivery (generation sha256:cbb9924411268bae16211aa476004c565dc3a530d93de6741dfb517130a1db54) publishes the enriched APA-899 selection: all 7,766 retained search rows cell-identical plus 3,693 bulk additions (APA-classified rows in all three publisher spellings, and same-case sibling records of retained cases gated on name/date agreement). 31 sealed and 18 docket-number-reuse rows are refused and witnessed; 12 prior sealed rows are preserved as captured. Parties/attorneys/firms stay NULL on bulk-derived rows. Public download and both actual MCP modes verify 11,459 rows. A derived court_docket_groups side table (generation sha256:32fdbde6041d5aa6b2a5ca19e206a2f7d42beda8bd356c5d2225373c8a9b769f, 901 rows) maps the same-case doppeldocket/refiling groups: 403 parents and 498 child links with a confidence tier, parent meant to be the lowest pacer_case_id among published group members because the publisher's own parent_docket_id is blank in the edition. That generation compared the ids as strings, so 21 groups (67 rows) carry a different parent; the repository's build_court_docket_groups compares them numerically and otherwise reproduces every value (candidate independent-rederivation/corrected-groups-candidate/, SHA-256 54b38474…), awaiting republication. Public download and both MCP modes verify the published bytes; every parent id resolves in court_dockets. About 765 selected rows (recounted 2026-09-23) are criminal, magistrate, petty-offense or miscellaneous dockets the publisher codes 899 (383 from the prior search selection, 382 added); excluding non-civil docket types is an open scope decision. Receipts in court-dockets-qualification/; reviews in reviews/court-dockets-enrichment-review.md (approve-with-findings, findings fixed) and reviews/court-dockets-independent-addendum.md (independent re-derivation: enrichment approved exactly, groups parent change requested).

  • Actual scheduled evidence: ordinary fork runs 35689287614 and 35689357725 now pass complete source, public-byte and both actual MCP-mode audits, with independent approval. Members generation sha256:c4b489004884825c499e885a6453290bf6c0957adbbaad5b9feb0faa1ad9b4a1 freshly qualifies all 12,770 members and 45,535 terms. Report generation sha256:f5f169482bfd2903155ebe109268d52fa9be6081061193f980bcd4c537ab129a qualifies all 105 reports, 1,246 sections and the empty root MODS COVER link selection across all 19 hearing parents. Original responses and evidence are retained on the fork. The 19 inherited hearing capture/checkpoint clocks remain unqualified; their native content remains verified. Earlier missing capture evidence is not recreated. See scheduled-retention-live-qualification/final-qualification.json and its independent reviews.
  • Remote court output: the implementation passes focused/full host checks and an isolated real-storage input-to-publication-to-readback experiment. Independent review approves code integration and bounded experiments. Commit 413b3ab includes bounded row/payload batching. A retained-row layout experiment and independent review verify reduced footer overhead and unchanged values; full host and dictionary gates pass. Opinion bodies were withdrawn on 2026-09-22, so no full opinion execution is planned; the reviewed remote path remains available for other large outputs. The experiment creates no real publication-index entry. See court-body-remote-probe/ and reviews/remote-generations-independent-review.md.
  • Native docket cache: the reviewed private build completed the full verified 71,677,647-row source while preserving all 54 literal columns, with a PASS verdict, all IDs unique and the semantic digest matching the expected value. Host mapping, prior-population reconciliation and public delivery remain open. See court-dockets-qualification/.

  • Earlier members/terms qualification: generation sha256:a0140c2f993e032ab53ceec6cd6e57de4b0dda8048ee3ece259d7e009ff8e758 publishes 12,770 legislators and 45,535 terms from both complete September 22 UTC community crosswalk captures. All 639,675 declared cells agree with raw JSON, prior identities/native values survive, and public bytes plus both actual MCP modes pass. The independent literal FEC join agrees on all 1,126 candidate/committee/member links, covering 889 legislators. This is not official-roster completeness; 29 native within-term party histories are not represented by the single-party term value. See members-qualification/ and its independent review reports.

  • Historical scheduled refresh reconciliation: at 03:56 UTC, new member and committee-report generations replaced the previously qualified current pins: sha256:0b0588cc96e965a914094889e23e4d4b359439317d727de05cbccdb8790c7c05 and sha256:925b245110c59ae3a605da56c0ac95fdb79a75aa74c0ecb21868af8398bb07b5. The earlier generations retain their qualification. Full byte/identity/field comparisons pass: every member/term native value is unchanged, and only its capture timestamp differs. All 105 report values are preserved except capture time; 1,246 sections remain byte-identical for the same report parents. All thirteen prior hearings remain unchanged, with six added hearings now matched against 24 retained source responses. All nineteen non-timestamp fields, body/text digests and event IDs agree. Literal package-root MODS inspection proves zero COVER links for all six added parents, so the expanded empty-link selection is qualified within that root-only scope. Changed capture/checkpoint metadata remains outside the earlier qualification because workflow artifacts retain invocation metadata only. The section table keeps its qualification; other current native content and selected root-COVER results are verified, while scheduled capture/checkpoint provenance remains open. See scheduled-members-reports-audit/full-comparison.json, native-content-audit.json and congressional-status/publication-0355.json.
  • Retained source disagreements: hearing packages CHRG-119hhrg63968 and CHRG-119hhrg64154 identify Congress 119 in structured metadata but print “ONE HUNDRED EIGHTEENTH CONGRESS” in their front matter. The output preserves the structured identity and retains both observations; no inferred correction is applied. Exact source hashes and locators are in scheduled-members-reports-audit/native-content-audit.json and MANUAL-AUDIT.md.
  • (Superseded by the September 23 continuation.) Laws/rosters: scheduled generations sha256:8a79810a8dfdfcd77928ff6a03b39f48c3e40caad09a6274695c77185d9b0f3d and sha256:bb48f2816894bf67ad7176aa6ec3ceb6d9f6aa4b2eab95db6e48b94bb61dd3d2 are publicly downloadable. Exact byte/digest/schema/count checks pass for 113 laws, 3,655 law/code rows, 65 Table III rows, 236 committees and 2,966 assignments. The fresh complete chamber files now reproduce every assignment identity and 56,354 comparisons across nineteen columns; the earlier observed_at remains outside this audit. Every member joins the qualified crosswalk, but 138 assignments across fourteen codes lack a committee-table match. Declared conversion equality does not prove canonical committee identity. Committee list/details and law source audits remain open. These families remain published awaiting source audit. See congressional-status/ and rosters-qualification/MANUAL-AUDIT.md.
  • (Superseded by the September 23 continuation.) Amendments: scheduled generation sha256:d48c2567fd20bc911664042cb4b6aba8efa54b3505fb9eb3db2525df2be0b0b4 publishes 7,014 rows. Full public bytes, schema and count agree with the index; native-field/body coverage remains unqualified. See congressional-status/scheduled-amendments-byte-audit.json.
  • (Superseded by the September 23 continuation.) Bill family: the ordinary owner replay now produces a local 18-output candidate with 419,839 bill rows, retaining every prior identity and unrelated value, held body record and diff. It replays all 16,213 receipt-pinned 118th HR/S source records. Independent review accepts partial native qualification: 1,178,929 exact native comparisons pass; 11,021 URLs are absent in raw BILLSTATUS and inherited from the same prior bill identity, with proven lineage. The full-native equality verdict remains FAIL. Broader source provenance, original text retention, missing bodies, models and historical backfills remain open; this broad candidate remains private. A separate scheduled 119th Congress generation now publishes all 18 outputs and 18,956 bills under sha256:15d547a92e4433896ceb015d2f05e410634556be5accd5db5a525aa9bf782ca2; full public-byte checks pass but native source qualification does not. Reconcile that newer generation before replacing any family member. The prior 119th population has 22,064 printing records but only 600 marked captured, and retained 118th HR/S has 19,687 with four marked captured. Those 604 records represent 601 unique packages/digests; three enrolled-bill/public-law aliases share source bodies. The known 40 XML originals are different printings with no package overlap. These flags do not establish possession or new qualification of all originals. See bill-family-continuation/.
  • (Superseded by the September 23 continuation.) Votes: generation sha256:80028c183a1daceee3ac5a6caa6c27b17000abf5781bd609255f49307609fcbf now publishes all 1,573 selected votes and 381,936 member rows. Complete native-field, tally/roster and prior-identity audits pass; public downloads and both actual MCP modes pass independent review. This delivers the additional 86 votes and 20,922 member rows. All 7,865 derived-link cell comparisons match the pinned bill-reference table; its original BILLSTATUS evidence remains unqualified, so roll_call_votes retains that parent gap. member_votes is generated and verified within the frozen selection. Original vote evidence is retained locally but not yet attached to the public generation. Independent full join replay confirms all voter identities resolve, with 18 half-open term-date gaps; inclusive ends leave three gaps and create 1,855 ambiguous House matches. The three remaining literal source observations say “Not Voting.” No term policy or inferred crosswalk is silently adopted. See votes-qualification/publication.json, mcp-audit/summary.json, complete-member-join-replay.json and their independent reviews.
  • Press releases: generation sha256:c3c056ff09697aca92a3c10ab6544114212862aa06a861420cbf4bbe51d759ec serves 28 rows from complete September 19 and September 22 UTC House/Senate Appropriations feed windows. All 952 mapped-field comparisons, public bytes and both actual MCP modes pass. Three rotated-out items survive. That generation had NULL bill links. A later scheduled generation repeated the possessive-s 2027 false match. The reviewed bounded repair sha256:f9bc0f4cedb002c87ece6fd92d1acc9f8b8d53663f402b470dbfef8c193f57e8 removes only that relation, preserves four literal H.R. links and every other current value, and passes full public-byte and both MCP readbacks. Newer scheduled capture metadata is not freshly raw-qualified; the full-source claim remains tied to the earlier captured window union. See press-link-repair/.
  • Unified Agenda: the existing sha256:ea589343f8dcb5ffe134c5a2ac2fbf5d8856f105f296d838fe59e6330ee075b2 generation passes all 67,218 mapped-field comparisons against the complete retained 202510 XML. Actual replay is byte-identical; public and MCP reads agree. This qualifies the retained edition, without a latest-edition or historical-series claim.
  • (Superseded by the September 23 continuation.) Comments: six complete agency source cohorts yield a local 23,889,665-row candidate: all original parent identities plus four BOP records. Every selected native field matches, all 23,862,187 unrelated rows survive unchanged, and the independently reconciled index conserves every row including 1,555 unknown dates. Wider source repair and physical partitions remain open. ACF's fully enumerated next cohort contains 129,052 source comment objects; it has not been acquired in this batch.
  • CourtListener acquisition: initial loading prioritizes the complete June 30, 2026 main export plus unique supplements; historical duplicates remain catalogued. All 46 selected objects (67.3 GB) passed full-file SHA-256 and exact source ETag verification; the detached sequential transfer recorded in opinions-resume-state.json finished the 54.6 GB opinions object, and verified-manifest.json holds zero pending keys. The qualified cluster metadata and the APA docket selection are published. Opinion bodies were withdrawn on 2026-09-22: readers reach the text through each cluster's absolute_url, so the body build and its capacity plan (about 73.3 GB beyond the 100 GiB floor) no longer apply; the retained original is kept. See courtlistener-clusters-qualification/NEXT-WORK.md.
  • Court clusters: generated and verified for the complete June 30, 2026 edition. Generation sha256:7a2cbdb72ad6a7b9c79aa7f25dfa03c4ae3fe752cdb33900cb5d6b346d299b7c publishes 10,070,727 unique clusters in a 3,952,823,520-byte file, SHA-256 c4189ac700bfa4610171a0ccf1642e9f27b7fdfc19c88f801e7c2c970bb927b6. Every prior identity and all 36 old fields are preserved. The full native source agrees, and all three added court fields match the qualified 71,677,647-row docket map and native court reference. The reviewed preflight proves the one court with blank jurisdiction has no docket references; it retains that source row. Full public-download verification and actual MCP reads in both download/direct-fork modes pass, including the count, schema and all 39 fields of six native witnesses. Independent candidate, publication and final readback reviews approve this scope. Opinion bodies are withdrawn (text links out through absolute_url); newer catch-up and hosted MCP deployment remain separate. See courtlistener-clusters-qualification/ and reviews/courtlistener-clusters-public-readback-review.md.
  • CourtListener source behavior: the retained search walk stops at 55 pages/1,100 unique rows. Its reported count of 7,811 is approximate. The corrected provider and host preserve exact-count checks for small docket selections and opinions, and require a completed cursor walk for large docket selections. The independently approved correction was adopted through SpicyDocs 0.26.1 and remains in 0.26.3; full provider/host gates and offline raw/output replay pass. Bulk mapping preserves RECAP source bitmasks and descriptive nature-of-suit text. Standard exports omit party/attorney relationship tables; the dated bulk selection is not assumed equal to the current search index.

The current parallel work assigns integration, bills, court acquisition/builds, vote auditing and independent review. Its September 22, 2026, 05:30 UTC checkpoint distinguishes completed reviews from unfinished transfers and publication. Package integration remains qualified at b50a9eb. Source CI fixes at 51c87a1 and isolated branch b0a8e6e passed GitHub checks. The 0.26.3 adoption is committed at c23a15d: the isolated source gate passes 7,330 tests and the installed host passes 2,426 tests. Only the reviewed vote reader and schema change from the prior wheel. Exact installed-byte, raw regression and independent adoption checks pass. The machine ledger links the same receipts and records the next action per output.

Receipts: members-qualification/, congressional-status/, native-vote-variants-adoption/, bill-family-continuation/, press-release-qualification/, unified-agenda-qualification/, full-comments/source-campaign/, courtlistener-bulk/ and courtlistener-clusters-qualification/ under the linked execution directory. Independent reviews state the exact approved scope and remaining limits.

September 23 continuation

  • Court docket groups republished. The repository builder's rebuild equals the reviewed candidate exactly and differs from generation 32fdbde6… only in the 21 groups (67 rows) whose parent had been chosen by string order. Checked against the raw 71,677,647-row native docket edition: all 901 members exist; each of the 403 groups has one court and docket number and one caption; every parent is the numerically lowest PACER case id; and each of the 21 old parents was the string-lowest. Generation sha256:0f855eb1b094d7405eac1321b276d20f8ca114165ee4c0f09409247126082d49 is live and its public bytes equal the validated build. See court-dockets-qualification/republish-groups/.
  • rulemaking_lifecycles withdrawn (decision 4). Its daily workflow published generation sha256:2f001194cff3819621847f4b5dadb48a9592414dceca2392ffbd184b6641c975 (26,519 rows) at 2026-09-22 23:19 UTC. The family was removed from publication.json by an If-Match write that left the other 31 families unchanged and the immutable objects in place; the workflow is disabled on the fork and no longer scheduled in code. See lifecycles-withdrawal/.
  • Opinion bodies removed (decision 6). The builder, rollup, workflow, tests and catalog/MCP registration are deleted; the push on 2026-09-23 deleted the fork workflow. court_opinions (text-free) now carries the opinion-to-decision map, and the generic remote writer stays for other large outputs.
  • Court citation tables and opinion index published (ledger rows): generations f1e2e523… (court_citations, court_citation_map, court_parentheticals) and f7cc67cc… (court_opinions), the 2026-06-30 edition, public row counts equal to the builds; receipts in court-bulk-tables/.
  • Congress.gov source audits (laws, committee rosters, amendments). Each family was rebuilt from its live source with no published prior (R2 unset), then compared with the live generation by contract identity, cell by cell. Laws (113), Code sections (3,655), Table III (65) and committee assignments (2,966) match in every native cell; only capture times differ. All 236 committees match except publisher activity counts that grew (15 bill counts; none fell). Amendments exposed a completeness defect: an offset walk over updateDate order repeats about as many records as it skips, so it met the declared count by rows while holding 7,013 distinct of 7,066; the live table was 52 short. The builder now pools alternating-order passes until the distinct count reaches the declared count (00f1144), and the complete table (generation 52d60a8f…) is a strict superset of the old one with no shared cell changed; three added amendments were confirmed on the detail route. Detail-only fields (sponsors, amended bill) were empty on every row because the list route does not carry them; the continuation's detail republication fills them. See source-audit-2026-09-23/.
  • USAspending recipients. A fresh capture of the same top-100-page selection through the builder's reader (raw pages retained) against generation 18e51cf5…, whose public bytes match the index: 9,740 of 10,000 recipients remain in today's ranking, with UEI and level identical for all, DUNS for 9,731 and name for 9,730 (15 publisher updates: cleared DUNS, punctuation, one renamed recipient); 3,787 all-time amounts moved as awards accrued; 260 left the top ranks and 114 entered. The ranking itself drifts during a walk (146 ids seen twice on 2026-09-23 06:40 UTC), which the builder refuses rather than publishing. See usaspending-qualification/.
  • CRS reports and FCC proceedings and filings published (T12). All three 2026-09-22 scheduled runs had refused on the source's own data. Each table was replayed from its live source with no prior (R2 unset) and compared with raw pages walked separately, restating the column mapping from the publisher's field names:
  • CRS reports. All 14,137 ids and cells match, except 15 update_date values the publisher re-stamped between the two reads (03:08 and 03:23 UTC).
  • FCC filings. All 5,137 ids and cells match.
  • FCC proceedings. All 21,676 single-document docket names match cell for cell.
  • Causes.
    • CRS and filings repeated an identity within one offset walk; the list shifts while it is read, the amendments defect. pool_passes now pools passes until the distinct count reaches the publisher's count; amendments moved onto it. It compares counts, not identity sets; consolidation plan B6 moves it into SpicyDocs and pools by set.
    • ECFS states no total, but every response carries term aggregations. A window's count is its express_comment (filings) or bureau_name (proceedings) buckets plus the records lacking the field, which matched every walked window exactly.
    • ECFS proceedings are not unique by name. One 2017 document has no name and is left out. Seven dockets hold two documents: 13-84, 15-91, 15-94 and 02-378 were re-created on 2026-09-21 as sparser copies, and 21-62, 24-89 and 25-12 differ in closing or status. The table keeps the original docket, then the last-edited document; the validation confirms it chose the original for all seven.
    • Proceedings are now walked whole every run, because a creation-date increment never refreshed closings and would publish a lone re-created document over its original.
  • Status. Code b2534aa. Generations 6c15aac2…, a1116f71… and c0c1aa4e… are live, with public bytes equal to the validated files. The fork's scheduled workflows run the pushed code, so they kept refusing until this commit was pushed (2026-09-23). See local-candidates-2026-09-23/.
  • Federal Register source audit (T11) and FR docket links (T15).
  • Pin. The live generation 731984ca…'s public bytes match its pin (1,009,005 rows, 1994-01-03 to 2026-09-22).
  • Completeness. The publisher's facet endpoint states true counts where the list endpoint caps at 10,000. All 393 months' row counts equal the monthly facet, and the total equals it (1,009,005).
  • Identity and cells. 74 whole publication days, 60 random (seed 20260923) and the last 14, were fetched with raw pages kept and mapped from the publisher's field names: identity sets and every cell match, 8,683 documents with JSON strings byte-exact.
  • Derived columns. rin equals the array's first element on every row. modify_date is null on every row, because the REST API does not expose it. 474 numbers appear on two dates, the builder's intended (number, date) key.
  • FR docket links. Generation 92f99b00… equals a separate re-derivation from this parent (Python json, not DuckDB UNNEST) as a row multiset in both directions: 899,227 rows, all 16 columns.
  • Dictionary. The stale "2000 floor" coverage on both tables is corrected. See fr-audit-2026-09-23/.

  • Rulemaking dataset bootstrapped (T17). With federal_register and fr_docket_links audited, all five parents qualified. The manifest pins them exactly: FR 47ad1212…, links 2c5f941a…, agenda 52775a37…, and the qualified dockets 27a2ed4a… and documents 7907ebe3… base objects.

  • Date defect fixed. The first local build exposed one: the builder read regulations.gov comment-window instants by their UTC date, but every end stamp is 11:59:59 PM Eastern (03:59:59/04:59:59Z). Every document-sourced close_date was therefore one day late: over the 133,006 documents that also carry an FR comments_close_on, the Eastern day matched for 127,361 and the UTC date for 118. Fixed in 0a898be (actor comment-periods:v5).
  • Stage profiled. rule_targets took 8.5 minutes, quadratic re-serialization of each edge's references. e8d3882 brings it to 38 seconds with byte-identical output.
  • Rebuild checks. In the rebuild every check is zero: primary-id duplicates; unresolved docket, RIN, FR-record, proceeding and agenda-item references; FR evidence lacking the RIN it claims; inverted or unanchored periods. Each period's bounds equal its evidence's earliest open and latest close, recomputed in SQL with the Eastern rule.
  • Live-source samples. 30 of 30 document periods match regulations.gov, 30 of 30 FR periods and 23 of 23 agenda-RIN links match federalregister.gov.
  • Publication. The validated files were published through the pipeline's own gate and publish step, with no rebuild. The public pointer, manifest and five artifacts equal them.
  • Fork workflows. Materialize — rulemaking join surface and Rollup — fcc_proceedings are disabled on the fork. Their pushed code predates 0a898be and b2534aa and would republish the one-day-late closes, and the ECFS re-created dockets over their originals. Re-enable both after pushing (pushed 2026-09-23; both await re-enabling). See rulemaking-2026-09-23/.
  • Vote terms (decision 2). member_vote_terms assigns every member_votes row the term it counts toward: half-open term_start <= vote_day < term_end, then a unique inclusive end, with the term type following the chamber and a Senate LIS id resolved through members. Built from the published votes (3204b8fc…), members and terms, it has 382,136 rows: 382,118 half-open, 15 inclusive-end and 3 unmatched (G000578 on 2025-01-03, S001157 twice on 2026-04-22, all native Not Voting). The non-half-open rows equal the independent replay's 18 exceptions, and all 1,855 boundary-day House rows take the new term. Generation 890481eb… equals the validated build (c113329). Its fork workflow has been active since the push. See vote-terms-2026-09-23/.
  • Comments cohort and T15 rollups (decision 7). The six-agency candidate (DOL, FMC, ITA, BOP, ONCD, PCSCOTUS) was approved by independent review (reviews/comments-six-agency-review.md). On 2026-09-23, 300 random rows re-matched their raw originals on all twelve native fields. The fork bucket held no comments objects, so this was a first publication. comments.parquet (a72a08e9…, 23,889,665 rows) and comments_index.parquet (28e5cf9c…) are live, with S3 read-back and public download digests equal to the reviewed files. The mirror workflow still fails daily for want of its catalog and does not overwrite them. That unblocked three T15 rollups, each rebuilt and recounted independently from the parents:
  • feed_summary (1a833d52…): comment counts per docket equal a direct count over comments.parquet, not the index, and document dates equal documents.
  • agency_stats (d4fa809a…): every docket, document and comment count matches; the totals are 279,124, 2,001,531 and 23,889,665.
  • org_committee_links (7dd4e95a…): every committee resolves with its stated name, and every organization's comment and docket counts match.

80 comments in five dockets absent from dockets are counted by agency but cannot appear in feed_summary, which lists dockets. See full-comments/publication-comments.json and t15-2026-09-23/. - Bill family (decision 1). Generation a846cb44… publishes all eighteen family tables and 419,866 bills. - Inputs. It reconciles the reviewed 419,839-bill candidate (bill-family-continuation/) with the live scheduled generation 83d20e77… (18,998 mostly 119th-Congress bills). The reconciliation used the family's own merge helpers and parameters, with the candidate as prior and the live rows as fresh. At day grain, no live row is older than its candidate counterpart: 1,839 are newer and 17,132 same-day rows are identical in content. (A lexical comparison had flagged 132, which were timestamp against date-only spellings of the same day.) - Checks, every table. No duplicate keys. Every candidate and every live identity survives, and no row appears in neither input. Every live-keyed row equals the live row, and every candidate-only row equals the candidate's. The one exception is 1,752 congress_bills rows, where live NULLs were filled from the candidate by the column-wise merge (1,749 of them url). - URL labels. SpicyDocs 0.28.0 appends congress_bills.url_source, the inherited-provenance label the decision requires; spicy-regs adopted it in 6d34a1b. 12,767 URLs are labelled inherited: 11,018 of the 11,021 lineage bills, plus the 1,749 live rows above. The other three lineage bills were restated by the live narrow writer on 2026-09-18, so they are fresh statements. 5,005 are labelled billstatus, matching the raw archives. Rows that predate the label are NULL. - Publication. All 18 public members equal the reconciled files. - Fork workflows. Rollup — bill family and Rollup — congress_bills are disabled on the fork until 6d34a1b is pushed. Their pushed contract lacks url_source, and their merge drops prior columns it does not know. Pushed 2026-09-23; both await re-enabling. See bill-family-publication-2026-09-23/. - SAM initial load (decision 10). A one-record request confirmed that the workspace SAM key reaches the Entity API: 788,978 active registrations. Retained raw responses then exposed four defects in the bulk-extract path, each fixed in SpicyDocs and adopted here: - Filters dropped. httpx params= replaced the trigger's query, so every selection filter was lost and the API answered its unfiltered default page (0.28.0). - Wrong trigger shape. The trigger answers a plain-text sentence, not JSON with a count (0.28.0). - In-progress read as refusal. A generating file answers HTTP 400 with errorCode FSP (0.28.0). - Wrong identity and count. Registrations are keyed by UEI and EFT indicator: 187 UEIs in 2026 carry several. The file is written while registrations change, so its declared count is a floor: 147,250 declared, 147,256 rows, 147,254 registrations, two held twice (0.28.2; host ebf5c7e adds entity_eft_indicator and a composite merge key).

The bounded load is every active registration dated 2026: one retained 71.5 MB extract (sam-initial-load-2026-09-23/raw/, plus a 508-record one-day extract that matched its paged count exactly). The table equals an independent mapping of the raw file in all 147,254 keys and every cell. Generation 56dd0f65… is live, and its public bytes equal the build. - ACF comments cohort (decision 7). - Capture. All 129,052 comment objects in the complete ACF listing were captured (380,321,010 bytes), and admitted as 128,965 records with 87 discarded observations. - Admission fix. The first admission refused on ACF-2026-0199-0536: two byte-identical Mirrulations refetch files at one modifyDate. A census found 23 such groups, all identical and none differing, so SpicyDocs 0.28.1 (6df4ae2) extended the docket/document rule to comments: identical records collapse, differing ones still refuse. - Repair. The native source fills fields the parent held as NULL on 108,369 existing rows, and adds 738 comments in ACF-2026-0595. No existing value changed, no row was lost, and every other row is preserved. - Review. Independent review (reviews/comments-acf-review.md) approved. It checked all 129,052 objects against S3's own checksums, all 23,761,438 other rows by whole-row hash, and every ACF row against its newest original. - F1, acted on. 0.28.1 had changed the comment policy without moving its version, so current readers refused the published six-agency 1.2 releases with a bare digest error. 0.28.3 (8474b5c) moves comments to policy 1.3; replaying a 1.2 release needs the retained 0.28.0 wheel. Under 1.3, ACF re-admitted to byte-identical staging, and the rebuilt candidate and index equal the reviewed ones as whole-row hash multisets (merge output layout is not byte-stable). - F2, recorded. 161 ACF comments exist on Mirrulations only as _UNAVAILABLE markers, so the cohort is complete over the JSON Mirrulations serves. - Publication. comments.parquet (fca7afb7…) and comments_index (ba31f99f…) replaced the six-agency objects after confirming the live bytes were still those. - T15 refresh. The three T15 tables were rebuilt and recounted on the new parent, with zero mismatches: feed_summary 4c2e2cd9…, agency_stats a9b3d6fc…, org_committee_links e3d05e5b… (1,186 links). feed_summary's ordering broke modify_date ties arbitrarily, so equal inputs gave different bytes; 87cf257 orders by docket_id as well. See full-comments/acf-campaign/. - CFR part ancestry (T11): the live parts are wrong in three shapes; the fix is designed. - What the rule does. Since cac7615, the builder takes no part from a section granule id ("a section token alone does not establish its enclosing part"). Re-applying that rule to every row of the live cfr_sections (8fb97150…) changes 253,758 of 319,507 rows across 260 packages, nulling part and cfr_ref and rewriting section. - Why that is wrong. Most of those citations are correct. GovInfo's granule summary states the literal designation: CFR-2026-title10-vol2-sec100-1 has granuleNumber § 100.1 (part 100), while CFR-2025-title14-vol4-sec19-8-1 has 19-8.1, no § and no part. The granule list does not carry granuleNumber, and the id cannot tell the two shapes apart. - What the volume XML shows. The 262 annual volume XMLs are retained, and one streaming scan compared each section's printed number with its enclosing PART heading. (A 263rd download, GPO-CFR-INDEX-2025.xml, is an HTML page the fetch recorded as a 200 with a digest; it is not a volume.) The live parts are wrong in three shapes: - Title 43 numbers sections by subpart: § 1601.0-1 is in part 1600, not 1601 (3,018 sections). eCFR's ancestry API confirms it. - Title 41's compound parts are cut at the hyphen: 50 for 50-201 (4,732 rows). Every GovInfo-derived field splits there too; the heading and eCFR keep 50-201. - The title 14 vol 4 correction nulled 1,444 parts, but 1,434 of them matched the heading; only the ten 19-8.x sections (part 241, not 19) were wrong.

A first rule that took the part from the section number's prefix was
rejected: it keeps title 43's error and title 41's truncation, and agreeing
with RefSpec's grammar proved nothing because both read the same prefix.
  • The fix. A section-ancestry scan in SpicyDocs' annual CFR reader returns each section's heading part, and build_cfr_sections reads it: 9,250 part changes against the live table: the three shapes above plus six publisher typos and two ranges (consolidation plan A8; ruling 5 decides title 43's cfr_ref). See cfr-ancestry-2026-09-23/ and the parsing survey (spicy-docs docs/research/parsing-survey-2026-09-23.md) §4.
  • Hazard. Until then, a scheduled rebuild with the current code would publish nulls where the table now holds correct parts. Rollup — cfr_sections is therefore disabled on the fork, and stays off after the push until the fix lands: the pushed code includes cac7615. Its last run, on 2026-09-22, failed.
  • Lobbying initial load (decision 10). Keyless lda.gov throttles after 15 quick requests, so the builder now paces keyless pages at 4 s (60443e7). The real builder ran window by window with its own 30-day bound and 7-day overlap: 2026-07-01, 07-23, 08-14 and 09-05, declaring 25,385, 2,543, 803 and 409 filings. Every window's walk matched its declared count, and all 1,218 responses are retained. The table's 27,863 filings equal the distinct raw records and a fresh single-query count of the whole range, with zero cell differences across the mapped columns. Generation a9fd5de6… is live and equal to the build. See lobbying-initial-load-2026-09-23/.
  • Amendments with detail (T08). The builder now reads each listed amendment's detail record (SpicyDocs 0.29.0's amendment-detail route) and takes sponsor, chamber, submitted date and amended bill or amendment from it. A full 119th walk retained all 7,155 responses (two retried 502/503s). The table's 7,095 rows equal the declared count, the distinct listed amendments and the detail responses; mapping from the publisher's field names finds no differing cell, and every one of the 7,066 live identities survives. Sponsor now fills 7,006 rows, amended bill all 7,095, amended amendment 1,094. The other 89 are House amendments whose detail names the Rules Committee, which the member-sponsor columns cannot hold. Generation c943cb37… is live; its table (sha256:f8c3943ace0281c542e7798188f8d64216081b5bc597e4084ee909aea6ce4873) equals the build. Rollup — amendments stayed disabled until the push, whose code carries the detail overlay; it awaits re-enabling. See amendment-details-2026-09-23/.
  • Validation sweep (2026-09-23 afternoon). Four independent scouts re-checked every open item against the live index, the fork's workflows, secrets and receipts, and the raw sources.
  • Carried forward: members 017366cc… equals the qualified c4b48900… except capture times; unified-agenda 5ec1daa4… equals ea589343… exactly. 202510 is still reginfo's newest Agenda edition.
  • Grown and checked: committee-reports 52f30e7c… is a strict superset of f5f16948…. Its 8 new reports and 4 new hearings match GovInfo's summaries, and each body is byte-identical to a fresh download.
  • New votes: 119-senate-2-239 and -240 match the Senate's raw XML in every tally and all 100 positions.
  • Corrections:
    • Record issues are not missing 360 detail bodies (every row has detail).
    • The live bill family has 1,205 captured bodies of 42,448 printings and 276 diff parents, not 604 and 128.
    • The APA non-civil docket count is 765, not about 730.
    • The comments table already covers all 133 comment-bearing agencies; six plus ACF are the natively repaired cohorts.
    • The mirror path is comments/agency/agency_code=<agency>/part-0.parquet.
    • pdf_extraction_results_json is null on every comment, but 70,845 comments carry text.
  • Operations: 10 fork workflows are disabled on purpose; decision 11's later change says which resume after the push (35 spicy-regs and 15 spicy-docs commits were unpushed at the sweep). GEMINI_API_KEY and R2_CATALOG_* are absent, Zyte is unwired, Pages is not enabled, and every recent dictionary deploy fails. See rerun-check-2026-09-23/.
  • Parsing survey (2026-09-23 evening). Six read-only scouts surveyed every parser in the stack. SpicyDocs' docs/research/parsing-survey-2026-09-23.md records the findings, and the consolidation plan beside it (docs/research/consolidation-path-2026-09-22.md) carries them as items A5–A12 and B6 onward. The ones that touch this fork:
  • A key in local logs. bill_subjects sends the Congress.gov key in the query, and its logged 4xx errors carry the URL. No retained log holds a key, and GitHub Actions masks it in CI (A5). The fix was not in the 2026-09-23 push.
  • Published tables with limits. rule_targets, proceedings and comment_periods drop labelled FR docket values (48,169 of 899,227 FR–docket link rows join today; a label-aware reader joins 96,803–106,771 more), and rule_targets misses zero-padded FR numbers (A7). Comment attachment text can be out of order (A6). document_citations keys and bill_vote_references days are recorded on their rows (A10, A11). The CFR parts are in the CFR bullet above (A8).
  • Where the fixes go. SpicyDocs is the only place shared parsing can live without a dependency cycle, and spicy-regs cannot install RefSpec today because RefSpec pins SpicyDocs 0.26.6. Each table change waits on a ruling in the plan (5–7). The scripts and outputs are in parsing-survey-2026-09-23/.

  • Derived comment text (plan A6, decision 19). The Mirrulations derived-text reader in SpicyDocs 0.30.0 (list_docket_derived_text, fetch_derived_text; numeric attachment order, one tool per comment by the pinned order, strict listing, a byte cap) replaced spicy-regs' string-ordered join. A first candidate had labelled 4,764 spicy-regs PDF extractions as mirror rewrites; the reviewer caught it by re-extracting the PDFs, and the staging was rebuilt to 65,994 rows (200c9e55…), every one of which passes the retained validation. The merge needed a configurable DuckDB memory limit (the 4 GB default OOMs on the 23.9M-row parent; 16 GB, 20.2 GB peak, 86 s). The candidate was published at 23:05 UTC after a HEAD and a full authenticated read confirmed the live object was still fca7afb7…; readback from S3 and the public URL equals the candidate. Scripts, build.json, validation.json and publication-comments.json are in comment-text-repair-2026-09-23/.

  • Rulemaking joins (plans A7, A12; decision 18). The rulemaking tables were rebuilt on SpicyDocs 0.31.0: docket values read through the one label-aware reader (the literal fallback was deleted after a measurement showed it changed no join: 279,260 of 279,261 held ids read unchanged, the exception GSA-NA-2005 named by no link), Federal Register numbers keyed by the comparison key that unpads and folds (a match between two sequences padded to different widths is refused: all five on the parents named another document), and instants placed on the Eastern calendar day through regulations_gov_day (equal to the old rule on all 1,488,432 distinct values). A first candidate on the pre-release wheel was reviewed (folds named in status, the predecessor test, the CFR entries the parser drops counted); the release candidate differs from it only where 0.31.0 admits more (-NONRULEMAKING dockets, three more numbers unpadded) and by the 2026-09-23 Federal Register issue. Published at 01:20 UTC on 2026-09-24 through the pipeline's own gate after the live pointer was confirmed twice; readback equals the candidate. Scripts, validation.json, samples.json, verify.json, spot-check.json, differences.json and publish.out are in rulemaking-joins-2026-09-23/candidate-0.31.0/.

  • CFR replace_all on 0.31.0. Run 35940291235 re-placed every volume: 321,010 rows, equal cell for cell to f6192c07… except cfr_ref on 2 rows (cfr-replace-all-2026-09-23/audit.json).

September 24: correction publication and catalog setup

The continuation receipt is ~/Work/corpora/fork-execution-2026-09-21/ledger-continuation-2026-09-24/. SpicyDocs 4310647 (0.32.1) and the adoption in SpicyRegs 2ddc313 were pushed; both CI runs passed. The current output pins and remaining source limits are in the ledger.

  • Meetings and nominations. Existing production builders read Congress.gov through the source-owned pooled reader, retaining their raw responses. The independent field maps from the earlier audit checked every refreshed cell. Both candidates preserve every prior identity. The guarded publisher admitted their source evidence and generation bytes in R2 before moving each family pointer; complete direct public readbacks match the local candidates.
  • Print corrections. The ordinary package cap reprocessed every held print. All changes match the grammar correction receipt, plus one further GSA prospectus number in CRPT-118hrpt974. A second PDF extractor confirms that 0072–OK24 belongs to POK–0046/0072–OK24, not a RIN. All parent and action changes are accounted for, and successful-read checkpoints cover all affected outputs. Publication and direct public readback passed. The local MCP tools also queried all three corrected public families and reported their new generation pins (public-mcp-check.json); this checks the consumer implementation against R2, not a hosted MCP deployment.
  • Docket search. The current public gzip contains the same documents and fields as the qualified delivery; its generated timestamp changed. The retained comparison now records the current digest and ETag.
  • Hosting and credentials. The owner enabled the R2 Data Catalog and GitHub Pages. Dictionary deployment 36016172994 succeeded, and the public site returned HTTP 200. Wrangler's OAuth login cannot administer API tokens (403); the owner's R2_ADMIN_TOKEN from the ignored workspace environment verified as active and accessed the catalog. Its value was installed as R2_CATALOG_TOKEN locally and in GitHub without printing it. URI and warehouse settings are also configured. Catalog setup does not establish seed or ETL completion; follow the ledger and seed runbook for those states.
  • Seed preflight. The ETL was disabled and old pending runs 35966200887 and 35903614693 were cancelled before seeding. The docket seed's dry run used to create its namespace and table before checking the flag. 7c1431a moves writes after the preflight and dry-run return. Tests cover an absent catalog, existing rows, repeated and null source keys, CLI dry runs and a refused undersized source. The live dry run reported the complete source population; a separate catalog read confirmed it created no namespace or table. Docket seed run 36018620120 then loaded the catalog successfully with public upload disabled, preserving the manifest seed's input ETag.
  • Seed completion. OMB smoke run 36018820705 loaded 284,591 comments. Full run 36019107910 then loaded the other agencies, retaining OMB, and ended with 23,890,403 comments. Its freshness check found no lagging agency and no duplicate comment IDs. The manifest check independently proved that every seeded docket and comment ID exists in the catalog, with no duplicate IDs and all public input ETags unchanged. The publisher repeated those checks, uploaded manifest 28439568… and verified its direct public readback at 15:35:34 UTC (manifest-publish.json beside the seed). ETL was re-enabled and run 36021389999 dispatched over all batches with a 240-minute per-batch limit. That sweep's runtime and resulting outputs still require qualification.
  • Verification gap observed before the workflow repair. The live dictionary check stopped at the intentionally withdrawn lifecycle table's 404, and Pages permitted that incomplete result. The continuation below resolves this gap; the earlier deployment alone did not establish complete live schema agreement.

September 24: local catch-up and workflow repairs

This entry records the 17:30–17:45 UTC checks. It supersedes the earlier claims that catalog credentials were missing, the first hosted sweep was still active, and the live schema checker remained blocked by the withdrawn lifecycle table. The output ledger owns the current publication measurements and qualified pins; fork generation owns the remaining work.

  • Catch-up moved locally. The first seeded sweep 36021389999 completed batch 0 and published its manifest last. It was cancelled at 16:36 UTC while batch 1 was still staging; the logs show ongoing CFPB downloads and no merge for that batch. A detached checkout at cded33d started the remaining batches locally at 16:38 UTC, using the existing pipeline with eight agency workers and attachment enrichment enabled. It retains each completed checkpoint and stops at the first failed batch. At 17:32 UTC batch 1 was still running. Hosted ETL, mirror publication, dedupe and comments monitoring were verified disabled during the local writer.
  • The public surfaces are temporarily at different stages. Batch 0 advanced dockets, documents, the comments index and manifest. The public comments monolith kept its earlier ETag and now trails the index's advertised population. The ledger records the exact counts and ETags. This is pending mirror publication; source qualification of the new base outputs also remains open.
  • Batch 1 stopped before publication. At 17:42 UTC, after staging and catalog merges completed, index validation refused two CFTC comments with missing docket_id values. The failure occurred before public upload and manifest save. Public ETags were unchanged at 17:45 UTC, while the catalog had already advanced according to the batch log; the ledger records those counts and the affected IDs. The completion watcher stopped on the failed batch, and all four held workflows remained disabled. Preserve the staging and checkpoints, resolve the missing values against source evidence, validate staging before another catalog write, and reconcile the partial merge before retrying. Neither supervisor is still running.
  • Recovery and publication repairs. a5bb31d adds an append-only dedupe journal, records candidate readiness before dropping live data, and recovers from a complete candidate even when the replacement live table is partial. Normal writes and exports refuse unfinished recovery. Scheduled dedupe now audits only and fails on duplicates; repair and the partition probe require explicit inputs. Failure-injection tests pass. A real R2 scratch-table probe interrupted the first agency copy and recovered every expected ID on retry; its tables were removed. This establishes the repaired recovery path, without implying that production data had suffered the reproduced loss. The mirror now checks retained IDs, uniqueness, docket/month counts and agency partitions before publication; public and raw-catalog checks share that logic.
  • Dependent refreshes. 5715163 connects successful ETL publication to the shared mirror, regulatory summaries, organization links, rulemaking and readback. Base ETags are retained and compared after dependent jobs finish. Vote terms wait for votes and members; Federal Register links wait for their source refresh. Independent source schedules and manual repair entries remain. Shared writer queues retain pending work. The retired lifecycle workflow is removed, while its research producer remains available locally. Execution of the complete new regulatory chain is still pending after catch-up.
  • Cloudflare serving path. Wrangler verified the active catalog, enabled r2.dev endpoint and absence of a custom bucket domain. The configured MCP Worker was absent at that check. The purge credential check now skips this uncached serving path and keeps credential checks for custom domains. No new purge token or Worker deployment was needed for these repairs.
  • Code checks and Pages. The repairs and their documentation were pushed through e944365; CI 36033160292 passed lint, type checks, the full suite and dictionary validation. Strict documentation generation/build also passed. Pages deployed successfully in 36033160021, while its independent live-schema job correctly failed on the then-unpublished report-part fields. The local Actions linter passed with a narrow exception for its unsupported concurrency.queue field; GitHub's documented queue policy is recorded in the workflow review.
  • Report migration and schema closure. The existing report workflow 36032871739 at cded33d published 52325038…. Anonymous reads verified the complete family's declared schemas and row counts, including the multipart identities for CRPT-119hrpt494 and CRPT-119hrpt811. The generation includes more reports and hearings than the earlier migration estimate, and its hearing links are now nonempty. Source and conservation audits remain pending. Attempt 2 of the live schema check passed at 17:28 UTC after this publication. Active schema agreement is now established; the ledger retains the earlier qualified generation until the new source audit completes.
  • Return to incremental scheduling. After recovery, restart the local completion watcher. It requires successful batch receipts, successful CI and the reviewed fork revision. It then dispatches the existing manual mirror/consumer refresh and requires public/catalog readback before restoring scheduled ETL and its checks. A failed check or changed fork revision stops that handoff for reconciliation. This is a prepared continuation; catch-up and the full refresh have not yet completed. Follow the runbook and retained status files.

Receipts under ~/Work/corpora/fork-execution-2026-09-21/:

  • etl-local-catchup-2026-09-24/: pinned runner, cancelled hosted log, batch-*.log, progress.json, completion.json and completion-gate checks.
  • workflow-audit-2026-09-24/: initial reproductions, real-catalog recovery probe, Wrangler observations, local/hosted validation and public readbacks. report-refresh-public-readback.json records the new report family; report-refresh-verification.json records the successful schema recheck; base-public-readback-during-catchup.json records the base-file snapshot; local-catchup-batch-01-failure.json records the failed batch, affected staged IDs and unchanged public ETags.

The workflow review retains the original findings and implementation evidence. Source qualification, historical document enrichment, retained court-edition automation and hosted MCP deployment remain separate work; this operational repair does not close T20.

September 24: CFTC source investigation

The 17:55–18:00 UTC investigation establishes that the missing docket IDs come from the archived source records. It supersedes the earlier uncertainty about whether extraction lost them. Production data, code and workflow state were unchanged during this investigation; catch-up was stopped at that checkpoint.

  • Both CFTC-2026-0595-0003 and CFTC-2026-0595-0005 explicitly contain attributes.docketId: null in Mirrulations. The first is titled as a test; the second describes testing an ex parte meeting display. All mapped native fields equal the staged rows when replayed through the actual host extractor. The complete batch has exactly these two missing docket IDs. A read-only catalog query found one physical row for each, preserving the source nulls.
  • Both comment detail requests to the official Regulations.gov API return 404. Their parent document CFTC-2026-0595-0002, docket CFTC-2026-0595 and neighboring comment …-0006 return 200. The parent and docket are titled Test 4-14-26; the docket is Nonrulemaking. The 404 answers establish current unavailability through those routes, not why the records disappeared.
  • Both comments explicitly name the parent through commentOnDocumentId and commentOn; the live parent matches both IDs and states its docket ID. This supports a separately evidenced parent-derived relationship. It does not replace the null docket field on the source comment, and no ID-prefix guess is required.
  • The immediate failure is the null-docket refusal in transforms/comment_partitions.py, invoked by the index builder after catalog writes. The new comments_health.py check also treats a missing docket relationship as an incomplete identity. Offline replay reproduces both refusals; ordinary null-preserving grouping retains both comments in one index group with exact coverage.

The initial recommendation was unknown-docket support across the application. Review of the actual content led to the narrower decision below: retain the source evidence and exclude these reviewed publisher test payloads from application comments and counts. Missing docket IDs remain a validation error for unreviewed records. No derived docket relationship is needed for this fix.

Evidence: ~/Work/corpora/fork-execution-2026-09-21/cftc-missing-docket-2026-09-24/ contains the exact mirror JSON with listed ETags and hashes, official API responses, staged-row comparisons, selected catalog readback, an offline replay and the findings in README.md. The source mapping and coordinate validator are unchanged between failed checkout cded33d and reviewed code e944365.

September 24: reviewed test exclusions and catch-up retry

The user authorized the exclusion fix after reviewing the source findings. Commit 3953175 separates retained source evidence from application admission and validates staged comments before persistent dataset merges.

  • Bounded admission rule. reviewed_comment_exclusions.json records each reviewed ID, source locator, exact raw-byte digest, canonical JSON digest, reason and evidence pointer. The transform excludes a record only when its ID and complete canonical payload match. A changed payload under either ID fails for review. It does not filter on titles, HTTP status or missing docket fields. Successful batches checkpoint the consumed source keys so deliberate exclusions do not repeat indefinitely. The exact JSON is retained both in the acquisition receipts and as bounded regression fixtures.
  • Validation before writes. The pipeline checks staged comment coordinates before merging documents, dockets or comments. Direct catalog comment merges perform the same check. The existing missing-docket refusal stays in force. Tests cover ordinary and chunked ingestion, consumed-key accounting, changed reviewed payloads and refusal before any catalog merge. The full test suite, lint, types and dictionary checks pass. A built wheel includes the decision file and loads the transform successfully through an isolated import.
  • Catalog cleanup. A real R2 scratch-table probe verified the bounded deletion across a reopened connection; the scratch table was then removed. This does not establish general DELETE/INSERT replacement behavior. The recovery checked every mapped native field against retained source JSON and saved the complete catalog rows before deleting only the reviewed IDs with their observed agency, null docket and modification date. A separate reopened connection at 18:18 UTC confirmed 23,972,434 rows and distinct IDs, neither excluded ID present and no null docket IDs. Neither row was in the public comments file. The original observations remain available for audit.
  • Retry from a verified checkpoint. The failed attempt's local manifest matched a fresh public download exactly, SHA-256 db964dd4accddf007ae4050468d424f44b0d02c428ad2a9f20661fdcac5e3b59. The detached checkout is pinned to 3953175; retry started at 18:21 UTC. It runs batches 1 through 14 through the ordinary CLI with eight agency workers and enrichment enabled. output-retry-02/ starts with only the verified manifest; the pipeline downloads its other public parents normally. Original staging remains in output/, with the first runner, status and logs preserved in attempt-01/. At 18:28 UTC, the live retry logged both exact reviewed exclusions while processing CFTC. The retry stops at its first failure; this observation does not establish completed batch publication.
  • Completion remains gated. Hosted ETL, mirror publication, dedupe and comments monitoring remain paused during the local writer. The completion watcher requires every batch receipt, successful CI at its reviewed fork revision and successful public/catalog readback from the existing refresh workflow before restoring schedules. Catch-up, the full dependent refresh and source qualification are still open; this fix does not close T20.

Recovery receipts are in ~/Work/corpora/fork-execution-2026-09-21/cftc-test-exclusion-repair-2026-09-24/: delete-probe.json, recovery-preflight.json, excluded-catalog-rows.json, removal-started.json, removal-verified.json and the verified public manifest. The local catch-up directory holds current progress.json, completion.json and batch-*.retry-02.log files.

September 24: source and successful-run audits

The user requested delegated bill publication and report/hearing qualification, then a complete review of successful Run rollup logs over the preceding day. Independent agents owned those tasks; the root reconciled their evidence and updated the campaign ledger.

  • Bill metadata repair published and qualified. Ordinary workflow 36045860325 at 71b78b6 published 4a1949a9… at 19:10 UTC. Both Congresses 118/119 and all bill types were selected, with new body fetches capped at zero. All 18 public members pass full-byte admission. Every prior identity survives; no duplicate, orphan bill reference or changed row outside those Congresses is found. The family has 419,978 bills. All 3,044 known damaged dates and 3,095 damaged URLs match current raw BILLSTATUS, including five affected 118th URLs. Thirty-one date repairs reflect later native publisher updates. Independent native checks cover all 38,376 captured bill records and their action and publisher-summary fields. Held sections, diffs and model outputs stay byte-identical.
  • Legacy URL provenance repaired at the owner. e69e483 makes a non-null bill URL with missing provenance trigger a source reread, alongside the retired list-reader label. Final publication repairs 38 additional URLs and labels 546 source-present legacy rows. Of these, one has no native URL and retains its prior value as inherited. Six reserved House bill identities (6, 9, 11, 13, 16 and 19) have no record in current archives and remain entirely unchanged. The initial population check conflated absent records with absent URL fields; its failed qualification and corrected scope are retained. Regression tests and full CI pass; the final source/public gate passes at 19:13 UTC.
  • Bill limits remain explicit. The larger-value merge retains 389 later prior timestamps. Fresh Congress.gov detail for 119-hr-3446 and 119-hr-940 states September 19 while current BILLSTATUS states July 17; an ordinary API delta path remains open. Historical rows, missing original bodies, models, broader interpretation rules and routine source retention retain their earlier qualification limits. The failed earlier run 35946820267 had correctly refused to drop eight archive checkpoint rows under its body-fetch cap. 24eac7d keeps every visited folder's row while recording completion separately; these successful ordinary publications verify that repair.
  • Report migration qualified within its captured scope. Public generation 52325038… and source evidence 35ab12ad… pass byte admission and independent JSON/XML/HTML checks. All 137 report parts match their own native metadata and bodies; all 1,757 sections have unique identities and contiguous source coordinates. Every earlier body and section content survives. The 108 hearings read by this run match their captured sources, while the other 25 retain every prior cell. All 269 read checkpoints reconcile to sources or unchanged prior rows. These checks advance the qualified pins for reports, sections, transcripts and read status, with the documented empty-heading absorption rule preserved.
  • Hearing-link qualification stays partial. All 78 COVER relationships match native MODS. One, CHRG-117shrg56721 / 117-s-2792, has 11 native dates and wrongly presents the first as a unique held_date. Its printed cover supports the relationship; deleting the link would discard source evidence. Retain all dates in SpicyDocs, publish a scalar only when unique, and invalidate the CHRG checkpoint rule after adopting that source change. Six detail 404s remain retryable with their GovInfo bodies retained. Native historical IDs CHRG-79jhrg79716p19 and CHRG-79jhrg79716p11 remain excluded by the package grammar before acquisition/checkpointing. Fresh summaries confirm both exist. No supported package was deferred by the run cap.
  • Every successful rollup step in the fixed daily window was read. The interval is September 23 at 18:44:20 UTC through September 24 at 18:44:20 UTC, using final job completion times. Full pagination yields 85 successful executions: 34 have the exact step, comprising 2,489 retained lines, and 51 have no such step. Root independently enumerated the same run IDs and checked the complete log hashes. The audit report distinguishes current gaps, later repairs and expected reuse or caps.
  • Current acquisition and metadata gaps. Two private-law XML bodies were captured and then refused for absent native Statutes at Large citations, leaving public rows falsely marked not_requested. Fresh native bodies and digest-verified public rows reproduce the problem. Table III's repaired traversal still needs qualified publication; the successful bill-subjects build had uploads disabled. Communications and print discovery have explicit bounded backlogs. Most ordinary external-reader paths still need routine response retention. These findings do not establish loss of prior identities.
  • USAspending wording corrected. The column description still called the endpoint's trailing-12-month amount “all-time.” Commit 150dabc corrects the source description and generated catalog, pages and metadata; 71b78b6 restores the required coverage-kind prefix after CI caught its omission. Dictionary tests and full CI pass at that revision. The current 10,218-row table retains 218 recipients outside the latest 10,000-row selection with older amounts and no per-row observation date; that schema limitation stays open.

Receipts under ~/Work/corpora/fork-execution-2026-09-21/:

  • bill-family-publication-2026-09-24/: source archives and listings, public bytes, identity/value comparisons, repair qualification and explicit limits; pass-1/ retains the earlier f1d04b73… publication and its audit.
  • report-hearing-qualification-2026-09-24/: admitted public/source artifacts, independent verification, fresh source captures and precise remaining limits.
  • rollup-success-log-audit-2026-09-24/: coverage manifest, every complete step/job log, per-run dispositions, raw private-law checks and final hashes.
  • parallel-delivery-2026-09-24/: independent successful-run inventory and log-coverage checks.

September 24 comment-text refactor

Batch 1 completed at 19:23:44 UTC and batch 2 at 19:48:52 UTC on 3953175; batch 3 then began on that same pinned checkout. Batch 1 spent 57m53s of its 62m27s in source staging, with CFPB forming the long agency tail. That phase combines raw downloads and inline text reads; the log alone does not isolate their costs.

Inline ETL and backfill now share bounded comment-level concurrency, with a worker-owned S3 resource and a shared once-per-docket listing cache. Independent text retry state survives raw-key checkpoints. Local batch drivers can reuse a committed manifest, avoiding its repeated rebuild. Progress and phase timings make the remaining work visible. The fixed CFPB comparison showed identical text/provenance and a 6.4-fold median text-fetch improvement at eight workers; see the comparison and qualification limits.

Local tests cover normal and chunked recovery through a fresh workspace, metadata preservation, access refusal, failed writes, global worker bounds and manifest reuse. The handoff runs the new code from a separate pinned checkout after a committed batch. etl-local-catchup-2026-09-24/handoff.json and progress.json distinguish the queued transition from the actual active revision. Mirror publication, dependent refreshes and source qualification remain separate gates.

September 25 local comments export and publication

The catch-up completed every batch at 01:41 UTC. Its hosted completion run 36083102561 then failed during the full comments export: DuckDB could not allocate another block within its 6 GB memory budget. The failure preceded mirror upload, and scheduled ETL remained paused.

The existing publisher now accepts an export memory budget and thread count. The local recovery used a 16 GB budget and two threads, with disk spilling and the existing retained-ID, unique-ID and coverage checks. It ran from an isolated checkout of 03541c7 plus the retained export-memory.patch; progress.json records the patch digest. The full unit suite, Ruff, type check and data dictionary check passed. These resource changes and documentation remain local; the hosted export has not been re-qualified.

Publication completed at 10:11 UTC. The combined file holds 26,303,691 comments, 2,413,288 more than its predecessor, with every prior ID retained. Its SHA-256 is b90e1105332ec7e0c399708724ed479c6c30c6ff066525e8bd10103c48df9845 and its public ETag is 45071b2b74ccec68498c23f35d1b3675-509. The index has 143,367 groups, including three genuine source-null docket relationships. The agency files cover the same population across 180 agencies. Every public object matches its local size and ETag; the receipt retains full file digests. Anonymous public and raw-catalog checks both passed at 10:16 UTC.

The existing feed, agency-statistics, monthly-volume, docket-search, discovery, organization-link and rulemaking commands then completed their refreshes. Base ETags remained unchanged throughout. Rulemaking published snapshot_6d3dc0f22923d60e5d6e78f4e60c49ea; its public manifest and all output digests match the local generation. Docket-search bytes also match. This records successful publication and integrity checks; source qualification of newer dependent generations remains a separate ledger task.

Comments monitoring and the read-only duplicate audit are enabled again, and the manual mirror entry is available. Scheduled ETL stays paused until the hosted exporter completes within its runner limits. One empty FWS source response and three derived texts above the 64 MiB limit remain in the published retry checkpoints. The prior failed completion receipt is archived, and the catch-up status now points to the successful local recovery.

Evidence: ~/Work/corpora/fork-execution-2026-09-21/comments-local-export-2026-09-25/: completion.json, public-object-verification.json, consumer-progress.json, consumer-public-readback.json, refresh-inputs.json, the publication-index snapshot, workflow states, validation log, code patch and operation logs.

September 25 comments publication efficiency — local implementation

The accepted design is now implemented locally. The catalog supplies one pinned snapshot scan into agency staging. Each agency sorts once; the compatible monolith streams its final files. Shared row/byte targets and resource settings replace independent large buffers. The index recount moves from each ingestion batch to the final mirror. The sweep discovers agencies and loads its Bloom membership once, with portable byte storage and unchanged membership semantics.

Publication retains the exact prior IDs, uniqueness, null relationships and agency/docket/month checks. An agency move that empties an old agency writes an empty replacement to its public URL. A completion receipt advances only after all public bytes match local digests and stable storage versions. Unchanged snapshot/schema/exporter identity plus matching object versions skips payload work. A failed batch stops the sweep before later checkpoints can retire its pending keys. Manual CLI and workflow finalization use the same publisher.

The full unit suite, Ruff, type check and generated data dictionary checks pass. Local full-corpus receipts and comparisons are under comments-efficient-publisher-2026-09-25/. Hosted qualification and the browser workload remain separate gates. These changes and evidence remain local; no public objects, workflow state or ETL schedule were changed in this work.

September 25 parallel source/output audits

Three medium-effort auditors checked congressional people/votes, congressional documents and external sources while the parent agent replayed regulatory outputs. The frozen index, native inputs, public bytes, prior-population comparisons and independent cross-reviews are retained under parallel-rollup-audit-2026-09-25/. The audit report records the method and scope; audit-summary.json preserves normalized dispositions and result-file hashes.

The new finding is four omitted Table III rows whose native act-section labels are blank. The audits also reproduce the false private-law not_requested state and the compiled hearing's false unique date. Current report sections, expanded bill/subject/print populations and several external refreshes remain partial. Empty model/backfill outputs are explicitly uncomputed.

The ledger advances only the scopes that passed. Full independent regulatory transformation replays pass against the exact catch-up parents; wider parent source qualification remains separate. The newer SAM scheduled acquisition succeeded, and former hearing detail refusals now have valid retained responses. At the 17:36 UTC end check, CRS alone had advanced beyond the audit freeze; that new generation remains unaudited. Base ETags and the rulemaking pointer were unchanged. This work changed documentation and retained local audit evidence; it made no production repair, publication or schedule change.

The final complete local run used the shared 3 GB DuckDB budget and one thread. It completed in 527 seconds at 9.62 GB peak process RSS, retaining the exact 26,303,691-row population and schema. Every index group and prior ID passed; all-column fingerprints match the retained public parent. A read-only check of all retained public/storage versions took 5.6 seconds, and a sample byte readback matched. That probe did not publish a receipt. See qualification.json, output-3GB/comments-build.json, layout-verification.json and transport-verification.json in the implementation evidence directory.

The 6 GB/two-thread diagnostics reached about 12.8 GB process RSS, even after connections were separated. Those results motivated the lower shared defaults; the configured budget is not a whole-process cap. The hosted ETL state was checked through the API and remains disabled_manually pending qualification.