Retained FEC corpus qualification — September 30, 2026
The selected retained non-PDF FEC data is qualified locally and queryable through the existing tables and MCP tools. The later research correction below closes omissions in the earlier selection; the earlier claim that it covered all retained non-PDF research was too broad. The current manifest, counts, source package and generation are recorded in the research-closure receipt. Independent checks preserve prior output cells and relationships and verify the added source fields and reference contexts. PDF corpus processing remains deferred by the user; the combined generation contains no PDF parsing or ingestion scopes. One historical API response body remains unavailable.
The selection includes official FEC source families, retained agency reports about FEC and their exact archive members. Unacquired official history and adjacent IRS, state, local and advertising research remain separate. The PDF deferral followed earlier native-text and OCR checkpoints; their artifacts and original pins remain retained. Filtered test suites nevertheless exercised PDF fixtures after that instruction, as recorded below.
The durable evidence root is
fec-corpus-completion-20260930/.
Its input manifests, qualification receipts and final verification record own
changing counts and exact identities. A discovered URL is not a retained file;
a byte-verified capture is not automatically a parsed or received population.
Research correction after the initial qualification
The older research audit found official FEC captures under academic/ and
repositories/, directories the earlier inventory had excluded wholesale.
The correction examines every research JSON/JSONL metadata file across all
directories and selects captures by source identity. It closes the originally
identified 99 missing non-PDF files, adds the loose bulk-download page found
outside capture folders, and accounts for retained file ranges and an archived
FEC page. The exact denominator and dispositions live in
inventory.json
and the adjacent capture-ledger.json. Explicitly FEC-related third-party
references retain their separate authority; adjacent political-data research
remains inventoried outside this FEC selection. A directory name does not decide
whether its contents belong to the official FEC corpus.
Retained API responses become native source observations through the existing document profile. Guidance, RSS, dictionaries and other reference bodies become literal, digest-pinned collection context. Byte ranges and legacy text without a qualified encoding retain reversible bytes and their source limitations. These references support source interpretation and coverage queries; they do not create financial transactions, inferred relationships or complete API histories. The added API observations retain every raw top-level response field.
The source fix preserves integer identifiers outside the artifact JSON safe
range as exact strings. Original response bytes retain their original numeric
spelling and type. The installed wheel is now
0.53.0+fec.bcdde5431fac, from source commit
bcdde5431fac89799f0f3f76400dfacbc661399f; see wheel.json and
wheel-change-scope.json beside the correction receipt. The shared FEC number
helper is its only changed Python module relative to the earlier wheel.
The ordinary receiving builder produces the additions. The combined build
compares every output batch with the earlier verified tables plus those
additions, preserving every prior cell and occurrence. Its first sealed result
is retained under pre-loose-original/; the final result adds the loose page's
contexts and reuses the verified native table bytes. Independent checks compare
original API fields, source hashes, reference text/bytes and every selected
context with receiving output. Final generation and local MCP checks are in
verification.json
and adjacent mcp-check.json. Source and consumer focused test logs are retained
there too. These explicitly selected checks perform no PDF processing.
The missing anonymous committee-API response remains
unresolved-original-body. The exact-size/digest search is retained in
missing-body-search.json;
a new live response cannot substitute for the historical bytes. This correction
accounts for that exception without claiming recovery. PDF originals and the
format-guide PDF member remain deferred without opening their bodies. There
was no new acquisition, push, main merge or remote publication in this correction.
Preserved selection and native additions
The September 21 original selection is preserved inside the later gap234
selection. The independent declaration and literal-cell baselines are retained
under independent/baseline-membership.json and independent/baseline-cells.json.
The combined manifest preserves prior declarations, source records and
relationship observations. Exact duplicate input declarations resolve to an
existing collection; differing source bytes, observation versions and archive
members remain separate. Receipt aliases and associations remain in coverage
metadata.
The native additions use the existing source-observation tables:
- Retained bulk master snapshots and every individual-contribution member, including the main stream and its date partitions. Independent row-hash multiplicities establish their overlap; physical observations remain separate.
- The assessed PostgreSQL committee-history table. Maintained
pg_restoreregenerates the pinned schema and COPY stream. Native column names, SQL NULL, empty strings, array text, raw escapes and source coordinates survive. - Retained filing originals and daily/paper archive members. Native bodies keep
their exact source byte ranges. These references are recoverable from pinned
originals; they do not make every filing body inline or searchable in SQL.
Malformed Senate
48.fecretains its earlier physical-line fallback and the native-reader refusal. - Earlier successful agency releases, FOIA.gov archive members and the Word Flat OPC original. Word package, XML and text/control observations do not establish rendered-page or OCR fidelity.
- Individual API response fields, records, assets and body text. Each declares
query_completeness: not-asserted; incomplete responses do not become complete query releases.
Reference dictionaries, workbook cells, listing facts, authority and exceptions
use digest-pinned caller context in fec_collections. They remain separate from
provider outcomes and from transaction observations. Credential-bearing strings
in caller context have explicit redaction locations and original-string digests;
retained originals remain unchanged. Metadata-only rows emit no source records and retain an unknown source-record count. A retained_unparsed
disposition is open work; it is never a completion substitute.
Evidence and limits
The standard rollup retains its exact input manifest and seals output membership,
schemas and the implementation digest. Source originals stay at their retained
locations and are verified before reading; the build does not acquire data.
Original acquisition dates remain distinct from later local-verification times;
where only an acquisition day is known, context retains that limitation.
The source wheel and consumer code identities are pinned in the run receipt.
The initial generation used an isolated wheel build from source commit
14b98e9db91644728b0fa7f924f22095dc719bf7, version
0.53.0+fec.14b98e9db916; its two builds are byte-identical. The wheel digest
and byte-for-byte comparison against committed package files are retained under
source/wheel.json and independent/wheel-source-comparison.json.
A session-scoped verified archive inventory avoids repeated decompression while
still checking each opened original's digest and size.
The coverage ledger distinguishes original captures, selected archive members, source reference metadata, derived bodies/OCR evidence and runtime files. Initial short inventory deadlines included recursive metadata processing; they were not proof of unreadable files. A staged retry resolved those deadlines before reconciling source-byte pins. The one remaining anonymous API access receipt has no located original body; it remains receipt-only and unresolved. The two synthetic committee-baseline fixtures and unrelated EXIM research are excluded. Publisher-origin fixture bytes that match selected archive members remain associated with those members without duplicate native rows.
Official family and bulk-group censuses are distinct axes. Neither establishes complete official history. Positional-only financial data retains unresolved publisher field definitions. Literal historical arrays do not produce current committee relationships. Financial amendment, deletion and summation policies remain consumer choices.
No remote publication, hosted MCP deployment, new acquisition, external message or primary DocSpec/OCR mutation is part of this campaign. The earlier agency-only generation is not a replacement for the larger public observation family. Publication of a combined candidate requires the independently verified preservation results and the ordinary conditional publication guard. The current remote population was not rechecked for this retained-corpus campaign.
Initial generation verification record
This section preserves the initial selection's evidence. The research correction above owns the current selection and package; these earlier receipts and output files remain unchanged.
The frozen input is consumer/combined-inputs.json, SHA-256
f722e2060f9fbf026b0456e789355b36260c4a2310e4e6943fd4a9cb4be44900.
The ordinary build-fec-observations rollup completed with upload disabled,
exit code zero, in 5,358.07 seconds. It sealed generation
0e2ae1be40d4036245b071939a31754d90c1307b7cca6e5c3c5cde07c605a41d.
The Parquet footers contain 7,915 collection rows, 80,413,092 native source
observations and 951,374 relationship observations. These are measured physical
observations, not unique financial transactions. Independent content comparisons
and local MCP checks passed; their receipts are recorded below.
The external run wrapper records process state, output growth and free disk
space. The final output tree, including staged tables, the sealed copy and retained
manifest evidence, occupies 19,108,964,346 bytes, within the 20 GiB output budget; 56,256,225,280 bytes remained free.
consumer/combined-run.json owns the build outcome; execution-summary.json
separates final files from sampled growth and the temporary original spool. Its initial dirty-diff pin
and executed implementation digest remain unchanged when later documentation or
test edits are committed. consumer/final-verification.json maps the final
consumer commit to those unchanged reviewed product-file hashes and the sealed
generation; a later commit identity does not replace the executed code identity.
| Boundary | Evidence under the campaign root |
|---|---|
| Retained denominator and unresolved originals | consumer/selection-closure.json, coverage-ledger.jsonl, coverage-census.json |
| Frozen selection, unique IDs and context hashes | independent/final-input-audit.json |
| Source-native cells, bytes and overlap | bulk/qualification.json, documents/completion.json |
| Complete bulk/SQL receiving fields and population hashes | bulk/receiving-native.json |
| Final native/relationship coverage reconciliation | independent/final-corpus-coverage.json |
| Every native collection assigned an output comparison | independent/receiving-coverage-plan.json |
| Prior main and agency output cells, contexts and final counts | independent/receiving-final.json |
| Executed build and resource limits | consumer/combined-run.json, combined-build.log, combined-resources.jsonl |
| Reviewed consumer functions and product hashes | bulk/consumer-review.md, consumer-review-checks.json, consumer-review-closure.json |
| Verified generations and actual local MCP tools | consumer/sealed-generations.json, combined-mcp.json, native-filing-body-recovery.json |
| New master relationship fields and multiplicities | bulk/new-bulk-relationships.json |
| New filing/API/agency cells and filing relationships | documents/receiving-comparison.json, filing-relationship-comparison.json |
| All declared relationship counts, including zero scopes | independent/relationship-count-readback.json |
Independent receiving checks preserved every prior main-selection cell and
relationship occurrence, and every prior agency-only cell. All caller-context
facts and pins match their retained JSON; metadata-only dispositions keep NULL
provider fields and the expected zero emitted rows. The first context checker
mistakenly compared the documented {pin, facts} object directly with the facts
file. Its failed receipt remains retained; the corrected checker independently
verifies both components and passes. See independent/receiving-final.json and
context-oracle-correction.json.
The bulk/SQL comparison matched all 77,935,935 selected row occurrences across 127 receiving collections, including literal native fields, mapped values, source identity and coordinates. Exact population digests match the independent source qualification. This overlaps the preserved baseline and is not an additional 77,935,935 rows beyond the final table total.
The document/filing comparison matched every selected native field, body range and API asset, including all 107 new filing relationships across the 757 checked filing selections. The bulk relationship comparison matched all 750,338 new master relationship observations. Both checks compare every output field and exact occurrence counts, including missing/empty states and source identity.
The existing local MCP server successfully described all five FEC tables and
queried actual rows. Checks cover API raw response fields and assets, financial
workbook cells, PostgreSQL NULL/empty/native-array distinctions, candidate-cycle
coverage, explicit PDF deferral and the Senate fallback. The existing native
filing-body reader also recovered a pinned body from its exact source range.
The server accurately reports local_unversioned; the separate generation
verification binds the files' sealed artifact pins. No hosted-access or
publication claim follows from these local checks.
The initial consumer suite filtered with -k "not pdf" passed 3,677 tests, skipped two and deselected
57, but failed three checks: two coverage-description format assertions and
one stale source-evidence exemption. That nonzero run remains in
consumer/gate-non-pdf.json and pytest-non-pdf.log. After correcting those
issues and restricting the new asset handling to its intended profile, the
focused affected checks passed: 153 tests in post-review-focused.log. A further
single test in manifest-mutation-focused.log verifies that a changed manifest
prevents generation sealing while retaining the original bytes. The initial
suite is not relabeled as an all-green rerun. The filter also had a limitation:
an unchanged GovInfo hearing test without pdf in its name still exercised a
small synthetic PDF after the user's deferral. consumer/test-filter-miss.json
records the exact command, test witness and suite-log time window. That execution
is retained in the original log; the run must not be described as entirely free
of PDF tests. The identified witness did not process retained FEC PDF originals.
Ruff, type checking and generated-dictionary validation passed with the pinned
installed wheel. The source package's separate filtered gate and API-reader
qualification are recorded under source/. Static review found that the source
gate also selected Senate PDF extraction tests whose names lacked pdf,
including test_document_capture_xml.py::test_fresh_senate_adapter_capture_round_trips.
The API-focused source test
test_fec_documents.py::test_api_error_and_invalid_scope_refuse also generated a
synthetic PDF before refusing its invalid backend scope. The filtered/focused
exclusions therefore did not ensure PDF-free test runs. The static source audit is retained in
documents/source-pdf-deferral-audit.json. These misses occurred after the
deferral and remain explicit; the consumer synthetic-PDF witness was
not the only exception. Further PDF extraction tests and corpus
processing remain paused under the user's instruction. Pre-deferral results are
preserved separately and do not count as PDF adoption in the combined generation.