Skip to content

court_opinion_clusters

Court decisions and their opinion groups

Decision metadata from CourtListener opinion-clusters bulk data and opinion-search catch-up. cluster_id joins court_opinions.cluster_id (and through it the citation map and parentheticals); cl_docket_id joins court_dockets.cl_docket_id only where both tables hold the same publisher docket object: an appellate decision can name a scraper-created docket while court_dockets holds the RECAP one, so match on the docket number when the id does not join. Court identity comes from a bulk docket map or a search result; jurisdiction classification uses CourtListener's court reference data. This table does not hold individual opinion bodies. All columns are VARCHAR; native numeric and boolean spellings are strings.

Coverage. Not a range. The complete 2026-06-30 CourtListener opinion-clusters export, every mapped source field checked against the original export; the court fields agree with the same-edition docket map and court reference data. A search catch-up adds every cluster whose id is above the export's highest, whatever its filing date; a correction to an exported cluster arrives with the next export. Opinion bodies and broader historical coverage are separate. Per-court gaps are the publisher's: cadc has no unpublished cluster between 2001-04 and 2020-07. See docs/research/fork-output-ledger-2026-09-21.md and receipt join-gaps-2026-09-26/i/. (measured 2026-09-22)

Data quality. Bulk and search records expose different fields. ingest_source identifies the selected row's route; unavailable fields are NULL and metadata empty strings are normalized to NULL. A search row never replaces a bulk row of the same export, whose ids it lies above, and a later export's bulk row replaces it. Opinion search does not index RECAP trial-court clusters, so one created after the export arrives only with the next export. Missing court scope is unknown, not proof of a non-federal court. court_dockets is a narrower selection, so unmatched docket IDs are expected. Names, judge strings and citations do not establish cross-source identity. One decision can be several clusters. On the 2026-06-30 export (measured 2026-10-03), 131 of the 140 Supreme Court decisions reported in 600-606 U.S., grouped by cl_docket_id and U.S. citation, have two to five clusters, none with an scdb_id. That grouping is a heuristic: it cannot see a decision with no reporter citation yet.

  • Parquet file: court_opinion_clusters.parquet
  • MCP query_sql support: Configured; requires an available artifact.
  • Publication status: Not established by this schema page or its measurement date.
  • Row count: Not stated here; the MCP describe_table reply gives the live count under publication.
Column Type Description
cluster_id VARCHAR CourtListener opinion-cluster ID: this table's primary/dedup key, one per CourtListener record of a decision, not per decision. The publisher can hold several clusters for one decision, a Supreme Court slip opinion and its later U.S. Reports print above all, and later opinions cite either (Loper Bright: 9986254 and 10600041). Count decisions by scdb_id, the publisher's decision key, where it is set; it is unset on every Supreme Court cluster filed after 2019. Otherwise group by cl_docket_id and reporter citation (court_citations), a heuristic that misses a decision with no reporter citation yet. Join to court_opinions.cluster_id and court_citations.cluster_id; not an opinion_id.
cl_docket_id VARCHAR CourtListener docket ID, renamed from docket_id. Joins court_dockets.cl_docket_id only where both tables hold the same publisher docket object (an appellate case often has two: the court-website scraper's, named here, and RECAP's, held there), so match on the docket number when the id does not join; unrelated to regulations.gov docket_id.
court_id VARCHAR CourtListener court identifier from the docket map or search record; NULL when unresolved.
court_jurisdiction VARCHAR CourtListener jurisdiction code from the court reference lookup; NULL when unavailable.
court_is_federal VARCHAR t where CourtListener's jurisdiction code for the court starts with F (F appellate, FD district, FB bankruptcy, FBP bankruptcy appellate panel, FS special), f otherwise, as a string; NULL when the classification is unavailable. FS covers courts outside the circuit system, Article I courts and executive adjudicators among them (cit, uscfc, cavc, bia, mspb), so t does not mean an Article III court: filter court_jurisdiction for that.
case_name VARCHAR Source case caption; not a stable identity or affiliation link.
case_name_short VARCHAR Source short case caption, when supplied by bulk data.
case_name_full VARCHAR Source full case caption when supplied.
date_filed VARCHAR Source decision filing date; used by search catch-up, not an observation or correction date.
date_filed_is_approximate VARCHAR Source flag stating whether date_filed is approximate, stored as a string.
judges VARCHAR Source judge text; names are not normalized person identifiers.
nature_of_suit VARCHAR Source nature-of-suit text when supplied; this cluster table is not limited to the docket rollup's 899 query.
precedential_status VARCHAR CourtListener's source precedential-status value; this table does not independently assess legal authority.
citation_count VARCHAR The publisher's count of citing opinions at observation time, as a string: opinions, not decisions, and only decisions carrying a reporter citation are ever cited.
scdb_id VARCHAR Supreme Court Database identifier as supplied by CourtListener: the publisher's decision key for a Supreme Court decision where it is set. NULL on every Supreme Court cluster filed after 2019 (2026-06-30 export), whose decisions only the heuristic cluster_id names can group.
scdb_decision_direction VARCHAR Source Supreme Court Database decision-direction code; retained without interpretation.
scdb_votes_majority VARCHAR Source Supreme Court Database majority-vote count, stored as a string.
scdb_votes_minority VARCHAR Source Supreme Court Database minority-vote count, stored as a string.
source VARCHAR CourtListener's source value for the cluster; distinct from the ingest_source route.
procedural_history VARCHAR Source procedural-history text when supplied.
attorneys VARCHAR Source attorney text; does not retain structured person identity.
posture VARCHAR Source procedural-posture text when supplied.
syllabus VARCHAR Source syllabus text when supplied.
headnotes VARCHAR Source headnotes text when supplied.
summary VARCHAR Source summary text when supplied; not a generated SpicyRegs summary.
disposition VARCHAR Source disposition text when supplied.
history VARCHAR Source history text when supplied.
other_dates VARCHAR Source other-dates text, without date normalization.
cross_reference VARCHAR Source cross-reference text, without inferred identifier resolution.
correction VARCHAR Source correction text when supplied; not proof that later corrections have been acquired.
arguments VARCHAR Source arguments text when supplied.
headmatter VARCHAR Source headmatter text when supplied.
blocked VARCHAR Source blocked flag as a string; not a current access check.
date_blocked VARCHAR Source date of blocking when supplied.
slug VARCHAR Source URL slug; not a stable identity.
absolute_url VARCHAR CourtListener decision page URL, constructed from cluster ID/slug for bulk rows or supplied by search.
date_created VARCHAR Source record-creation timestamp when supplied by bulk data.
date_modified VARCHAR Source record-modification timestamp when supplied by bulk data.
ingest_source VARCHAR Route for the retained row: bulk or search. The table does not itself pin the original capture bytes.