Skip to content

court_docket_groups

Court docket same-case groups

Which published court_dockets rows are records of one case, with one representative row per group. CourtListener can hold one case as several docket objects and leaves its own parent link blank in its bulk editions, so the grouping is inferred from the records: rows under one court and docket number are one case when their captions read the same (ignoring case and punctuation, with U.S. and Dept. spelled out), or when they were filed the same day and one caption is contained in the other (v. Trump within City of New York v. Trump). Count each group once. confidence_tier says what the rows are: per_defendant_dockets, a criminal case's dockets for its several defendants, which PACER opens together with consecutive case ids; or duplicate_records, several CourtListener records of one filing (RECAP copies carrying 14–15-digit ids, an id-less record beside a RECAP one, an older docket object reused for the case). The representative row is the member whose PACER case id is the lowest of those in its court's sequence, meaning within 10% of the median id of the same court's dockets numbered within five of it. Where no such neighbour exists, as for appellate numbers, it is the lowest numeric id, which can be an older object reused for the case. It is a row to show for the group; the publisher marks no main case. Join on cl_docket_id; the representative's own row has parent_cl_docket_id equal to cl_docket_id. This is an inferred mapping, not publisher data, and edition-scoped.

Coverage. Not a range. A derived mapping over the published court_dockets selection, computed from CourtListener's bulk docket edition named in edition. The grouping evaluates every published row whose court and docket number hold two or more published rows and which that edition holds. Such a row has no group when the rows under its number show no sign of one case (different captions: a court reusing a number) or none has a numeric PACER case id; a row published after the edition has no group until the grouping is recomputed. (measured 2026-10-03)

  • Parquet file: court_docket_groups.parquet
  • MCP query_sql support: Configured; requires an available artifact.
  • Publication status: Not established by this schema page or its measurement date.
  • Row count: Not stated here; the MCP describe_table reply gives the live count under publication.
Column Type Description
cl_docket_id VARCHAR CourtListener docket id of a grouped row in court_dockets. Primary key.
parent_cl_docket_id VARCHAR The group's representative row: the member with the lowest PACER case id in its court's sequence, else the lowest numeric id (see the summary). Equal to cl_docket_id on the representative's own row. Not a main case the publisher marks; it marks none.
confidence_tier VARCHAR What the group's rows are: per_defendant_dockets (a criminal case's per-defendant dockets: the ids the representative was chosen from run consecutively with at most four missing, or a member carries CourtListener's defendant number) or duplicate_records (several records of one filing). Says nothing about refiling: a refiled case gets a new docket number, which this grouping cannot see.
group_size BIGINT Number of published court_dockets rows in the same-case group.
edition VARCHAR The CourtListener bulk edition the grouping was computed from (2026-06-30).
rule_version VARCHAR Version of the grouping rule: 2, the same-case test, the representative in its court's sequence, and the tier by what the rows are. Rule 1 (2026-09-22) grouped identical captions only, took the lowest PACER id as the parent and called a wide id spread refiled.
match_basis VARCHAR What showed the rows are one case: same_caption (every member's caption reads the same, ignoring case and punctuation, with U.S. and Dept. spelled out) or same_date_contained_caption (filed the same day, with one member's caption contained in another's).