ExCoR
ExCoR · Exact Coalitions for Rankers

Why is this document ranked above that one?

ExCoR explains the order of a dense retriever's whole ranked list. It groups the passages of the ranked documents into named topics and measures how much each topic holds the order together. Every masking is scored exactly from cached vectors, without re-encoding.

Scroll to follow one real query, step by step ↓
How it works

From a ranked list to its explanation, in seven steps

We follow one question from the FinQA benchmark through the whole pipeline. Every chunk, group name, score and Shapley value shown below is the actual output of ExCoR for this query.

Query

  1. Input

    A query and its ranked list

    The ranker returns the best documents for the query; the first are shown. ExCoR explains this order: the whole list, not one document at a time.

    1
  2. Chunk & encode

    Chunks, encoded once

    Each document is cut into non-overlapping chunks of tokens, each encoded independently and cached. With MaxP, a document's score is the similarity of its best chunk (outlined) to the query.

    Hover a chunk to read it.

    2
  3. Group & name

    Topics shared across documents

    k-means clusters the chunks of the documents into k = groups, local to this query and shared across documents. An LLM names each group from its most central chunks.

    In colour: the groups that turn out most important. In grey: the other groups.

    3
  4. Coalition

    Remove groups, rescore exactly

    A coalition is the set of groups we keep. Here we drop and rescore each document from its remaining chunks. Since chunks are encoded independently, deleting text is the same as dropping cached vectors: the new scores are exact, with no re-encoding, and the list reorders.

    4
  5. Value

    How much of the order survives?

    The value v(S) of a coalition is the NDCG of the new order against the original one: 1 when the order is intact. For the coalition of step 4, where the list reorders, v(S) = .

    The scale does not start at 0. With no group kept, every document is empty and their order is arbitrary: v(∅) = is the average NDCG of a random order of the documents, since NDCG penalises lower ranks only logarithmically.

    5
  6. Attribute

    100,000 coalitions, one Shapley value per group

    Each coalition costs a few vector operations, so ExCoR scores 100,000 random coalitions per query. KernelSHAP turns them into a Shapley value φ per group: its average contribution to keeping the order.

    Our experiments show that this budget is needed for stable attributions: with 100,000 coalitions, the group ranking agrees with a 1,000,000-coalition reference at Kendall τ = 0.93 to 0.95. Explainers that re-encode the documents already need minutes per query for 5,000 coalitions, where τ is only 0.72 to 0.79.

    6
  7. Explain

    Read the explanation

    The groups with the largest φ explain the order of the list. A pairwise question, , is read from the same attributions: compare the groups present in one document and absent from the other.

    A reading of the listwise attributions, not a separate pairwise method.

    7
Results

Why ExCoR

Five findings from the paper, on four long-document benchmarks (legal, financial, narrative) and ten embedders. Each one comes with its evidence.

+0.152
Fidelityb over the best ChunkGroupSHAP result on AILA (+0.052 on FinQA, +0.022 on FinanceBench).
E5-small, mean aggregation · Table 3
0.3s / query
against 168–411 s for ChunkGroupSHAP on the same GPU: about three orders of magnitude faster.
One H100 · Table 3
0.93–0.95
Kendall τ with a 1,000,000-coalition reference at 100,000 coalitions, against 0.72–0.79 at 5,000.
10 encoders · Figure 2
78–91%
of groups pass the intruder test, against 18–23% for random groups (chance: 20%).
10 encoders · Figure 3
Faithfulness

The explanation recovers the ranking

Fidelityb asks whether an explanation suffices to rebuild the ranking: keep the seven groups with the highest Shapley values, score each document by the groups it contains, and compare that order with the real one (Kendall τb).

With mean aggregation, ExCoR is more faithful than ChunkGroupSHAP on all three benchmarks they share, by +0.152 on AILA, +0.052 on FinQA and +0.022 on FinanceBench over its best result, reported or re-run on the same queries.

E5-small ranker · Table 3 of the paper

Show the full table
MethodAILAFinQAFinanceBenchTime (s)
RankSHAP (words, reported)0.0370.2090.221–
ChunkGroupSHAP, k = 200 (reported)0.1920.3130.346–
ChunkGroupSHAP, k = 500 (reported)0.2320.3220.421–
ChunkGroupSHAP, k = 200 (re-run, variant)0.1850.3030.338171–411
ChunkGroupSHAP, k = 500 (re-run, variant)0.1400.3230.403168–395
ExCoR (mean)0.3840.3750.4430.3
ExCoR (MaxP)0.2570.3640.4100.3

Reported rows come from the ChunkGroupSHAP paper; their times, measured on other hardware, are not shown. ExCoR: mean over three explanation seeds.

Cost

Seconds become fractions of a second

Re-encoding explainers run the encoder for every coalition: our reproduction of ChunkGroupSHAP needs 157,000 to 196,000 encoder passes per query. ExCoR needs none: a coalition only drops cached chunk vectors.

On the same H100, ExCoR explains a top-50 list in 0.3 s with 100,000 coalitions, against 168 to 411 s for ChunkGroupSHAP with 5,000: about three orders of magnitude faster.

E5-small ranker · Table 3 and Section 3.4 of the paper

Stability

Enough coalitions for stable attributions

Shapley values are estimated from sampled coalitions, and the estimates converge slowly. Against a reference from 1,000,000 coalitions, the order of the groups agrees at τ = 0.93 to 0.95 with 100,000 coalitions, for every encoder and both aggregations, and more than 95% of the reference top-10 groups are recovered.

At the 5,000 coalitions that re-encoding explainers can afford, τ is only 0.72 to 0.79.

10 encoders, 4 benchmarks · Figure 2 of the paper

Kendall tau with the 1,000,000-coalition reference as a function of the number of coalitions N, for mean and MaxP aggregation: the curves rise and level off near 0.95 at N = 100,000.
Kendall τ with the reference as a function of the number of coalitions N; line: median over the ten encoders, band: min–max. Dashed: 5,000 (ChunkGroupSHAP) and 100,000 (ExCoR).
Coherence

Groups are real, nameable topics

An explanation is only useful if its groups mean something. In an intruder test, an LLM sees four chunks of a group and one from another group, and must find the intruder.

It succeeds for 78 to 91% of ExCoR's groups, against 18 to 23% for random groups of the same sizes (chance: 20%), for every encoder and dataset. The groups are also listwise: on AILA and NarrativeQA, a group spans 6.6 and 16 documents on average.

MaxP, k = 100, 10 encoders · Figure 3 of the paper

Intruder-test accuracy per encoder and dataset: k-means groups around 0.8 to 0.9, random groups around 0.2, the chance level.
Intruder-test accuracy of ExCoR's groups (blue) and of random groups of the same sizes (grey), per encoder and dataset. Dashed: chance level.
Retrieval quality

Exact explanations, at no cost in ranking quality

ExCoR explains embedders used chunk-wise. Used this way, they match or outperform the same embedders used flat on three of the four benchmarks.

With MaxP, chunk-wise encoding is the best choice for all ten encoders on FinQA (+0.187 nDCG@10 on average) and NarrativeQA (+0.400), where truncation discards most of each book. AILA is close; only FinanceBench, whose documents fit in one window, favours flat encoding.

nDCG@10, mean over the ten encoders · Table 2 of the paper

See explanations for real queries

4 long-document benchmarks, 10 encoders, selected example queries with their named groups, documents and pairwise readings.

Explore examples