The explanation recovers the ranking
Fidelityb asks whether an explanation suffices to rebuild the ranking: keep the seven groups with the highest Shapley values, score each document by the groups it contains, and compare that order with the real one (Kendall τb).
With mean aggregation, ExCoR is more faithful than ChunkGroupSHAP on all three benchmarks they share, by +0.152 on AILA, +0.052 on FinQA and +0.022 on FinanceBench over its best result, reported or re-run on the same queries.
E5-small ranker · Table 3 of the paper
Show the full table
| Method | AILA | FinQA | FinanceBench | Time (s) |
|---|---|---|---|---|
| RankSHAP (words, reported) | 0.037 | 0.209 | 0.221 | – |
| ChunkGroupSHAP, k = 200 (reported) | 0.192 | 0.313 | 0.346 | – |
| ChunkGroupSHAP, k = 500 (reported) | 0.232 | 0.322 | 0.421 | – |
| ChunkGroupSHAP, k = 200 (re-run, variant) | 0.185 | 0.303 | 0.338 | 171–411 |
| ChunkGroupSHAP, k = 500 (re-run, variant) | 0.140 | 0.323 | 0.403 | 168–395 |
| ExCoR (mean) | 0.384 | 0.375 | 0.443 | 0.3 |
| ExCoR (MaxP) | 0.257 | 0.364 | 0.410 | 0.3 |
Reported rows come from the ChunkGroupSHAP paper; their times, measured on other hardware, are not shown. ExCoR: mean over three explanation seeds.