article · 1 min read · Pandexis Research · Mar 14, 2026

AI consensus in document review: why many agents beat one

A single LLM's confidence score is not evidence. Consensus across many independent agents - measurable, auditable, defensible - is.

The individual-model problem

A single large language model producing a confidence score for a document classification is producing an opinion. An opinion, however calibrated, is not a reviewable record. Opposing counsel can challenge it; a judge can ask what exactly produced the number; a partner cannot defend what she cannot audit.

Consensus as evidence

Run the same document past N independent agents. Measure agreement. Where they agree - with a score that exists because we took a vote, not because one model asserted it - you have something closer to evidence. Where they disagree, you have a signal for a human reviewer to spot-check. The disagreement rate itself is a metric you can report.

What Pandexis surfaces

Per-document agreement rate, per-agent reliability score, the specific passages each agent cited, and the consensus math (Dawid-Skene iterations, convergence) - all visible in the product and exportable into the review record.

The call to action

The next time you are asked 'how did you decide this document was responsive', the honest answer should be longer than 'the model was confident'. Pandexis gives you the longer answer.