Abstract
This paper describes the multi-agent consensus architecture Pandexis uses for document review. Agents run in parallel across multiple LLM providers, emit structured extractions, and are reconciled through Dawid-Skene consensus. We report on agreement-rate measurement, per-agent reliability estimation, and the effect of cross-provider redundancy on production defensibility.
1. Motivation
The per-model confidence score is not a reviewable artifact. Consensus produces a measurable, auditable agreement statistic that maps onto the defensibility standard the review workflow actually needs.
2. Method
N independent agents are run per document chunk, with configurable redundancy_level. Each emits a schema-validated extraction. A Dawid-Skene aggregator iterates per-agent reliability and per-document class probability until convergence. Disagreement is surfaced to the reviewer.
3. Results
Preliminary internal benchmarks show substantial defensibility improvement over single-model baselines at comparable review cost. Full benchmarks and dataset descriptions available on request under a customer NDA.
4. Discussion
Cross-provider redundancy matters. Running a red-team agent from a different model family produces a meaningfully different set of false positives and negatives than running a second agent from the same family at higher temperature.