ForensicBench
Evaluating agentic LLMs on journal-entry fraud detection
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
LORIA, CNRS, Universite de Lorraine · LIPN, CNRS · Forvis Mazars
| Date | Doc | Account | Debit | Credit |
|---|---|---|---|---|
| 2024-01-12 | JE-80412 | 606300 Purchases | 18,250.00 | |
| 2024-01-12 | JE-80412 | 401000 Suppliers | 18,250.00 | |
| 2024-01-14 | JE-80977 | 606300 Purchases | 42,000.00 | |
| 2024-01-14 | JE-80977 | 401000 Suppliers | 42,000.00 | |
| 2024-01-31 | JE-81530 | 641000 Payroll | 96,400.00 | |
| 2024-01-31 | JE-81530 | 421000 Staff pay | 96,400.00 | |
| 2024-02-03 | JE-82291 | 401000 Suppliers | 42,000.00 | |
| 2024-02-03 | JE-82291 | 512000 Bank | 42,000.00 |
Red rows belong to one fraud scheme. The ledger gives no such marker: the agent has to find them.
One common harness runs every model of the paper, so scores are comparable. It is deliberately simple, with no replanning and no reflection: the leaderboard measures a model together with its scaffold. Use it as a baseline, or plug in your own harness and submit the flags.
Four phases, one budget
The agent orients itself on the ledger, plans once, investigates every hypothesis with a dedicated worker, then reports. Each run has a fixed 20M-token budget.
Inside an investigation worker
A worker receives one hypothesis with its exit criteria and budget. It alternates reasoning, tool calls (SQL, Python) and observation until the criteria are met or the budget is spent, then returns a verdict.
Fraud catalogue and prompts
The harness gives the agent a conceptual catalogue of the five scheme types: the normal business process and the kinds of breakdown that can indicate manipulation. It names no GL account and no injection parameter, and contains no label. The catalogue and the prompts of the reference harness are downloadable with the data.
Where the harness fits
The harness only sees the read-only ledger. Its flags are scored against labels that never leave the private store.