ForensicBench
Evaluating agentic LLMs on journal-entry fraud detection
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
LORIA, CNRS, Universite de Lorraine · LIPN, CNRS · Forvis Mazars
| Date | Doc | Account | Debit | Credit |
|---|---|---|---|---|
| 2024-01-12 | JE-80412 | 606300 Purchases | 18,250.00 | |
| 2024-01-12 | JE-80412 | 401000 Suppliers | 18,250.00 | |
| 2024-01-14 | JE-80977 | 606300 Purchases | 42,000.00 | |
| 2024-01-14 | JE-80977 | 401000 Suppliers | 42,000.00 | |
| 2024-01-31 | JE-81530 | 641000 Payroll | 96,400.00 | |
| 2024-01-31 | JE-81530 | 421000 Staff pay | 96,400.00 | |
| 2024-02-03 | JE-82291 | 401000 Suppliers | 42,000.00 | |
| 2024-02-03 | JE-82291 | 512000 Bank | 42,000.00 |
Red rows belong to one fraud scheme. The ledger gives no such marker: the agent has to find them.
Leaderboard
Entries are ranked by Entry-F1, then Type-F1, Recall, Precision, Coverage and Consistency. There is no composite score. Rows using the reference harness are the paper results.
| # | Model | Harness | Runs | Entry-F1 | Type-F1 | Recall | Prec. | Cover. | Consist. |
|---|---|---|---|---|---|---|---|---|---|
| 1 | MiniMax-M2.7230B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 34.7 | 21.7 | 29.0 | 48.4 | 21.7 | 50.7 |
| 2 | Qwen3.5-397B397B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 26.3 | 14.1 | 20.5 | 41.3 | 18.0 | 61.4 |
| 3 | Qwen3.5-122B122B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 25.2 | 13.5 | 19.8 | 40.6 | 14.2 | 46.9 |
| 4 | Qwen3.6-35B35B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 15.0 | 9.4 | 9.8 | 36.6 | 9.3 | 28.1 |
| 5 | Mistral-Medium-3.5-128B128B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 13.6 | 9.1 | 10.3 | 21.8 | 9.5 | 43.8 |
| 6 | Gemma-4-31B31B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 12.1 | 8.7 | 7.5 | 37.4 | 6.8 | 76.1 |
| 7 | Gemma-4-E4B4B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 8.5 | 3.8 | 5.2 | 29.5 | 4.2 | 31.2 |
| 8 | GPT-OSS-120B120B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 8.0 | 5.2 | 4.6 | 39.9 | 4.4 | 55.8 |
| 9 | Mistral-Small-4-119B119B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 3.8 | 2.4 | 2.3 | 13.4 | 2.1 | 28.5 |
| 10 | Qwen3.5-9B9B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 3.5 | 2.2 | 2.0 | 18.5 | 2.0 | 18.0 |
| 11 | Granite-30B30B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 2.8 | 1.5 | 1.5 | 16.9 | 1.6 | 38.4 |
| 12 | Llama-3.3-70B70B Open-weightVerified | Reference harness ForensicBench authorscode | 25 runs | 0.3 | 0.0 | 0.1 | 3.2 | 0.2 | 80.4 |
Submission formats
Run your agent on the five unlabelled ledgers and upload one CSV of flags, one row per flagged journal entry. A file must cover all five sectors.
Single run
One run per sector. Scores are reported without a standard deviation or Consistency, and the entry is marked as 1 run.
sector,document_id,scheme_type energy,9f2c...,fictitious_ap_disbursements
Multi run
Several replicates per sector, the same ones in every sector (the full protocol uses 5, so 25 runs). Reports the mean, the standard deviation across replicates and Consistency.
sector,replicate,document_id,scheme_type energy,1,9f2c...,fictitious_ap_disbursements
- Scheme types: fictitious_ap_disbursements, revenue_manipulation, vendor_collusion, shadow_payroll, inventory_manipulation. Use unknown when the type is not decided (the entry still counts for Entry-F1).
- Sectors: energy, healthcare, luxurygoods, manufacturing, transport.
- Rows with the same sector, replicate and document are counted once.
Labels
The ground-truth labels are held out so that the ranking stays meaningful, and scoring is automatic. If you need them for research evaluation or reproduction, write to guywaffo@gmail.com.
Harness code and verification
Every submission must name its harness and include its code: a repository URL with the commit hash, or an archive. Entries start as unverified. A maintainer reviews the harness before an entry is marked verified.
- Automatic checks: valid identifiers, duplicate rows, share of the ledger flagged per run, identical replicates, implausibly high precision and recall. A flagged file goes to review.
- The harness must only use the read-only ledger. Labels, ledger-specific lookups and hand-labelling are not allowed.
- Rate limit: 3 submissions per email per day.