ForensicBench
Evaluating agentic LLMs on journal-entry fraud detection
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
LORIA, CNRS, Universite de Lorraine · LIPN, CNRS · Forvis Mazars
| Date | Doc | Account | Debit | Credit |
|---|---|---|---|---|
| 2024-01-12 | JE-80412 | 606300 Purchases | 18,250.00 | |
| 2024-01-12 | JE-80412 | 401000 Suppliers | 18,250.00 | |
| 2024-01-14 | JE-80977 | 606300 Purchases | 42,000.00 | |
| 2024-01-14 | JE-80977 | 401000 Suppliers | 42,000.00 | |
| 2024-01-31 | JE-81530 | 641000 Payroll | 96,400.00 | |
| 2024-01-31 | JE-81530 | 421000 Staff pay | 96,400.00 | |
| 2024-02-03 | JE-82291 | 401000 Suppliers | 42,000.00 | |
| 2024-02-03 | JE-82291 | 512000 Bank | 42,000.00 |
Red rows belong to one fraud scheme. The ledger gives no such marker: the agent has to find them.
The task
Find the scheme, not just the entry
An LLM agent gets read-only SQL access to a company ledger. It must flag the fraudulent journal entries and assign each one a scheme type. Nothing in the data says which entries are fraudulent: the agent has to build the evidence itself.
| 606300 | Purchases | Dr | 42,000 |
| 401000 | Supplier | Cr | 42,000 |
| 401000 | Supplier | Dr | 42,000 |
| 512000 | Bank | Cr | 42,000 |
Both entries balance and look routine on their own. Only their joint structure (same supplier, same amount, weeks apart) reveals the scheme.
Why it matters
Fraud is a pattern, audits need patterns
- Real accounting fraud rarely shows up as one suspicious posting. It is a coordinated sequence of entries, each individually plausible.
- Auditors search for scheme-level evidence and name the process behind it, such as fictitious vendors or ghost-employee payroll.
- Agentic LLMs can query databases, run code and keep investigative state, so they could assist this work. Whether they can is an open question.
- Real ledgers are confidential. A synthetic, label-rich ledger makes the question testable and reproducible.
The problem
No benchmark covers this setting
Financial-statement datasets
Entry-level detectors
General agent benchmarks
No existing work combines scheme-level granularity with a live relational ledger, which is what an auditor works with.
Contributions
What we release
Forensic Ledger
Realistic synthetic ledgers across five sectors, with five injected multi-entry fraud schemes and labels at both entry and scheme level.
Evaluation protocol
Entry-F1, Type-F1, Coverage and Consistency, over 25 runs per model, with no composite score.
Reference agent
A plan-then-investigate agent that queries the ledger with SQL and Python, reproducible and released as a baseline.
Open leaderboard
A public ranking of open-weight models, with a hidden label store and an automatic scorer for new submissions.
Method
Ledger, agent and protocol
The Forensic Ledger
We build on DataSynth, an open-source enterprise data generator that produces balanced entries, Benford-compliant amounts and realistic calendars under the French chart of accounts (PCG). Its labels mark single entries and it ships only two multi-stage schemes, so we add a scheme-injection layer. Schemes run as multi-stage processes with forensic traces, and every injected entry gets a label that links its document to a scheme type and a scheme instance.
Five sectors (Energy, Healthcare, Luxury Goods, Manufacturing, Transport), about 300K entries each over three years, and five scheme types:
The reference agent
One common harness runs every model, so scores are comparable. It is deliberately simple, with no replanning and no reflection, which means the leaderboard measures a model together with its scaffold.
Orient
The agent profiles the ledger with SQL before any fraud reasoning, using about 10% of the budget.
Plan, once
A single call produces ranked, falsifiable hypotheses with exit criteria and a token budget each. The plan is then fixed.
Investigate each hypothesis
A worker runs a bounded loop of reasoning, tool calls (SQL, Python) and observation until its exit criteria are met or its budget is spent, then returns a verdict.
Report
Suspicious entries are flagged with a scheme type through report_suspicion. Each run has a fixed 20M-token budget.
Evaluation protocol
Entry-F1 (primary)
Type-F1
Coverage
Consistency
Each model runs on 5 sectors with 5 replicates each, so 25 runs per model and 300 in total. Models are ranked lexicographically, with no composite score.
Results
Even the best model recovers about a third
Entry-F1 (%) by model
Macro-average over 25 runs per model. Type-F1 for the best model is 21.7, and it trails Entry-F1 on every model, by up to 13 points.
A low ceiling
Detecting is not typing
Some schemes are hard
The task is solvable
Leaderboard
Submit your agent
The unlabelled ledgers are released on Hugging Face and scoring is automatic. Run your own agent on the ledgers, upload your flags (document id, scheme type) as a single run or as several replicates, and get your Entry-F1, Type-F1 and Coverage. Labels stay private, and each entry comes with the code of its harness for verification.
Citation
@inproceedings{waffo2026forensicbench,
title={ForensicBench: Evaluating Agentic LLMs on Journal-Entry Fraud Detection},
author={Waffo Dzuyo, Guy Stephane and Guibon, Ga{\"{e}}l and Cerisara, Christophe and Belmar-Letelier, Luis},
booktitle={EMNLP 2026, Industry Track},
year={2026}
}