EMNLP 2026 Industry TrackPublished

ForensicBench

Evaluating agentic LLMs on journal-entry fraud detection

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
LORIA, CNRS, Universite de Lorraine · LIPN, CNRS · Forvis Mazars

General ledgerwhat the agent queries
DateDocAccountDebitCredit
2024-01-12JE-80412606300 Purchases18,250.00
2024-01-12JE-80412401000 Suppliers18,250.00
2024-01-14JE-80977606300 Purchases42,000.00
2024-01-14JE-80977401000 Suppliers42,000.00
2024-01-31JE-81530641000 Payroll96,400.00
2024-01-31JE-81530421000 Staff pay96,400.00
2024-02-03JE-82291401000 Suppliers42,000.00
2024-02-03JE-82291512000 Bank42,000.00

Red rows belong to one fraud scheme. The ledger gives no such marker: the agent has to find them.

Benchmark architecture
ForensicBench architecture: ledger, agent harness, flags, scorer with private labels, leaderboard
1.51M
journal entries
5
sector ledgers
5
fraud scheme types
12
open-weight models
34.7%
best Entry-F1

The task

Find the scheme, not just the entry

An LLM agent gets read-only SQL access to a company ledger. It must flag the fraudulent journal entries and assign each one a scheme type. Nothing in the data says which entries are fraudulent: the agent has to build the evidence itself.

Jan 14 · Fictitious invoice
606300PurchasesDr42,000
401000SupplierCr42,000
Feb 03 · Matched disbursement
401000SupplierDr42,000
512000BankCr42,000

Both entries balance and look routine on their own. Only their joint structure (same supplier, same amount, weeks apart) reveals the scheme.

Why it matters

Fraud is a pattern, audits need patterns

  • Real accounting fraud rarely shows up as one suspicious posting. It is a coordinated sequence of entries, each individually plausible.
  • Auditors search for scheme-level evidence and name the process behind it, such as fictitious vendors or ghost-employee payroll.
  • Agentic LLMs can query databases, run code and keep investigative state, so they could assist this work. Whether they can is an open question.
  • Real ledgers are confidential. A synthetic, label-rich ledger makes the question testable and reproducible.

The problem

No benchmark covers this setting

Financial-statement datasets

Fraud is labelled at company-year level, far above the scheme level needed to investigate a ledger.

Entry-level detectors

They flag isolated statistical outliers, with no scheme type and no accounting-process context.

General agent benchmarks

They cover web and code tasks, and contain no forensic accounting setting.

No existing work combines scheme-level granularity with a live relational ledger, which is what an auditor works with.

Contributions

What we release

1

Forensic Ledger

Realistic synthetic ledgers across five sectors, with five injected multi-entry fraud schemes and labels at both entry and scheme level.

2

Evaluation protocol

Entry-F1, Type-F1, Coverage and Consistency, over 25 runs per model, with no composite score.

3

Reference agent

A plan-then-investigate agent that queries the ledger with SQL and Python, reproducible and released as a baseline.

4

Open leaderboard

A public ranking of open-weight models, with a hidden label store and an automatic scorer for new submissions.

Method

Ledger, agent and protocol

The Forensic Ledger

We build on DataSynth, an open-source enterprise data generator that produces balanced entries, Benford-compliant amounts and realistic calendars under the French chart of accounts (PCG). Its labels mark single entries and it ships only two multi-stage schemes, so we add a scheme-injection layer. Schemes run as multi-stage processes with forensic traces, and every injected entry gets a label that links its document to a scheme type and a scheme instance.

Five sectors (Energy, Healthcare, Luxury Goods, Manufacturing, Transport), about 300K entries each over three years, and five scheme types:

Fictitious AP disbursementsVendor collusionRevenue manipulationInventory manipulationShadow payroll

The reference agent

One common harness runs every model, so scores are comparable. It is deliberately simple, with no replanning and no reflection, which means the leaderboard measures a model together with its scaffold.

1

Orient

The agent profiles the ledger with SQL before any fraud reasoning, using about 10% of the budget.

2

Plan, once

A single call produces ranked, falsifiable hypotheses with exit criteria and a token budget each. The plan is then fixed.

3

Investigate each hypothesis

A worker runs a bounded loop of reasoning, tool calls (SQL, Python) and observation until its exit criteria are met or its budget is spent, then returns a verdict.

4

Report

Suspicious entries are flagged with a scheme type through report_suspicion. Each run has a fixed 20M-token budget.

Evaluation protocol

Entry-F1 (primary)

Did the agent flag the right entries?

Type-F1

Right entry and right scheme type. Always at most Entry-F1.

Coverage

Share of each fraud family recovered, averaged over families.

Consistency

Stability across the five replicates.

Each model runs on 5 sectors with 5 replicates each, so 25 runs per model and 300 in total. Models are ranked lexicographically, with no composite score.

Results

Even the best model recovers about a third

Entry-F1 (%) by model

SmallMid-sizeLargeFrontier-scale
MiniMax-M2.7 230B
34.7
Qwen3.5-397B 397B
26.3
Qwen3.5-122B 122B
25.2
Qwen3.6-35B 35B
15.0
Mistral-Medium-3.5 128B
13.6
Gemma-4-31B 31B
12.1
Gemma-4-E4B 4B
8.5
GPT-OSS-120B 120B
8.0
Mistral-Small-4 119B
3.8
Qwen3.5-9B 9B
3.5
Granite-30B 30B
2.8
Llama-3.3-70B 70B
0.3

Macro-average over 25 runs per model. Type-F1 for the best model is 21.7, and it trails Entry-F1 on every model, by up to 13 points.

A low ceiling

The best model, MiniMax-M2.7, reaches 34.7% Entry-F1. More than two thirds of fraudulent entries are missed.

Detecting is not typing

Agents often flag fraud entries but misidentify the coordinated pattern they belong to.

Some schemes are hard

Revenue manipulation is the accessible scheme (11 of 12 models detect it). Fictitious AP and shadow payroll set the ranking, and seven models score 0% recall on shadow payroll.

The task is solvable

A hand-written rule oracle recovers 78 to 91% of the schemes with the same database access, while an uninformed audit-test battery reaches only 0.36% Entry-F1. The gap is model capability, not label noise.

Leaderboard

Submit your agent

The unlabelled ledgers are released on Hugging Face and scoring is automatic. Run your own agent on the ledgers, upload your flags (document id, scheme type) as a single run or as several replicates, and get your Entry-F1, Type-F1 and Coverage. Labels stay private, and each entry comes with the code of its harness for verification.

Citation

@inproceedings{waffo2026forensicbench,
  title={ForensicBench: Evaluating Agentic LLMs on Journal-Entry Fraud Detection},
  author={Waffo Dzuyo, Guy Stephane and Guibon, Ga{\"{e}}l and Cerisara, Christophe and Belmar-Letelier, Luis},
  booktitle={EMNLP 2026, Industry Track},
  year={2026}
}

Back to apps