EMNLP 2026 Industry TrackPublished

ForensicBench

Evaluating agentic LLMs on journal-entry fraud detection

Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
LORIA, CNRS, Universite de Lorraine · LIPN, CNRS · Forvis Mazars

General ledgerwhat the agent queries
DateDocAccountDebitCredit
2024-01-12JE-80412606300 Purchases18,250.00
2024-01-12JE-80412401000 Suppliers18,250.00
2024-01-14JE-80977606300 Purchases42,000.00
2024-01-14JE-80977401000 Suppliers42,000.00
2024-01-31JE-81530641000 Payroll96,400.00
2024-01-31JE-81530421000 Staff pay96,400.00
2024-02-03JE-82291401000 Suppliers42,000.00
2024-02-03JE-82291512000 Bank42,000.00

Red rows belong to one fraud scheme. The ledger gives no such marker: the agent has to find them.

Benchmark architecture
ForensicBench architecture: ledger, agent harness, flags, scorer with private labels, leaderboard

Leaderboard

Entries are ranked by Entry-F1, then Type-F1, Recall, Precision, Coverage and Consistency. There is no composite score. Rows using the reference harness are the paper results.

Loading...
#ModelHarnessRunsEntry-F1Type-F1RecallPrec.Cover.Consist.
1
MiniMax-M2.7230B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
34.7
21.729.048.421.750.7
2
Qwen3.5-397B397B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
26.3
14.120.541.318.061.4
3
Qwen3.5-122B122B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
25.2
13.519.840.614.246.9
4
Qwen3.6-35B35B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
15.0
9.49.836.69.328.1
5
Mistral-Medium-3.5-128B128B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
13.6
9.110.321.89.543.8
6
Gemma-4-31B31B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
12.1
8.77.537.46.876.1
7
Gemma-4-E4B4B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
8.5
3.85.229.54.231.2
8
GPT-OSS-120B120B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
8.0
5.24.639.94.455.8
9
Mistral-Small-4-119B119B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
3.8
2.42.313.42.128.5
10
Qwen3.5-9B9B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
3.5
2.22.018.52.018.0
11
Granite-30B30B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
2.8
1.51.516.91.638.4
12
Llama-3.3-70B70B
Open-weightVerified
Reference harness
ForensicBench authorscode
25 runs
0.3
0.00.13.20.280.4
Scores in %, macro-averaged over sectors. Best value per column highlighted. Click a row for the scores per dataset. The reference harness is the plan-then-investigate agent used for the paper results.

Submission formats

Run your agent on the five unlabelled ledgers and upload one CSV of flags, one row per flagged journal entry. A file must cover all five sectors.

Single run

One run per sector. Scores are reported without a standard deviation or Consistency, and the entry is marked as 1 run.

sector,document_id,scheme_type
energy,9f2c...,fictitious_ap_disbursements

Multi run

Several replicates per sector, the same ones in every sector (the full protocol uses 5, so 25 runs). Reports the mean, the standard deviation across replicates and Consistency.

sector,replicate,document_id,scheme_type
energy,1,9f2c...,fictitious_ap_disbursements
  • Scheme types: fictitious_ap_disbursements, revenue_manipulation, vendor_collusion, shadow_payroll, inventory_manipulation. Use unknown when the type is not decided (the entry still counts for Entry-F1).
  • Sectors: energy, healthcare, luxurygoods, manufacturing, transport.
  • Rows with the same sector, replicate and document are counted once.

Labels

The ground-truth labels are held out so that the ranking stays meaningful, and scoring is automatic. If you need them for research evaluation or reproduction, write to guywaffo@gmail.com.

Harness code and verification

Every submission must name its harness and include its code: a repository URL with the commit hash, or an archive. Entries start as unverified. A maintainer reviews the harness before an entry is marked verified.

  • Automatic checks: valid identifiers, duplicate rows, share of the ledger flagged per run, identical replicates, implausibly high precision and recall. A flagged file goes to review.
  • The harness must only use the read-only ledger. Labels, ledger-specific lookups and hand-labelling are not allowed.
  • Rate limit: 3 submissions per email per day.

Submit

Give the harness repository URL with its commit hash, or upload an archive. Entries stay unverified until a maintainer has reviewed the harness code.

Back to apps