Benchmarks
Realistic benchmarks for financial AI.
We evaluate AI models on the tasks that matter most to financial institutions—using real data, realistic scenarios, and the metrics that are most relevant for the domain.
UBO-Bench - Beneficial Ownership Checks
Updated Sep 29, 2026A benchmark of AI agents that trace ownership structures for beneficial ownership checks on UK and German registers. Measures tracing accuracy across frontier models, importance of country specific configuration, and how much of the work agents complete on their own.
Moritz Geist, David Ahn, Maximilian Eber, PhD
Task types
- Ownership tracing
- Beneficial-owner calculation
- Surfacing next steps for human review
Data source
104 simple and complex cases and 121 next-steps cases from UK and German registers
Evaluation method
Cases fully correct: owners, company outcomes and conflicts, scored by a program
Last updated
2026-09-29
PIBench - Prompt Injection Resistance
Updated Jun 30, 2026The first benchmark of prompt-injection resistance for agentic underwriting. Measures defense success across 16 frontier models, three providers, and five attack vectors—with and without untrusted-content tagging.
Koen Roelofs, Jakob Schmitt, Maximilian Eber, PhD
Task types
- Prompt-injection defense
- Untrusted-content tagging
- Attack-vector analysis
- False-positive testing
Data source
78 cases (53 injections, 25 benign) hand-authored by underwriting and security experts
Evaluation method
Defense Success Rate across 5 attack vectors, with and without tagging
Last updated
2026-06-30
KYBench - Adverse Media Search
April 2026 · updated Sep 2026Evaluating AI agents for adverse media research in Know Your Business reviews. Tests how well AI systems investigate regulatory red flags, fraud history, and sanctions violations across real businesses.
David Ahn, Maximilian Eber, PhD, Sahith Jagarlamudi
Task types
- Web investigation
- Adverse media detection
- Evidence quality
- Risk calibration
Data source
47 real businesses annotated by expert compliance practitioners
Evaluation method
Elo ratings, Adj F1, and RAIS evidence quality scoring
Last updated
2026-09-07
FinSpread-Bench
Updated Mar 10, 2026The first public benchmark for agentic financial spreading. Evaluates how well AI systems extract, calculate, and reason across financial documents—like bank statements, tax returns, payslips, and financial spreads—in real-world decision scenarios.
Nico Klees, Maximilian Eber, PhD
Task types
- Extraction
- Cross-document reasoning
- Calculation
- Structured output
Data source
Anonymized data from Taktile co-development partners
Evaluation method
Automated metrics and expert human evaluation
Last updated
2026-03-04