Publications

Explore our work.

Benchmarks, technical reports, and in-depth research on reliable AI deployment in financial services.

Subscribe via RSS

BenchmarkSeptember 2026

UBO-Bench: Evaluating AI Agents for Beneficial Ownership Checks

Moritz GeistDavid AhnMaximilian Eber, PhD

Moritz Geist, David Ahn, Maximilian Eber, PhD

A benchmark of AI agents that trace ownership structures for beneficial ownership checks on UK and German registers. Measures tracing accuracy across frontier models, importance of country specific configuration, and how much of the work agents complete on their own.

BenchmarkJune 2026

PIBench: Prompt Injection Resistance in Agentic Underwriting

Koen RoelofsJakob SchmittMaximilian Eber, PhD

Koen Roelofs, Jakob Schmitt, Maximilian Eber, PhD

The first benchmark of prompt-injection resistance for agentic underwriting. Measures defense success across 16 frontier models, three providers, and five attack vectors — with and without untrusted-content tagging.

BenchmarkApril 2026 · updated Sep 2026

KYBench: Evaluating AI Agents for Adverse Media Research

David AhnMaximilian Eber, PhDSahith Jagarlamudi

David Ahn, Maximilian Eber, PhD, Sahith Jagarlamudi

The first public benchmark of AI-driven adverse media investigation. Evaluates detection accuracy, evidence quality, reliability across agent runs, and cost efficiency across frontier models.

BenchmarkMarch 2026

FinSpread-Bench: Evaluating Agentic AI for Financial Spreading

Nico KleesMaximilian Eber, PhD

Nico Klees, Maximilian Eber, PhD

The first public benchmark for agentic financial document processing. Evaluates extraction accuracy, cross-document reasoning, calculation correctness, and structured output quality across seven frontier models. Built on anonymized production data from financial institutions.

PaperMarch 2026

AI in AML: A guide to governance and implementation

Dustin EatonMaximilian Eber, PhD

Dustin Eaton, Maximilian Eber, PhD

Why AML teams must now apply model risk management standards to AI systems. Published in ACAMS Today, exploring how regulators are extending MRM frameworks to AI deployed in compliance functions — and what institutions need to do to prepare.