OneBench: AI Insights
OneBench evaluation observatory // Finance and banking
Finance AI evaluation intelligence

Know what an AI benchmark proves—and what it does not.

A banking-focused map of public benchmarks, assurance frameworks and institution-owned test packs for selecting, approving and continuously monitoring AI systems.

Evaluation assets19

Benchmarks, frameworks, standards and bank-specific test packs.

Publicly reusable17

Open datasets, code or published methodologies that can accelerate a bank evaluation.

Known public gaps2

Critical functions where a credible institution-owned golden set is still required.

Live intelligence12

Latest evaluation and benchmarking stories found in the OneBench evidence feed.

The OneBench evaluation position
A leaderboard is a screening tool, not a model approval.

A defensible G-SIB decision combines external benchmarks with a versioned use-case test pack, control tests, human baselines, cost and latency measures, and post-deployment monitoring. Scores must be reproducible against the exact model, prompt, retrieval layer, tools and policy configuration being deployed.

A bank-grade evaluation stack
01External screen

Use public benchmarks to shortlist models and expose obvious weaknesses.

02Use-case golden set

Test representative, difficult and high-risk cases drawn from the bank's work.

03Controls & adversarial

Probe leakage, entitlements, prompt injection, conduct, bias and unsafe actions.

04Operational fitness

Measure latency, cost, availability, fallbacks, auditability and human escalation.

05Continuous evaluation

Rerun on model, prompt, data, tool and policy changes; monitor production drift.

Coverage by bank function

Select a function to jump directly into the relevant benchmark set. The bar shows the number of publicly reusable evaluation assets; amber markers identify where a local bank test pack is essential.

Click any function to explore
Benchmark and framework catalogue

Filter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.

Primary sources linked
19 matching evaluation assets
Public benchmark

FinanceBench

Patronus AI

Established

Open-book financial question answering over public-company filings with evidence-linked answers.

Finance & investmentKnowledge & RAG
Access
Open dataset
Last reviewed
2026-07-25
Measures
4 dimensions

What it measures

  • Financial document retrieval
  • Numerical reasoning
  • Answer correctness
  • Evidence grounding

How a G-SIB should use it

Test research assistants, filing analysis, credit research and finance RAG before adding institution-specific documents.

Useful metrics

Exact / judged answer accuracyEvidence retrievalCitation correctness
Primary sourceOpen FinanceBench
Evaluation and benchmark watch

Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.

Updated with the main intelligence pipeline
Latent Space

[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)

Anthropic's rumored Claude Opus 5 offers 'Fable-level performance' at half the cost of the unreleased Fable model, according to expert commentary.

Open source ↗
Lenny's Newsletter

Claude Opus 5 review: this model is brilliant (but annoying)

A review of Claude Opus 5 benchmarked it against six leading models, with the reviewer expressing surprise at its performance.

Open source ↗
METR Research

Metrics of Agent Ability

METR Research is developing new metrics to evaluate the capabilities of AI agents, focusing on robust and reproducible measurement.

Open source ↗
Latent Space

[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model

Black Forest Labs claims its FLUX 3 multimodal flow models outperform Seedance 2.0, Gemini Omni, and Grok Imagine, and introduced a FLUX-mimic robotics model.

Open source ↗
arXiv cs.CL — Computation and Language

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

WildTrace is a new benchmark for evaluating LLMs' ability to integrate evidence from disparate parts of long documents for complex reasoning tasks.

Open source ↗
arXiv cs.CL — Computation and Language

MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

MetaHOPE is a proposed framework for evaluating metaphor translation errors in machine translation and LLM outputs, focusing on semantic and cultural complexities.

Open source ↗
arXiv cs.CL — Computation and Language

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

Research explores if Vision-Language Models (VLMs) express visual content with discourse-appropriate information structure, using Hungarian language testing.

Open source ↗
arXiv cs.CL — Computation and Language

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

FlyRoute, an arXiv paper, proposes a self-evolving framework for agent profiling that uses real traffic to dynamically update agent capabilities for task routing.

Open source ↗
arXiv cs.CL — Computation and Language

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

New benchmark, ImplicitBBQ, evaluates implicit bias in LLMs, addressing limitations of existing name-based proxies for detecting non-explicit identity biases.

Open source ↗
arXiv cs.CL — Computation and Language

Progressive in Principle, Centrist in Practice: LLM Political Bias Is Instrument-Dependent

Research finds LLM political bias depends on evaluation instrument, with abstract questionnaires showing left-leaning bias but concrete policy questions showing centrist views.

Open source ↗
arXiv cs.CL — Computation and Language

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

Research introduces MELLA, a multimodal dataset addressing culturally 'thin' descriptions from MLLMs in low-resource languages by focusing on native visual-textual alignments.

Open source ↗
arXiv cs.CL — Computation and Language

Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

Research finds LLM-based translation preserves moral semantics across languages, specifically Polish, easing cross-lingual moral classification.

Open source ↗
Methodology

OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.