Benchmarks, frameworks, standards and bank-specific test packs.
Know what an AI benchmark proves—and what it does not.
A banking-focused map of public benchmarks, assurance frameworks and institution-owned test packs for selecting, approving and continuously monitoring AI systems.
Open datasets, code or published methodologies that can accelerate a bank evaluation.
Critical functions where a credible institution-owned golden set is still required.
Latest evaluation and benchmarking stories found in the OneBench evidence feed.
A defensible G-SIB decision combines external benchmarks with a versioned use-case test pack, control tests, human baselines, cost and latency measures, and post-deployment monitoring. Scores must be reproducible against the exact model, prompt, retrieval layer, tools and policy configuration being deployed.
Use public benchmarks to shortlist models and expose obvious weaknesses.
Test representative, difficult and high-risk cases drawn from the bank's work.
Probe leakage, entitlements, prompt injection, conduct, bias and unsafe actions.
Measure latency, cost, availability, fallbacks, auditability and human escalation.
Rerun on model, prompt, data, tool and policy changes; monitor production drift.
Select a function to jump directly into the relevant benchmark set. The bar shows the number of publicly reusable evaluation assets; amber markers identify where a local bank test pack is essential.
Click any function to exploreFilter by bank function or evaluation type. Open any entry for the measured capability, recommended banking use, metrics, limitation and primary source.
Primary sources linkedFinanceBench
Patronus AI
Open-book financial question answering over public-company filings with evidence-linked answers.
- Access
- Open dataset
- Last reviewed
- 2026-07-25
- Measures
- 4 dimensions
What it measures
- Financial document retrieval
- Numerical reasoning
- Answer correctness
- Evidence grounding
How a G-SIB should use it
Test research assistants, filing analysis, credit research and finance RAG before adding institution-specific documents.
Useful metrics
Latest evidence from the existing OneBench source network that explicitly concerns model evaluation, benchmarks, assurance, red-teaming or measurable AI quality.
Updated with the main intelligence pipeline[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)
Anthropic's rumored Claude Opus 5 offers 'Fable-level performance' at half the cost of the unreleased Fable model, according to expert commentary.
Open source ↗Claude Opus 5 review: this model is brilliant (but annoying)
A review of Claude Opus 5 benchmarked it against six leading models, with the reviewer expressing surprise at its performance.
Open source ↗Metrics of Agent Ability
METR Research is developing new metrics to evaluate the capabilities of AI agents, focusing on robust and reproducible measurement.
Open source ↗[AINews] Black Forest Labs FLUX 3 - Multimodal Flow Models that beat Seedance 2.0, Gemini Omni and Grok Imagine, and FLUX-mimic video-action robotics model
Black Forest Labs claims its FLUX 3 multimodal flow models outperform Seedance 2.0, Gemini Omni, and Grok Imagine, and introduced a FLUX-mimic robotics model.
Open source ↗WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
WildTrace is a new benchmark for evaluating LLMs' ability to integrate evidence from disparate parts of long documents for complex reasoning tasks.
Open source ↗MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors
MetaHOPE is a proposed framework for evaluating metaphor translation errors in machine translation and LLM outputs, focusing on semantic and cultural complexities.
Open source ↗When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs
Research explores if Vision-Language Models (VLMs) express visual content with discourse-appropriate information structure, using Hungarian language testing.
Open source ↗FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing
FlyRoute, an arXiv paper, proposes a self-evolving framework for agent profiling that uses real traffic to dynamically update agent capabilities for task routing.
Open source ↗ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
New benchmark, ImplicitBBQ, evaluates implicit bias in LLMs, addressing limitations of existing name-based proxies for detecting non-explicit identity biases.
Open source ↗Progressive in Principle, Centrist in Practice: LLM Political Bias Is Instrument-Dependent
Research finds LLM political bias depends on evaluation instrument, with abstract questionnaires showing left-leaning bias but concrete policy questions showing centrist views.
Open source ↗MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
Research introduces MELLA, a multimodal dataset addressing culturally 'thin' descriptions from MLLMs in low-resource languages by focusing on native visual-textual alignments.
Open source ↗Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
Research finds LLM-based translation preserves moral semantics across languages, specifically Polish, easing cross-lingual moral classification.
Open source ↗OneBench includes a benchmark when its dataset, code or methodology is publicly inspectable, or when it provides a useful assurance pattern for a regulated enterprise. “Established” describes reuse and methodological maturity, not regulatory approval. Catalogue entries are editorial assessments; live stories link to the underlying publication. Public benchmark results should always be checked for dataset contamination, judge-model bias, version mismatch and unreported inference budgets.