Beyond static benchmarks: How Kensho evaluates AI retrieval across evolving financial data
Unstructured financial data never stops changing, so fixed benchmarks can't always tell us whether a retrieval agent found the right evidence. Our evaluation approach captures how retrieval actually performs in production, anchored by a custom harness that freezes each question in time to verify whether the agent returned trusted data.
Authors: Tayyibah Khanam, Adam Lineberry
Financial professionals work across a large and continually changing body of information, from structured datasets such as company financials and market data to unstructured sources like earnings call transcripts, investor presentations, and corporate filings. As more research and analysis shifts to AI applications and agentic workflows, the real challenge is surfacing precisely the right information grounded in trusted, verifiable data.
The S&P Global AI Data Portal solves for reliable, scalable data access with two retrieval methods, built to work with any AI and agent architecture. The retrieval path for more complex, open-ended research and agentic workflows is Adaptive Retrieval. This solution breaks down natural language questions into sub-queries, routing each to a specialized agent built with subject-matter expertise to navigate a particular dataset within S&P Global’s vast data estate and return citation-backed responses.
For the teams developing unstructured data retrieval agents at Kensho, this architecture reflects the complexity of agentic research workflows and raises a fundamental question: How do we know whether an agent retrieved the best available evidence and produced a trustworthy response?
Finding the answer to that question, at enterprise scale and over data that never holds still, has taken Kensho well beyond traditional retrieval benchmarks.
Evaluation surfaces
Our unstructured search agents return verbatim chunks from grounded documents, not finished answers. Evaluation can examine different surfaces, each offering a different view of retrieval quality:
Returned chunk content: The retrieved evidence is arguably the most important surface. Gold snippets, (the text spans a subject-matter expert labels as the evidence a question should return) provide a direct reference, but without special handling they quickly become stale as new documents are published.
Returned chunk metadata: Information in the metadata, such as expected companies mentioned, dates, fiscal periods, and document types, provide criteria that tends to remain stable even as the corpus changes. These checks assess whether retrieval stayed within the question’s scope, though not whether the content answers it.
Returned chunk identities: Comparing which chunks are returned across repeated runs reveals retrieval stability. Overlap measures consistency, not correctness: different chunks may provide equally valid evidence, especially for broad or ambiguous questions.
Agent trace: The agent’s queries, tool calls, and search decisions reveal whether it followed its instructions and used tools with appropriate arguments. Non-determinism makes a single prescribed gold trace (one exact sequence of steps labeled as correct) too restrictive, but developers can still evaluate whether the trace follows an expected general shape.
Generated answer: Generating an answer goes beyond the agent’s output contract within Adaptive Retrieval, but provides a useful proxy for the evidence it returned and its usefulness in downstream workflows. This can make comparisons easier in some circumstances.
Why established retrieval benchmarks are not enough
Structured data comes with its own answer key. A revenue figure is right or wrong, and the source data says which. Unstructured financial information has no such key. The evidence is scattered across prose, and reasonable people can disagree about which passages count. Relevant evidence may be found anywhere from management commentary on an earnings call, to a disclosure in a regulatory filing, to other sources of information spread across several documents. Identifying it can require interpreting time periods (whether explicitly mentioned or not), companies, document types, and financial terminology.
Traditional information-retrieval evaluation sidesteps this. The eval begins with a fixed collection of documents, a set of questions, and relevance judgments marking which passages should be returned. It then scores methods with repeatable metrics such as recall@k, mean reciprocal rank, and nDCG. Benchmarks like MS MARCO and BEIR have made this framework invaluable for research but it rests on two assumptions: the corpus is fixed, and the correct result can be defined in advance.
Production financial data satisfies neither. New documents arrive continuously, so any gold labels decay as the world moves. Many questions admit more than one reasonable interpretation, so "correct" is at least partly subjective—even subject-matter experts may not agree on the ideal answer.
Public benchmarks remain a useful starting point for comparing retrieval methods in the abstract, but a system can score well against a static reference set while performing poorly in the ever-evolving present of production. Evaluating these systems as users actually experience them takes more than a single metric on a frozen test set; it requires custom methods built around comparison, variability, evaluator bias, and time.
Comparing systems without a single correct answer
When gold labels aren't available, one option we explored is a form of unsupervised A/B testing. Because our agents return document chunks rather than finished answers to queries, we first have an LLM write an answer grounded on the retrieved chunks, then compare answers generated across two versions of the system. A judge reads both and picks the one it believes better addresses the question.
We use the answers as a proxy for the retrieved content because set-to-set comparison in this context is ill-defined. Two chunk lists differ in count, length, granularity, overlap, and order. Judging “which list is better” has no natural aggregation rule, and classical IR metrics like nDCG assume a human scanning top to bottom, which is not how a downstream LLM reads them.
The comparison can weigh aspects such as answer relevancy and whether the answer generated covers the important parts of the question. Across a broad enough set of questions, this estimates how often one system is preferred over another.
But answer quality evaluation trades one problem for another: the quality of the judge. LLM judges can be swayed by things unrelated to answer quality, such as verbosity, formatting, writing style, and even the order the two answers are shown in. Techniques like order reversal and length analysis help expose these biases, but a judge is only trustworthy to the extent that its verdicts agree with SMEs, and confirming that takes a set of labeled comparisons to calibrate against.
Calibrating the judge takes expert time to tune its prompt. Pairwise is also slower to run, since it leans on multiple LLM calls for quality analysis rather than a quick LLM check. Given this, we treat this approach as promising but not yet proven. A confident score is dependent on answer generation and judge behavior, not whether it agrees with human judgment, and until we close that gap it isn't something we rely on in practice.
Balancing consistency with flexibility
Agentic retrieval is not a single deterministic lookup; it's a chain of non-deterministic LLM-driven decisions. Even with generation settings tuned to minimize randomness, an ambiguous question can be interpreted in more than one reasonable way. So the same question, asked twice, can produce different agent behavior and results.
Maximum stability is not the goal. For open-ended questions, the freedom to interpret each question afresh and chart its own retrieval path is what lets the agent bring its full reasoning to bear on the best available evidence; forcing consistency would mean constraining that reasoning and lowering answer quality to buy lower variance. The aim is to balance the two.
We don’t want to eliminate variation, but to tell genuine flexibility apart from instability.
Freezing time in a moving corpus
Gold snippets provide one of the most direct ways to evaluate retrieval. Subject-matter experts identify the evidence a question should return, and we assess the retrieved content against that reference. But those annotations are expensive to produce and hard to keep valid when the corpus keeps growing and users expect answers from the newest information in it. New filings, earnings calls, and news quickly supersedes the selected evidence.
The question’s meaning can shift too, as events unfold or the companies and time periods it refers to change. Without controlling for this drift, a system retrieving better, newer evidence can appear to perform worse against yesterday’s gold. To preserve those annotations, we built a custom evaluation harness that supports a frozen date for each question, typically the date its labels were created.
The harness restricts retrieval to information available as of that date and gives the search agent and evaluator the same temporal reference. Unlike maintaining a duplicate or truncated database tied to a single cutoff, this approach supports arbitrarily many frozen dates within one evaluation set. New questions can be labeled against current information and added alongside older questions, each retaining the context in which its gold evidence was selected.
This solution addresses corpus drift, but not question drift: the evaluation set still needs periodic refreshing to represent new topics and changing user needs. Freezing the evaluation environment also cannot erase information a model learned during training, including knowledge of events after a question’s frozen date. These limits matter, but the central benefit remains: expensive expert annotations stay useful long enough to support repeatable evaluation, without forcing every question into the same historical snapshot.
Evaluation should reflect reality
Evaluating retrieval agents takes more than picking a benchmark or asking an LLM to score an answer. The design has to reflect how the system behaves in production and how users decide whether an answer is trustworthy.
Each technique addresses a specific gap. Clock freezing enables gold-labeled content checks. Answer quality comparison measures relative quality when there's no gold. Metadata checks ensure the response has the right shape. Trace-based evals tell us if the agent is behaving as intended. Stability analysis detects regressions in retrieval consistency.
Together, they support a more useful definition of retrieval quality: not whether the system produced a single convincing answer to the test, but whether it consistently found the right evidence and stayed within the scope and time frame of the question..
As the S&P Global AI Data Portal helps users work across structured and unstructured data, rigorous evaluation is how we ensure that changes to the system are real improvements — reflected in the experience users actually receive.