Cleverdon establishes Cranfield paradigm as standard for evaluation using set of queries with pre-determined relevant items. 1978: ACM creates Special Interest Group on Information Retrieval (SIGIR). They held their 49th annual conference in Melbourne. 1992: DARPA creates Text REtrieval Conference (TREC) focused on large-scale IR evaluation. It continues to this day, run by NIST.
test collections and benchmarks does not always extend beyond the lab, and it can be too incremental. The constraints of scalable, batch, offline evaluation make it difficult to take into account user behavior – or user effects in general. Industry researchers share results but obfuscate key details. Academic researchers tend to be more open, but less exciting. Nonetheless, incredible community to explore and discuss ideas.
core IR focus since Cranfield tests in 1950s. But the dependence on human judgements made it slow and expensive. Measuring recall using human judgments is prohibitively expensive. TREC pools results from the competing systems being compared. The possibility that LLMs could replace human relevance judgments has mesmerized and polarized the IR community.
frontier models offer cheaper, faster, and more scalable judgments than abysmally paid human assessors. But LLMs systematically assign inflated relevance scores, often confidently, to irrelevant results and can be highly sensitive to passage length and surface-level lexical cues. [Yu et al, 2026] Despite similar aggregate preferences, human and LLM judges exhibit low agreement on specific instances. [Gienapp et al, 2026]
proliferate, systems drift toward outputs the LLM judge recognizes, and meta-evaluation of LLM judges against a fixed human reference stops being meaningful. The same widespread adoption that makes LLM-judges effective as components in IR systems is what makes them unreliable as evaluators against a fixed reference.” [Dietz, 2026]
reject LLM-as-a-Judge for evaluation or systems, but to anchor every meta-evaluation cycle to fresh human judgment on modern systems.” [Dietz, 2026] Rather than rejecting LLM judges, Dietz and colleagues have introduced an Auto-Judge Benchmark as a way to evaluate them.
using human judgments out-performs zero-shot and few-shot LLMs. [Gienapp et al, 2026] Instead of just asking an LLM for relevance judgments, formalizing the information need improves reliability. [Keller et al, 2026] Use LLM to populate features in an LTR model, decomposing relevance into exactness, coverage, etc. [Farzi and Dietz, 2026]
that splits annotation effort between human assessors and LLMs. Humans judge the highest-ranked results while LLMs cover lower-ranked results. This approach achieves competitive evaluation reliability while drastically reducing cost. [Otero and Parapar, 2026]
solely on semantic similarity or world knowledge drift from the behavior that reveals actual user preferences, particularly for ambiguous queries. Behavior-grounded judges based on historical behavior with similar queries achieve stronger alignment with user preferences. [Vardasbi et al, 2026]
scores, but much less effort establishing that the evaluations actually measure what we care about. We need meta-evaluation to determine whether evaluation actually tells us about the metrics we care about. [Thomas et al, 2026]
lift for its App store using fine-tuned LLM-generated labels to supplement behavioral labels. [Christakopoulou et al, 2026] Amazon is using LLM judgments in combination with human judgments to evaluate e-commerce recommendation quality, mitigating position bias and serving as a scalable pre-screener. [Liu et al, 2026]
judging conversations, which in turn tends to require simulating users. IR has long struggled to evaluate anything beyond query-result pairs. The Cranfield paradigm focuses on retrieval, which is necessary but not sufficient, to make experiments cheap and statistically powerful. But meta-analysis shows critical user and query effects. Evaluating adaptive IR requires putting the user back into the abstraction. [Voorhees, 2008]
shifts from the ranked result list to the agent trajectory: what it browsed, rejected, and reasoned through. Supervision has to move inside the interaction loop, evaluating every intermediate step and not just the final answer. Retrieval becomes infrastructure: agents use retrieval as a tool, so agent–retriever co-design matters more than lexical vs. dense.
from trajectories of agent interaction data (browsing, rejections, reasoning traces). [Zhou et al, 2026] SmartSearch supervises the process, not just the answer, scoring every intermediate query a search agent issues. [Wen et al, 2026] FACE evaluates conversational search at the turn level and achieves 0.9 system-level correlation with humans. [Joko and Hasibi, 2026] There was even a debut AgentSearch workshop!
judges already ship real value. Judges can be made substantially better: classifiers, formalized needs, hybrid pooling, multi-criteria judging. There are systematic biases and failure modes: score inflation, human–LLM disagreement on preferences. The deeper problem is coupling / circularity: evaluation validity, the co-adaptation spiral, contamination.
Intent drives retrieval. Retrieval at scale still matters. We still have to think about indexing, candidate generation, and managing latency and compute cost. Ranking still matters. LLMs solve relevance, not overall ranking. And, while evaluation might be cheaper and faster with LLM judges, it’s still critically important and challenging.
the core problems of intent, scale, and candidate generation haven't changed. Evaluation is cheaper and faster with LLM judges, but still critically important and hard. Mind the pitfalls and embrace the tools. Conversational agents are becoming the new surface for search, and they move what we evaluate from ranked lists to trajectories.