Upgrade to Pro — share decks privately, control downloads, hide ads and more …

IR Today: Theory, Practice, and Agents

Sponsored · Ship Features Fearlessly Turn features on and off without deploys. Used by thousands of Ruby developers. →

IR Today: Theory, Practice, and Agents

This is a talk I gave as part of Doug Turnbull's "Retrieval Augmented Gathering" series.

https://maven.com/p/1f1f46/ir-today-theory-practice-and-agents

Avatar for Daniel Tunkelang

Daniel Tunkelang

September 29, 2026

More Decks by Daniel Tunkelang

Other Decks in Technology

Transcript

  1. Everything about search is changing! Retrieval and Ranking Hand-Tuned Scoring

    -> LTR Models -> Dense / Hybrid Retrieval Evaluation Human Judgments -> LLM-as-Judge -> Jev (!) Interface 10 Blue Links -> Rich Search Applications -> Conversational Agents
  2. IR research has been around for a while. 1950s: Cyril

    Cleverdon establishes Cranfield paradigm as standard for evaluation using set of queries with pre-determined relevant items. 1978: ACM creates Special Interest Group on Information Retrieval (SIGIR). They held their 49th annual conference in Melbourne. 1992: DARPA creates Text REtrieval Conference (TREC) focused on large-scale IR evaluation. It continues to this day, run by NIST.
  3. Academic research can be a bit…academic. The emphasis on curated

    test collections and benchmarks does not always extend beyond the lab, and it can be too incremental. The constraints of scalable, batch, offline evaluation make it difficult to take into account user behavior – or user effects in general. Industry researchers share results but obfuscate key details. Academic researchers tend to be more open, but less exciting. Nonetheless, incredible community to explore and discuss ideas.
  4. Everyone is talking about LLMs as judges. Evaluation has been

    core IR focus since Cranfield tests in 1950s. But the dependence on human judgements made it slow and expensive. Measuring recall using human judgments is prohibitively expensive. TREC pools results from the competing systems being compared. The possibility that LLMs could replace human relevance judgments has mesmerized and polarized the IR community.
  5. LLM judges are efficient, but biased. Even the most expensive

    frontier models offer cheaper, faster, and more scalable judgments than abysmally paid human assessors. But LLMs systematically assign inflated relevance scores, often confidently, to irrelevant results and can be highly sensitive to passage length and surface-level lexical cues. [Yu et al, 2026] Despite similar aggregate preferences, human and LLM judges exhibit low agreement on specific instances. [Gienapp et al, 2026]
  6. There’s a more serious issue: circularity. “As judge-style system components

    proliferate, systems drift toward outputs the LLM judge recognizes, and meta-evaluation of LLM judges against a fixed human reference stops being meaningful. The same widespread adoption that makes LLM-judges effective as components in IR systems is what makes them unreliable as evaluators against a fixed reference.” [Dietz, 2026]
  7. Serious, but not fatal. “The constructive path is not to

    reject LLM-as-a-Judge for evaluation or systems, but to anchor every meta-evaluation cycle to fresh human judgment on modern systems.” [Dietz, 2026] Rather than rejecting LLM judges, Dietz and colleagues have introduced an Auto-Judge Benchmark as a way to evaluate them.
  8. Can we improve on LLM judges? Fine-tuning an LLM-based classifier

    using human judgments out-performs zero-shot and few-shot LLMs. [Gienapp et al, 2026] Instead of just asking an LLM for relevance judgments, formalizing the information need improves reliability. [Keller et al, 2026] Use LLM to populate features in an LTR model, decomposing relevance into exactness, coverage, etc. [Farzi and Dietz, 2026]
  9. Combine LLM and human judges. Use a hybrid pooling strategy

    that splits annotation effort between human assessors and LLMs. Humans judge the highest-ranked results while LLMs cover lower-ranked results. This approach achieves competitive evaluation reliability while drastically reducing cost. [Otero and Parapar, 2026]
  10. Relevance vs. revealed preferences. Spotify found that LLM judgments based

    solely on semantic similarity or world knowledge drift from the behavior that reveals actual user preferences, particularly for ambiguous queries. Behavior-grounded judges based on historical behavior with similar queries achieve stronger alignment with user preferences. [Vardasbi et al, 2026]
  11. Relevance isn’t everything. IR has spent enormous effort improving evaluation

    scores, but much less effort establishing that the evaluations actually measure what we care about. We need meta-evaluation to determine whether evaluation actually tells us about the metrics we care about. [Thomas et al, 2026]
  12. Industry is leveraging LLM judgments. Apple achieved a 0.24% conversion

    lift for its App store using fine-tuned LLM-generated labels to supplement behavioral labels. [Christakopoulou et al, 2026] Amazon is using LLM judgments in combination with human judgments to evaluate e-commerce recommendation quality, mitigating position bias and serving as a scalable pre-screener. [Liu et al, 2026]
  13. What about agents? Evaluating agents is hard because it means

    judging conversations, which in turn tends to require simulating users. IR has long struggled to evaluate anything beyond query-result pairs. The Cranfield paradigm focuses on retrieval, which is necessary but not sufficient, to make experiments cheap and statistically powerful. But meta-analysis shows critical user and query effects. Evaluating adaptive IR requires putting the user back into the abstraction. [Voorhees, 2008]
  14. What do agents mean for IR? The unit of analysis

    shifts from the ranked result list to the agent trajectory: what it browsed, rejected, and reasoned through. Supervision has to move inside the interaction loop, evaluating every intermediate step and not just the final answer. Retrieval becomes infrastructure: agents use retrieval as a tool, so agent–retriever co-design matters more than lexical vs. dense.
  15. Agents were on the SIGIR agenda. LRAT trains agentic search

    from trajectories of agent interaction data (browsing, rejections, reasoning traces). [Zhou et al, 2026] SmartSearch supervises the process, not just the answer, scoring every intermediate query a search agent issues. [Wen et al, 2026] FACE evaluates conversational search at the turn level and achieves 0.9 system-level correlation with humans. [Joko and Hasibi, 2026] There was even a debut AgentSearch workshop!
  16. Main take-aways from SIGIR. Judges are useful and scalable: LLM

    judges already ship real value. Judges can be made substantially better: classifiers, formalized needs, hybrid pooling, multi-criteria judging. There are systematic biases and failure modes: score inflation, human–LLM disagreement on preferences. The deeper problem is coupling / circularity: evaluation validity, the co-adaptation spiral, contamination.
  17. The core search problems haven’t changed. Query understanding still matters.

    Intent drives retrieval. Retrieval at scale still matters. We still have to think about indexing, candidate generation, and managing latency and compute cost. Ranking still matters. LLMs solve relevance, not overall ranking. And, while evaluation might be cheaper and faster with LLM judges, it’s still critically important and challenging.
  18. Take-Aways The machinery of retrieval and ranking keeps evolving, but

    the core problems of intent, scale, and candidate generation haven't changed. Evaluation is cheaper and faster with LLM judges, but still critically important and hard. Mind the pitfalls and embrace the tools. Conversational agents are becoming the new surface for search, and they move what we evaluate from ranked lists to trajectories.
  19. One more thing… I’ve been working on a job search

    app and would love your feedback! https://dtunkelang-job-search.hf.space/