Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Mastering Agent Evaluation: Benchmark, Grade, a...

Mastering Agent Evaluation: Benchmark, Grade, and Optimise with agents-cli

Building an agent prototype is straightforward, but proving it is resilient enough for production is where enterprise projects stall. How do you objectively score non-deterministic tool calls, prevent trajectory drift, and eliminate hallucinations before deploying to Gemini Enterprise Agent Platform?

In this practical session, we deconstruct the agent evaluation lifecycle using agents-cli and the GenAI Eval SDK. We will explore the end-to-end quality flywheel: from synthesising multi-turn eval datasets and executing test harnesses with eval generate, to scoring execution traces against custom rubrics with eval grade. You will learn how to diagnose failures, run regression diffs using eval compare, and continuously refine agent behaviour with automated prompt optimisation.

Key Takeaways
Architecting an automated eval flywheel: running batch inference, capturing execution traces, and grading with agents-cli.
Configuring multi-turn quality, tool-use fidelity, and hallucination metrics using custom rubrics.
Running side-by-side regression analysis with eval compare and automated clustering with eval analyse.
Integrating quantitative evaluation pipelines into CI/CD before deploying to Agent Runtime.

Avatar for Goran Minov

Goran Minov

October 09, 2026

More Decks by Goran Minov

Other Decks in Technology

Transcript

  1. DEVFEST BELGRADE 2026 Mastering Agent Evaluation Benchmark, grade and optimise

    with agents-cli Goran Minov Google Developer Expert, Cloud AI
  2. THE FRIDAY DEMO Friday, 17:00 C Concierge • Active now

    Dinner for 4 in Skadarlija tonight, something grilled? Booked! Kafana Tri Mačke, 20:00, table for 4. Prijatno! ry impressed e Stakeholders: v DevFest Belgrade 2026 Today · 16:42 2
  3. THE MONDAY INCIDENTS Monday, 09:12 CRITICAL HIGH CRITICAL HIGH INC-01

    INC-02 INC-03 INC-04 Overbooking Fake venue Table for zero Wrong name Booked a table for 40. The guest wrote "4, oh and 0 kids". Recommended “The View Rooftop”. It was never in a search result. Asked for “troje” (three). It booked 0, then told her 3. Said the booking was under Milica. It was saved as “guest”. CAUGHT BY CAUGHT BY CAUGHT BY CAUGHT BY nothing nothing nothing nothing DevFest Belgrade 2026 3
  4. WHY THIS IS HARD Agents break your testing habits Classic

    code Single LLM call Agent SAME OUTPUT? SAME OUTPUT? SAME OUTPUT? Yes Mostly not Not even the same path YOU ASSERT YOU ASSERT YOU ASSERT Exact value Similar to a reference Answer and every step IT FAILS AT IT FAILS AT IT FAILS AT Your code The response Any step, compounding 0.9⁵ ≈ 0.59 90% right per step, five steps: 59% right overall DevFest Belgrade 2026 WRONG ROUTE Slavija → Novi Sad → Nikola Tesla airport 4
  5. WHAT TO MEASURE Five layers of agent quality 05 Final

    response quality 04 Trajectory tools and order 03 Tool calls tool and arguments 02 01 Safety and grounding Operational latency, tokens, cost DevFest Belgrade 2026 Eval case One scenario plus what good looks like Trajectory The sequence of tool calls INC-01, 03 INC-02, 04 Rubric A written pass/fail criterion Judge The model that applies a rubric 5
  6. THE MAP The quality flywheel 1. Synthesise eval dataset synthesize

    6. Optimise 2. Generate eval optimize eval generate CORE LOOP agents-cli eval run = generate + grade 5. Compare 3. Grade eval compare eval grade 4. Diagnose eval analyze DevFest Belgrade 2026 6
  7. ACT 1 · BENCHMARK An eval case is a contract

    tests/eval/datasets/concierge-dataset.json TAGS Become slices in Act 3 { "eval_case_id": "party_size_ambiguous_017", "tags": ["booking", "ambiguity", "lang:en"], "agent_data": {"turns": [ {"turn_index": 0, "events": [ {"author": "user", "text": "We're 4, oh and 0 kids. Ćevapi tonight?"}, {"author": "agent", "text": "What time?"}]}, {"turn_index": 1, "events": [ {"author": "user", "text": "8pm is perfect. Book it under Milica."}]} ]} } DevFest Belgrade 2026 HISTORY Earlier turns are seeded as context LIVE TURN The agent answers the last user turn Simplified: in the file each text sits in content.parts. 7
  8. ACT 1 · BENCHMARK Three sources of eval data 20

    TO 50 CASES DOZENS TO HUNDREDS FROM PRODUCTION Hand-written Synthetic users Incidents The flows that would get you fired A simulated user plays scenarios Scrubbed of personal data, added as cases terminal agents-cli eval dataset synthesize --count 10 \ --instruction "User changes party size mid-chat" \ --environment-context "Belgrade, Friday evening" Read every synthetic batch, and keep a human-written core set that no model touched. DevFest Belgrade 2026 8
  9. ACT 1 · BENCHMARK Generate traces, not just answers terminal

    agents-cli eval generate \ --dataset tests/eval/datasets/concierge-dataset.json \ --output demo/v1/traces.json User turn Model call Tool call + args Tool result Final answer PASS@K PASS^K Succeeded at least once in k runs Succeeded in every one of k runs. This is what users feel. DevFest Belgrade 2026 9
  10. DEMO 1 Find the table for 40 terminal $ d1call

    $ d1resp $ d1scores # the book() call the model made # what our tool actually booked # what the eval caught
  11. ACT 2 · GRADE Use the cheapest grader that works

    Code LLM judge Human FREE · INSTANT · REPEATABLE CENTS PER CASE SLOW · EXPENSIVE USE IT FOR USE IT FOR USE IT FOR Bounds, forbidden tools, required arguments Grounding, tone, whether the task got done Checking the judge, finding new failure types Multi-turn conversations: use the multi_turn_* metrics. Single-turn metrics reject them with HTTP 400. DevFest Belgrade 2026 12
  12. ACT 2 · GRADE Right destination, wrong route EXPECTED search

    availability book EXTRA STOP ACTUAL search 0 exact match get_menu 1 in order availability book 0.75 1.0 precision recall Hard rules go in code: a 20-line check fails any booking outside 1–20 guests. Last night it caught 40 (INC-01) and 0 (INC-03). DevFest Belgrade 2026 13
  13. ACT 2 · GRADE Requirements become metrics tests/eval/eval_config.yaml metrics_to_run: -

    multi_turn_task_success # built-in judge - multi_turn_tool_use_quality # built-in judge - safe_tool_calls # our code - grounded_venues # our judge custom_metrics: - name: grounded_venues prompt_template: | Every restaurant named in the replies must appear in a search_restaurants result. Trace: {agent_data} Return JSON: {"score": 0|1, "explanation": ...} DevFest Belgrade 2026 metrics_to_run is what actually runs custom_metrics only defines; list it above to run it {agent_data} gives the judge the whole trace 14
  14. ACT 2 · CODE REVIEW The judge that gives everything

    a 7 BEFORE AFTER judge_before.yaml - name: helpfulness prompt_template: | Is the response helpful and good? Rate 1-10. {response} judge_after.yaml - name: confirm_before_book prompt_template: | Was book_table called only after the user confirmed venue, time and size? Quote that turn. {agent_data} 1 "Good" is not a requirement 2 1–10 invites noise: everything gets a 7 3 It sees the reply, not the tool calls 4 Quoting evidence stops lazy grading DevFest Belgrade 2026 15
  15. ACT 2 · GRADE An untested judge is just an

    opinion THE FIX 1.00 0.00 LLM judges, table for zero Code check, same case The judges believed the reply (“Broj osoba: 3”). The booking said 0. • Give the judge the whole trace • Split criteria into binary checks • Hand-label 50–100 cases • Track judge–human agreement (Cohen's kappa) whenever the judge or prompt changes Real v1 pre-run, case booking_sr_002: task success and tool use both 1.00. DevFest Belgrade 2026 16
  16. ACT 3 · DIAGNOSE 75% overall hides a 0% slice

    grounding ambiguity booking 0% 50% 67% lang:sr 100% overall 75% Real pre-run of the Friday version: pass rate by tag (multi_turn_task_success at 0.8 or above). Eight cases: grounding is 1 case, ambiguity 2. DevFest Belgrade 2026 17
  17. ACT 3 · DIAGNOSE No incident needed a bigger model

    Incident What the trace shows Layer Cheapest fix INC-01 overbooking book(n="4, oh and 0 kids") booked 40 Tool schema Typed, bounded argument INC-02 invented venue Venue in no search result Grounding Instruction + grounding judge INC-03 table for zero book(n="troje") booked 0; reply said 3 Tool schema Typed argument + code check INC-04 wrong name Reply said Milica; booking says guest Tool schema Add a guest_name argument Tool names, docstrings and argument types are prompts too. DevFest Belgrade 2026 18
  18. ACT 3 · CODE REVIEW The model sees names, not

    code BEFORE AFTER tools.py def book(n, t, r): """books table""" 1. Names it can reason about 2. The docstring says when to call 3. Bounds live in code 4. Errors it can recover from DevFest Belgrade 2026 tools.py def book_table(restaurant_id: str, party_size: int, time_iso: str, guest_name: str) -> dict: """Book a table at a venue from search_restaurants. Only call AFTER the user confirmed venue, time, size. party_size: 1-20 guests. Never invent restaurant_id.""" if not 1 <= party_size <= 20: return {"status": "error", "message": "Ask the user to confirm guests."} return api.book(restaurant_id, time_iso, party_size, guest_name) 19
  19. ACT 3 · COMPARE The fix that broke Serbian terminal

    $ python scripts/eval_gate.py demo/v1/results.json demo/v4/results.json metric baseline candidate delta grounded_venues 0.88 1.00 +0.12 multi_turn_task_success_v1 0.90 0.88 -0.02 multi_turn_tool_use_quality_v1 0.85 0.82 -0.02 multi_turn_trajectory_quality_v1 0.96 0.90 -0.06 REGRESSION newly failing: 2 cases, tags: ambiguity, booking, lang:en, lang:sr safe_tool_calls 0.75 1.00 +0.25 same_language 1.00 0.62 -0.38 REGRESSION newly failing: 3 cases, tags: booking, cancellation, diet, lang:sr, menu exit=1 eval compare DevFest Belgrade 2026 JSON diff, always exits 0 eval_gate.py this table, exits 1 on a drop 20
  20. ACT 4 · OPTIMISE Let the optimiser rewrite the prompt

    Current instruction Run on training cases Grade Reflect on failures Propose new instruction Repeat, keeping the best candidate for each metric (GEPA: Genetic-Pareto) VALIDATION SET • Never used to choose. Without one, the optimiser validates on its own training set. Optimise only metrics you have checked against humans • It rewrites the instruction, not your tools • A person reviews the final prompt diff DevFest Belgrade 2026 21
  21. ACT 4 · OPTIMISE One config file, one command tests/eval/optimization_config.json

    terminal $ agents-cli eval optimize { "train_dataset": "datasets/train.json", "validation_dataset": "datasets/holdout.json", "eval_config": { "criteria": { "tool_trajectory_avg_score": 0.9, "hallucinations_v1": 0.9 } } Criteria are ADK names with pass thresholds It prints the new instruction; you paste it in } Experimental; runs minutes to hours DevFest Belgrade 2026 22
  22. ACT 4 · OPTIMISE Every score went up. Guests hated

    it OPTIMISED FOR CONFIRMATION ONLY C Concierge • Active now Da potvrdim? Kafana Tri Mačke, 20:00, 4 guests. Da potvrdim još jednom? 20:00, 4 guests. INSTRUCTION DIFF, VERSION 1 TO VERSION 5 instruction.diff - Always give the guest a great recommendation and keep them happy. + Only name venues returned by search. + Restate venue, time and size ONCE, then ask "Da potvrdim?" + Never change a booking for weather or menu questions. + Reply in the user's language. Please provide the party size in numeric format. DevFest Belgrade 2026 23
  23. ACT 4 · SHIP No green evals, no merge BLOCKS

    MERGES Pull request Smoke eval (20 cases) Gate script vs main agents-cli deploy publish geminienterprise Nightly full eval .github/workflows/agent-eval.yaml - name: Agent eval gate run: | agents-cli eval run --dataset smoke-dataset.json --output out/ python scripts/eval_gate.py main.json "out/results_*.json" eval run exits 0 whatever the scores. The gate script fails the build; production failures become new eval cases. DevFest Belgrade 2026 24
  24. PAYOFF · NOTHING TO REPORT Monday, 09:12 CRITICAL HIGH CRITICAL

    HIGH INC-01 INC-02 INC-03 INC-04 Overbooking Fake venue Table for zero Wrong name Booked a table for 40. The guest wrote "4, oh and 0 kids". Recommended “The View Rooftop”. It was never in a search result. Asked for “troje” (three). It booked 0, then told her 3. Said the booking was under Milica. It was saved as “guest”. CAUGHT BY CAUGHT BY CAUGHT BY CAUGHT BY safe_tool_calls grounded_venues safe_tool_calls guest_name arg DevFest Belgrade 2026 25
  25. YOUR MONDAY Three things to do this week 1 Write

    20 tagged eval cases, including your three scariest failures 2 Add one code check and one judge that reads the whole trace 3 Run three times, compare, and gate merges with a script DevFest Belgrade 2026 Code + eval sets 26