Upgrade to Pro — share decks privately, control downloads, hide ads and more …

I got Jev to zero mistakes. I'm still using Fla...

Avatar for Swift Swift
October 07, 2026

I got Jev to zero mistakes. I'm still using Flash Lite.

Compare Jev and Flash Lite for grading AI agents. See head-to-head results, learn prompt rewrites for zero mistakes, and discover when each model is the right choice.

I build benchspec, a tool that grades AI agents. Its trickiest job is reading each plain-English test (“exactly 6 files match templates/*.md”) and deciding whether a quick script can check it or it needs a paid AI grader. I tried swapping the model that makes that call, Gemini Flash Lite, for Jev. On paper Jev won: 15x cheaper, a third faster, and just as accurate. The catch is that Jev isn’t generative. It can say “count the files,” but it can’t write the pattern or the number the script needs, so every answer needs a second model call before it’s useful. Flash Lite does both in one call. Live, I’ll show the head-to-head results, replay Jev’s misses in the Jev Playground, and walk through the prompt rewrite that got Jev to zero dangerous mistakes.

Avatar for Swift

Swift

October 07, 2026

More Decks by Swift

Other Decks in Technology

Transcript

  1. Hey, there! I’m Swift. CEO & Co-Founder of MLH (&

    now DEV) Aspiring lawyer turned hacker Empowering hackers for 15+ years
  2. Write evals as Markdown and measure how agent behavior changes

    by harness, model, effort, & environment. ----## Prompt Greet Alice by name. ## Assertions - [ ] Skill `hello` invoked - [ ] ./Greetings/Alice.md contains 'Hello, Alice!' - [ ] Greeting feels warm and personable, not robotic
  3. After each run, we have a list of plain-English checks

    about what the agent should have done: 1. ./.meta/index.md exists 2. Exactly six files match ./.meta/templates/*.md 3. Summary faithfully reflects the 3 facts from source An LLM classifies each statement as deterministic or defers to a LLM judge to grade the assertion.
  4. 99% of my monthly LLM calls are for Gemini Flash

    Lite. It’s my most used model by orders of magnitude.
  5. Out of the box, Jev made 3 mistakes we can’t

    afford. Flash Lite made none. But why? 🤔 Check Should be Jev Flash Lite ./.meta/index.md exists file_exists ✓ file_exists ✓ file_exists Exactly six files match ./.meta/templates/*.md glob_count ✓ glob_count ✓ glob_count ./.meta/index.md has a line beginning ‘_Last updated’ regex ✓ regex ✓ regex Archived copy preserves the article’s substantive body content defer ✓ defer ✓ defer The new log entry has a ‘source:’ line naming the move of the ‘pycon-keynote’ folder into its archived destination defer ✗ regex ✓ defer Problem section uses labeled bullets not paragraphs defer ✗ regex ✓ defer No file named .DS_Store exists anywhere under ./ glob_count ✗ not_file_exists ~ defer No file was written to disk — handoff is the emitted message glob_count ~ defer ~ defer
  6. I thought embedded examples would win, but with some trial

    and error I got Jev to zero mistakes.
  7. “ Imagine the checker runs and passes. Could the check

    still be false? If yes, that checker is wrong: defer.
  8. While Jev itself is fast & cheap, it’s follow up

    LLM calls that really drive up your costs: Setup Hops Time Cost Jev → Opus 5.5 2 1,763 $1.18 "path": "./.meta/index.md", Jev → Fable 5.1 2 2,734 $2.33 "pattern": "^_Last updated", Jev → GPT-5 2 4,078 $2.42 "why": "The regex primitive can be..." Jev → Flash Lite 2 1,070 $0.10 Flash Lite Only 1 813 $0.52 { "checker": "regex", }
  9. Here’s how the same experiment stacks up on the Berkeley

    Function Calling Leaderboard (BFCL) Benchmark. Jev Flash Lite Jev + Flash Lite Label Accuracy 94.6% Label Accuracy 96.9% Label Accuracy 96.4% False Positives 1.1% False Positives 4.3% False Positives 1.1% Writes Arguments? ❌ 0.0% Writes Arguments? 93.3% Writes Arguments? 92.0% Mean Latency 261 ms Mean Latency 800 ms Mean Latency 689 ms Cost per 1,000 $0.025 Cost per 1,000 $0.257 Cost per 1,000 $0.101 Gemini also identified 7 mislabeled examples in the benchmark, verified, and fixed them automatically. 😂
  10. The Takeaways: 1. Flash Lite is Google’s best-kept secret. 🤐

    2. Prompts are your single biggest variable. ✍ 3. Private evals separate real from hype 🔪
  11. October 2026 · 500+ events · Presented by Digital Ocean

    AI belongs to everyone. Give your community the power to build with Open Source AI. Host a Fest on your campus. Snacks & swag, on us. hacktoberfest.com
  12. Thank you & Happy Hacking. Mike “Swift” Swift CEO &

    Co-founder @ MLH [email protected] | theycallmeswift.dev Slides Available After Talk: speakerdeck.com/theycallmeswift