Apple, Slab, TheScore, Superlist, etc. ✦ 16+ years of polyglot experience, focus on Web & Cloud ✦ StackOver ow: 75,000+ score (Top 5 in Pakistan) ✦ Author / Contributor of multiple famous libraries & tools ✦ Featured on popular developer communities fl ✦
a chatbot ✦ Struggled a lot with AI-assisted coding ✦ Code quality was extremely poor ✦ Often had to spent time xing it ✦ Or throwing it away and doing manually fi ✦
experienced developers were 19% slower with early-2025 AI. Google's DORA 2024 research found AI adoption reduced delivery stability, continuing into 2025 despite higher adoption & throughput. SOURC E S O UR CE
of modern LLMs as a Google replacement Fully delegating code to AI without reviewing output Accelerating professional software engineering with AI YO U AR E HE RE
Agent = Model + Harness ✦ "Everything other than the model" ✦ Prompt, Evals, Tool Calls, Docs, Context, etc. ✦ Even the GUI/CLI "agent" tool you use “ Agent = Model + Harness Vivek Trivedi (Researcher, LangChain)
42% → 78%, 46% → 80%, 23% → 45% ✦ ~22 point swings vs ~1 point swings ✦ Using frontier models S A ME HA R N ES S Different Model S A ME MO DE L Scaffold Changes ~1 PO IN T SW IN GS ~22 PO IN T SW IN GS
Built into your coding agent (CLI/GUI tool) ✦ System prompt, Tool calls, Orchestration M O DE L Outer Harness (User) I N N E R H ARN E SS ✦ Controls put in place by users ✦ User prompt, Agent rules, Output validation ✦ Our focus today O UTER HARNESS
ADR STYL EGUI DES REF EREN C E DO CS RUL ES SC RI PTS / CL I TO O LS CO DEM O D S L AN GUAGE SERVERS ... Feedforward Guides I NI TIAL G E N E RATI O N ✦ HUMAN AGEN T UNIT TESTS E2E TESTS STATI C AN ALYSI S REVI EW AGEN TS LO G S BROWSER L I N TERS SBO M VAL I DATI O N SEC URI TY SCAN N ERS ... Feedback Sensors SEL F - C O RRECTI N G
maintainability and architectural quality Focus on Deterministic controls rst ✦ I M PLE M ENTATIO N L AYE RS 1. L I N TI N G & STATI C C H EC KS 2. UNIT TESTS Fast, reliable, cheap 3. I NTE G RATI O N/ E2 E Implementation Layers 4. AI REVI EWS ✦ Fastest & accurate feedback early ✦ Goal: Push agents' reliable coverage as far up as possible fi ✦ 5. MANUAL Q A
security mandates, architecture patterns ✦ Keep AI out of writing tests, preserve double-bookkeeping Build Reusable Harnesses ✦ CI templates with common deterministic checks ✦ Inferential review agents for security, architecture, gap analysis, even PR reviews Scale via Service Templates ✦ Service-level AGENTS.md
templates ✦ Internal team guides ✦ Codemods & internal tools ✦ Boilerplate projects Embed harnesses directly in them ✦ Scaffold not just code, but AI knowledge and conventions from day one ✦ Inter-organization review agents fi ✦