Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Stop Paying for Easy Prompts: Architecting Hybr...

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest. →
Avatar for Abed Matini Abed Matini
September 10, 2026

Stop Paying for Easy Prompts: Architecting Hybrid Local-to-Cloud AI Systems

Not every AI request needs an expensive cloud model. This talk presents a practical architecture for routing predictable workloads to local models while reserving cloud LLMs for complex tasks.

Learn how to design explainable routing, response validation, controlled fallback, sensitive-data redaction, conversation memory, telemetry, and governed local worker pools using tools such as Ollama and LiteLLM.

Avatar for Abed Matini

Abed Matini

September 10, 2026

More Decks by Abed Matini

Other Decks in Technology

Transcript

  1. E S CA PE 2 02 6 · 1 0

    S E PT E M B E R · CA PE TOWN / J O H A N N E S B U R G Stop paying for easy prompts Architecting hybrid local-to-cloud AI systems 45 minutes · 14:00 SAST / 17:30 IST / 13:00 BST Abed Matini · escconf.com 1
  2. THE PROBLEM One model handles every request Summarise a standup

    Extract fields from an invoice Review a security architecture Different difficulty. Different risk. Same premium cloud path. That default adds avoidable cost, latency, and data exposure. 2
  3. THE DECISION Route by required capability Easy → local Hard

    → cloud Short, repeatable, testable, low-risk tasks. Ambiguous, high-risk, multi-step, or contextheavy tasks. Summarise · classify · extract · rewrite Architecture · threat models · complex decisions “Easy” describes the required capability—not simply prompt length. 3
  4. THE IDEA One application. Two execution paths. Application │ ▼

    Policy router ──► local SLM ──┐ │ │ └──────────► cloud LLM ─── ─► validated response ┴ Before cloud: detect / redact sensitive data After local: validate / fall back once if needed The router chooses the cheapest path that meets the quality and policy bar. 4
  5. MENTAL MODEL Packet routes 1 · INSPECT 2 A ·

    L O C A L PAT H 2 B · C L O U D PAT H Request enters policy Use local when evaluated Escalate when required Classify the task, sensitivity, required capability, output contract, and consequence of error. Routine and testable Approved for the local model No cloud API charge Complex or ambiguous Higher reasoning requirement Policy permits the boundary → Each request selects one path. “Easy” means the local model passed that workload’s evaluation. 5
  6. GUARDRAIL Local is not automatically better Small models have lower

    capability ceilings. Local hardware still has operational cost. Sensitive tasks may require stronger controls—not merely local execution. A 0.5B demo model is not a production recommendation. Quality is the gate. Cost optimization starts only after the output passes. 6
  7. ROUTING Easy vs hard is a policy decision L O

    C A L C A N D I D AT E C L O U D C A N D I D AT E Predictable workload Demanding workload “Extract these fields into the required JSON schema.” “Compare architectures and defend a migration strategy.” Known task class Low ambiguity and bounded output Local model passed the task eval Multi-step reasoning High ambiguity or consequence Local quality is not proven → Prompt length is only one signal. Task capability, sensitivity, risk, and evidence decide the route. 7
  8. ROUTING DECISION Easy is an evaluated workload class Inspect useful

    signals Task and required capability Sensitivity and consequence of error Size, ambiguity, and output contract Whether the local model passed this task’s evaluation def choose(request): if request.forced: return request.route if hard_markers or words >= 120: return "cloud" if easy_markers and words <= 80: return "local" if words <= 40: return "local" return "cloud" # ambiguity defaults safely The demo uses transparent heuristics so you can inspect every reason. Production thresholds should come from labelled traces and task evaluations—not prompt length alone. 8
  9. RELIABILITY Routing is a request lifecycle 1. Inspect the request:

    task, sensitivity, size, risk. 2. Choose local or cloud. 3. Validate schema, safety, and task-specific quality. 4. Fall back once when the local result fails the bar. 5. Return one response and record what happened. 9
  10. REQUEST LIFECYCLE Full journey: one path at a time 1

    · INSPECT 2 · EXECUTE 3 · VA L I D AT E 4 · R E COV E R Policy router Choose local Contract fails Cloud fallback once Read task, sensitivity, capability requirement, size, and risk. The router sends this request to one evaluated local model. The candidate misses schema, safety, taskquality, or timeout requirements. The replacement passes the same contract, then one answer returns. → The failed local candidate is replaced—not merged. Telemetry records route, reason, latency, and fallback. 10
  11. RESPONSE CONTRACT Validate before you respond model output │ ─

    valid schema? ─ enough evidence? ─ sensitive data handled? └─ within timeout? │ ─ yes respond └─ no → cloud fallback once ├ ├ ├ ├ → Fallback is a controlled reliability mechanism—not an infinite retry loop. 11
  12. RESPONSE LIFECYCLE Validate, then respond 1 · C A N

    D I D AT E 2 · R E S P O N S E CO N T R A C T 3 · O U T CO M E Model output Required checks Pass or recover Local or cloud returns a candidate result. Returning text does not mean the request succeeded. Schema and required fields Safety and data leakage Task-specific acceptance test Route timeout budget Pass: return one answer. Fail local: try cloud once. Fail again: expose the failure. → A model produces a candidate. The application’s concrete contract decides whether it may become the response. 12
  13. STACK Give each tool one job Layer Responsibility Ollama /

    llama.cpp Run open models on local hardware LiteLLM Give local and cloud models one API shape Policy router Decide where each request should run Validators Enforce schema and quality thresholds Telemetry Measure route, latency, quality, and cost Demo: qwen2.5:0.5b locally and Gemini 3.6 Flash in the cloud. 13
  14. G A T E W AY M E C H

    A N I C S LiteLLM keeps application code stable application → policy chooses an alias → LiteLLM → provider API ─ local Ollama └─ cloud → Gemini ├ Same call shape complete("local", messages) complete("cloud", messages) → What the gateway adds Provider translation and stable aliases Normalized responses and usage metadata Proxy options for retries, health, load balancing, and fallback In this project, the policy router chooses local or cloud before LiteLLM sends the request. LiteLLM is the gateway—not the business-policy judge. 14
  15. M U LT I - T U R N C

    H A L L E N G E When the route changes, who remembers? The challenge The solution Model APIs are stateless Local and cloud share no hidden state UI history is not model context Sending every turn increases cost and may expose private data The application owns the session A context builder selects recent and relevant turns Keep a raw local view and a redacted cloud-safe view Append every answer with route and model provenance session → context builder → policy router → local or cloud ↑ │ └──────────── append response ─────────────┘ You cannot transfer a model’s hidden state or KV cache between providers. Portable memory is explicit conversation, tool state, summaries, and retrieved facts. 15
  16. ROUTING MATURITY Start explainable. Learn from traces. 1. Rules: task

    type, token length, sensitivity, risk 2. Semantic routing: compare prompt embeddings with known task classes 3. Learned routing: predict whether the cheaper model will succeed 4. Continuous evaluation: update thresholds from production traces RouteLLM and Semantic Router are later steps—not prerequisites. Always log: decision · reason · model · latency · fallback · estimated cost. 16
  17. REDACTION MECHANICS Replace the identifier, preserve the task 1. DETECT

    Email [email protected] and use sk-demo-secret 2. REPLACE Email [EMAIL_1] and use [API_KEY_1] 3. ROUTE Only the redacted text may leave the boundary 4. RESTORE Reinsert locally only when the response requires it result = redact(text) result.redacted # safe text for the selected route result.replacements # {"[EMAIL_1]": "[email protected]"} — keep local Stable placeholders preserve references across the prompt. The demo detects structured patterns with regex and validates card numbers with Luhn; production needs tested DLP, access control, and fail-closed handling. 17
  18. EDGE Minimize data before cloud 1. Detect sensitive fields locally.

    2. Replace them with stable placeholders. 3. Summarise or trim unnecessary context. 4. Route only the minimum required text. [email protected] → [EMAIL_1] Redaction reduces exposure; it does not by itself guarantee GDPR/POPIA compliance. Production needs tested DLP and governance. 18
  19. EDGE PREPROCESSING Redact, then route 1 · D E T

    E C T L O C A L LY 2 · REPLACE 3 · ROUTE MINIMUM 4 · R E S T O R E L O C A L LY Find identifiers Stable placeholders Cross the boundary Reinsert if needed [email protected] sk-demo-secret [EMAIL_1] [API_KEY_1] Send only the redacted context required to complete the task. The placeholder mapping stays inside the trusted boundary. → Redaction reduces exposure. It does not replace tested DLP, access control, audit, or compliance review. 19
  20. LIVE DEMO Four requests. One routing policy. python -m demo.router

    --scenario A # easy → local python -m demo.router --scenario B # hard → cloud python -m demo.router --scenario C # PII → redact → cloud python -m demo.router --scenario D # force-local → fallback Watch the decision: route · reason · latency · cost · fallback . 20
  21. M E A S U R E T H E

    R E S U LT Optimize a system, not a token bill Gate Question Quality Does the routed output pass the same task eval? Reliability How often do local calls fall back? Latency What are p50 and p95 end-to-end times? Cost What cloud spend was avoided after local operating cost? Privacy What data crossed the network boundary? Savings are real only when quality and reliability stay inside the target. 21
  22. YOUR FIRST PRODUCTION ITERATION Start with one narrow task 1.

    Label one week of requests: task · risk · sensitivity. 2. Pick one high-volume, testable task. 3. Add a local path behind a feature flag. 4. Evaluate it against the cloud baseline. 5. Add validation, one fallback, and route telemetry. 6. Expand only after the numbers hold. A gateway is useful. A sophisticated router can wait. 22
  23. O PT I O N A L S CA L

    E- O UT Local can be a governed pool application → policy router → local alias → one approved worker ↘ cloud only when policy requires it LiteLLM can broker Known inference endpoints Load balancing and health checks Retries and controlled fallback You still operate Discovery and device enrolment Identity, authentication, and TLS Models, patches, ownership, and audit Why: add throughput and resilience without changing application code or the easy-versus-hard policy. This repo runs one Ollama node; the pool is an architecture extension. 23
  24. O PT I O N A L S CA L

    E- O UT Peer worker pool: capacity, not model sharding 1 · G OV E R N 2 · BROKER 3 · O P E R AT E Policy router Local alias Worker A · B · C Chooses local or cloud using the same task, sensitivity, capability, and risk policy. LiteLLM sees known inference endpoints and selects one healthy approved worker per request. You still own enrolment, identity, TLS, model versions, patches, health, ownership, and audit. → The pool adds throughput and resilience. It does not make a small model smarter, and it does not split one model across laptops. 24
  25. E S CA PE 2 02 6 · 1 0

    S E PT E M B E R · E S CCO N F.CO M Abed Matini Stop paying for easy prompts — hybrid local-to-cloud AI GITHUB · CODE AND DEMO github.com/abedmatini/ai-pilot-escape HASHNODE · FULL ARTICLE abedmatini.hashnode.dev/stop-paying-for-easy-prompts SLIDESHARE · DECK Connect on LinkedIn Link coming after upload 25