Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Chatless-AI - Realtime-Sprachinterfaces für We...

Avatar for Sascha Lehmann Sascha Lehmann
October 01, 2026
11

Chatless-AI - Realtime-Sprachinterfaces für Web und Mobile entwickeln

Mit Realtime-Sprachmodellen wie GPT-realtime oder Gemini-Live entsteht eine neue Generation von Interfaces: Sprache wird zum sofort reagierenden, latenzarmen Interaktionskanal – ohne Prompting, ohne Wartezeiten, hands-free.

In diesem Talk zeigt Sascha Lehmann, wie Realtime-Modelle technisch funktionieren, wie man Kontextgrenzen, Rollen und Sicherheit zuverlässig kontrolliert und wie sich Realtime-AI gezielt in Web- und Mobile-Anwendungen integrieren lässt – von Architektur über Kostenoptimierung bis hin zur UX, die Nutzer transparent durch den Dialog führt.

Avatar for Sascha Lehmann

Sascha Lehmann

October 01, 2026

More Decks by Sascha Lehmann

Transcript

  1. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Sascha Lehmann @derLehmann_S Consultant
  2. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Sascha Lehmann Consultant @ Thinktecture AG  @derLehmann_S  https://www.linkedin.com/in/sascha-lehmann  [email protected]  https://www.thinktecture.com/thinktects/sascha-lehmann/
  3. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile What ”real-time” really means Speech interaction until now • Each arrow represents 200-500ms of latency • Like a zoom call with huge delay Mic STT (Speech to Text) LLM TTS (Text to Speech) Speaker
  4. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile What ”real-time” really means With realtime model • Permanent available audio channel • It no longer feels like software anymore, but like a real counterpart • Average response time ~320ms (OpenAI figure for gpt-realtime); full-duplex models like gpt-live-1 listen while they speak Mic STT (Speech to Text) Realtime Model TTS (Text to Speech) Speaker
  5. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Available Realtime Models Property OpenAI gpt-realtime-2.1 OpenAI gpt-realtime-2.1-mini OpenAI gpt-live-1 Gemini 3.8 Live Architecture Speech-to-Speech native (speech, reasoning, tools in one model) Speech-to-Speech native Full-duplex voice layer, reasoning & tools delegated to backend model Multimodal Live (Audio/Video/Text) Latency ~320ms first response (per OpenAI) ~320ms Full-duplex: listens while speaking Sub-800ms Languages Multilingual, language change midsentence Multilingual Multilingual, launch voices skew English 70 languages Particular strengths GPT-5-class reasoning, best instruction adherence, tool calling Cost efficient, quick, reasoning included Natural interruptions, noise handling, free choice of backend model Video input, multimodal, cost efficient Barge-In support (native, full-duplex) Emotional adaption Configurable via prompt restricted Tone, pace & style via prompt Affective Dialog native Voices 10+ voices Like gpt-realtime 12 voices (new accents & dialects) 30+ HD Voices Latest models gpt-realtime-2.1 (Jul 2026); gpt-realtime retires Jan 20, 2027 gpt-realtime-2.1-mini (Jul 2026) gpt-live-1 (Sep 10, 2026, GA) gemini-3.8-live (stable since Sep 15, 2026)
  6. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Architecture High Level • Model processes audio directly (lowest latency) • Thinks and responds in speech • Doesn’t rely on transcript • Hears emotion and intent • Filters noise • Two paradigms since 2026: gptrealtime-2.x = one model for speech, reasoning & tools. gpt-live-1 = fullduplex voice layer, reasoning & tools delegated to a backend model https://developers.openai.com/api/docs/guides/voice-agents/
  7. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Architecture Transport – WebRTC vs. WebSocket Aspects WebRTC WebSocket Protocol UDP (packet loss tolerant) TCP (guaranteed delivery) Latency Very low (~50–150ms) Higher (TCP Head-of-Line Blocking) Features Echo Cancellation, Noise Suppression, Jitter Buffer - Ideal for Browser-Apps, Client-side Server-side, Phone-integration Downsides NAT Traversal complexity Latency peaks during packet loss Also available gpt-realtime-2.x, gpt-live-1 OpenAI realtime/live, Gemini Live; SIP for telephony https://developers.openai.com/api/docs/guides/voice-agents/
  8. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Realtime Connection RealtimeAPI-Key Backend System Prompt 1 Ephemeral Key 2 3 4 App RealtimeAPI WebRTC
  9. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile 1. What is my current state? • The user has no visual clue • Voice needs a state indicator • Other methods • Audio cues • Sound effects
  10. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile 2. Success / Failure feedback • Unique Audio cue sound • Vocal feedback from the agent • Visual indicator update
  11. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile 3. Error Handling • Automatic retry • Agent asks for clarification • Control to restart the voice agent session
  12. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile 4. Interruption (Barge-In) • User needs the possibility to cancel an action • Voice activity detection (VAD) • Also include manual cancellation
  13. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Realtime Models do not only support speech • Voice is not the only interface – it is an additional one • Realtime models are also ”Multimodal Models” • Speech • Text • Images (gpt-realtime-2.x, Gemini Live) • Gemini Live also supports Live Camera Input - So the model can see what you currently see • Exception: gpt-live-1 is audio-only – no image input; visual context has to go through the backend model
  14. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile GPT-Live Full Duplex: listening while speaking • Turn-based: model waits for silence, then answers • Full-duplex: one model hears and speaks at the same time • User silence → VAD interrupt cut off, wait Agent Handles interruptions and background noise natively, no VAD tuning • Turn-taking latency 0.8 s vs. 1.4 s on gpt-realtime-2.1 • Full Duplex Bench interactivity 80 % vs. 45 % • Turn-based (gpt-realtime-2.x) Same tech behind ChatGPT Voice since Jul 2026, API since Sep 10, 2026 time Full-duplex (gpt-live-1) interjects while agent talks User Agent no gap listens while speaking, adapts, continues time Turn-taking latency and Full Duplex Bench figures: OpenAI, Sep 2026 (not independently reproduced)
  15. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile GPT-Live Architecture: voice layer + delegated backend gpt-live-1 User audio in / out Conversation layer Full-duplex speech, turn-taking, tone & pace, noise handling $0.05 / min delegate result Backend Your agent or a text model GPT-6 Astra, GPT-5.6 or any harness: reasoning, tools, RAG, guardrails billed separately (tokens) Tools APIs, search vs. gpt-realtime-2.x Trade-offs Worth it when • Realtime: one model speaks, reasons and calls tools • Live: talking and thinking are split; it keeps talking while the backend works • Per-minute price instead of audio tokens + Most natural turn-taking; backend is freely swappable – No image/video input, no structured outputs on the voice layer – Two systems to design, monitor and secure • Phone and support agents with frequent interruptions • Long sessions: tutoring, reservations, scheduling • You already own an agent backend (tools, RAG, guardrails) Source: OpenAI, Introducing GPT-Live-1 in the API, Sep 10, 2026
  16. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Pricing model of each model • ~10 Token/Sek (Input) • ~20 Token/Sek (Output) • • Cost increase with conversation length Each turn resends the context Modell Input (Audio) Cached Input Output (Audio) ~cost per minute (conversation) gpt-realtime-2.1 $32 / 1M Tokens $0.40 / 1M Tokens $64 / 1M Tokens (Text: $24) ~$0.06–0.12 gpt-realtime-2.1-mini $10 / 1M Tokens $0.30 / 1M Tokens $20 / 1M Tokens ~$0.02–0.04 gpt-live-1 $0.05 / Min (VoiceLayer) – Backend-Modell separat ~$0.05 + Backend Gemini 3.8 Live $3 / 1M Tokens Kein Caching $12 / 1M Tokens ~$0.01–0.02
  17. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile How to keep the costs under control 1. Make use of prompt caching 2. Make agent not too talkative 3. Truncation & Retention Ratio konfigurieren 4. Manage Conversation history manually 5. Make use of mini model 6. Proactive Tool calling without confirmation 7. Design clear conversation flows with a final end
  18. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Example calculation • Average duration of a pocket depth measurement: ~ 3 min • With gpt-realtime-2.1-mini + caching: ~$0.04–0.08 per finding • With Gemini 3.8 Live: ~$0.03–0.06 per finding
  19. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile System Prompt Best Practices • Role & Objective — Clear identity + task definition ("You are a dentist assistant, your task is form-filling") • Instructions / Rules — Bullet points over prose, CAPS for critical rules ("NEVER GUESS", "ALWAYS use get_selected_tooth first") • Personality & Tone — Language, formality, response length, pacing explicitly set (e.g. "max 2 sentences", formal address, "professional & concise") • Conversation Flow — State machine for procedures (Pocket Depth: Init → Measurement Loop → Completion) • Reference Pronunciations — Phonetic guides for domain terms + tooth numbers that TTS would mispronounce ("one-eight" not "eighteen") • Tools / Function Calling — Per tool: trigger condition, preamble ("One moment..."), error handling (1x silent retry, 2x notify user, 3x fall back to manual) • Safety & Escalation — Domain constraint, exact refusal scripts, escalation after 3x off-topic • Examples — Concrete input/output pairs for ambiguous cases
  20. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile System Prompt Key Techniques • Bullet Points > Prose — Model follows short bullets better than paragraphs • Labeled Sections — Model navigates by heading; each section = one concern • Lock the Language — "Respond ONLY in German" stated explicitly, prevents drift • Unclear Audio Handling (3-Step) — 1. Ask to repeat → 2. Address audio quality → 3. Offer to skip • Rotating Confirmation Phrases — Provide a list of alternatives ("Understood", "Noted", "Alright") → prevents robotic repetition • Tool Preambles — Short sentence BEFORE the tool call so the user knows what's happening • Confirm Critical Data — For medical values always read back: "Tooth one-eight: three. Correct?" • Strict Domain Constraint — Exact scripted response for out-of-scope requests + escalation threshold
  21. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Security & Guardrails • No API-KEY in CLIENT! → Generate Ephemeral Token via Server • Clear definitions of what is allowed and what is not – System Prompt • Escalation-Tool: If there is uncertainty → Abort or handoff to a human agent • With gpt-live-1 the backend model is yours: guardrails, data access and logging live in your own agent, not in the voice layer • DSGVO: Check in which country the model is running (EU data residency via Azure OpenAI or Vertex AI regions)
  22. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Evaluation How do I know a prompt change actually made things better?
  23. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Evaluation Why it is hard • Two axes, not one — A response can be right and still sound broken. Grade both separately. • Content quality — Did it do the right thing? Correctness, tool choice, tool arguments, instruction following. • Audio quality — Did it sound acceptable? Naturalness, pacing, pronunciation, behaviour under noise. • A turn is a pipeline — speech start/stop, commit, response, audio deltas. Log every stage to find the real root cause. • Transcript is not ground truth — The audio signal is the truth; a transcript is just a model’s interpretation and can be wrong. • False fails & false passes — ASR drops a digit the model heard correctly, or a clean transcript hides clipped audio the model guessed at. • Grade on transcripts + traces — Run most automated grading on transcripts at scale; calibrate graders on noisy, production-like text. • Add an audio audit loop — Spot-check ~1–5% of sessions end-to-end by actually listening.
  24. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Evaluation Crawl / Walk / Run • Build complexity in steps — If your system cannot crawl, it will not run. Early evals must be diagnosable, repeatable and cheap. • Crawl — Synthetic (TTS) audio + single-turn. Tests intelligence: intent routing, tool choice, valid arguments. • Walk — Real noisy audio + single-turn. Tests perception: does it still hear "7pm", not "7" or "7:15"? • Run — Simulated user + multi-turn. Tests robustness: holds the goal, sequences tools, recovers from errors. • Example: "Change my reservation to 7pm" — Crawl grades the next turn only; Walk replays it under phone-bandwidth noise; Run adds messy follow-ups + an injected tool error. • Single-turn vs multi-turn — Single-turn = can you win the battle. Multi-turn = can you win the war (episode-level outcome).
  25. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Evaluation The loop: does a change actually help? • Dataset — Start with a gold seed set (10–50) of must-not-fail flows. Balance positives and negatives; tag by intent, audio, language, expected tool. • The iteration loop — Run evals, localise the failure to one behaviour, change ONE thing, rerun, confirm the fix improved it without regressions. • Graders — Layer them: deterministic (tool calls, JSON, patterns), LLM rubric (correctness, helpfulness), audio (silence, overlap, interruptions). • Regression suite — Hard cases you already fixed. Run on every prompt, model or tool change. Your "do not break" contract. • Harness — One job: make runs comparable. Pin audio bytes, chunking and VAD; prefer VAD off + manual commit for reproducibility. • Manual review is highest-leverage — Automation shows what you can measure; listening shows what you should measure. • Rolling discovery set — Fresh failures from production. Promote real failure modes into the offline dataset over time. • Holdout set — Untouched subset run occasionally. If test scores climb while holdout stays flat, you are training for the test.
  26. Chatless AI Architecting Realtime, Hands-Free AI Interfaces for Web &

    Mobile Summary • Realtime Voice ≠ Chatbot with microphone: It is a complete new interaction paradigm • Speech-to-Speech is the way for low-latency hands-free scenarios • UX is the real challenge – clear and precise communication of state to the human user is key • Cost is controllable – with caching, mini-models, truncation and short sessions • Start small – one scenario, one tool, one demo • And MOST IMPORTANT: The implementation needs to provide a REAL benefit for the user