Slide 1

Slide 1 text

One outage. Two LLM Features. Two outcomes.

Slide 2

Slide 2 text

Feb. 5th, 2026 Two LLM features hit the same outage. FEATURE A ~5min to decide: no action needed FEATURE B ~3hrs down in production

Slide 3

Slide 3 text

How Timee Delivers Day 1 Production Ready LLM Features June 10, 2026 Tomoyuki Saito

Slide 4

Slide 4 text

What happens when an LLM feature fails in production? What we learned. Why production readiness has to be default by Day 1.

Slide 5

Slide 5 text

MLOps Engineer · Timee Tomoyuki Saito Platform engineering for ML and LLM workloads Observability and production readiness Bird Watching

Slide 6

Slide 6 text

No content

Slide 7

Slide 7 text

No content

Slide 8

Slide 8 text

LLMs are becoming part of the product. Worker experience Better discovery, clearer information, and significantly less friction in finding the perfect next shift. Client experience Generation tools that help clients describe shifts faster, more accurately, and with higher conversion rates. Safety & Reliability Helping the platform be safer and more trustworthy across both sides through autonomous monitoring.

Slide 9

Slide 9 text

01 Two LLM features in production.

Slide 10

Slide 10 text

Designed for failure from day one. Feature A Asynchronous workflow Retries and failure handling Fallback path

Slide 11

Slide 11 text

Feature A

Slide 12

Slide 12 text

Feature B

Slide 13

Slide 13 text

Timee’s engineering organization API Call Deploy Deploy Platform Team (Platform Engineer) Stream Aligned Team (Product Engineer) Platform Team (MLOps Engineer) Complicated Subsystem Team (Data Scientist) ECS Cloud Run

Slide 14

Slide 14 text

Timee’s engineering organization at LLM era API Call Deploy Deploy Platform Team (Platform Engineer) Stream Aligned Team (Product Engineer) Platform Team (MLOps Engineer) Complicated Subsystem Team (Data Scientist) ECS Cloud Run

Slide 15

Slide 15 text

One team’s success Accessible to everyone

Slide 16

Slide 16 text

02 Production Readiness Checklist for LLM Application

Slide 17

Slide 17 text

A checklist for LLM applications. GENERAL ENGINEERING Stability & Reliability CI/CD automation, Release process Scalability Auto-scaling capacity planning Fault Tolerance Blast radius Control, Runbooks Standard Monitoring System metrics, Structured logging Security Access control, Encryption LLM-SPECIFIC ADDITIONS Prompt Management Version control & regression testing Token & Cost Control TPM/RPM rate limiting, Cost budgeting Resilience Model fallback, Circuit breakers Guardrails & Safety PII masking, Hallucination filters LLM Observability TTFT tracking, Response tracing Continuous QA Golden datasets, Offline evaluation

Slide 18

Slide 18 text

The Rails feature moved from PoC to production STEP 01 PoC Job description generation — a natural LLM use case. → STEP 02 Strong ROI Clear business impact. Direct user-facing value. → STEP 03 Production Strong product signal. The feature shipped. Note: Operationally, readiness wasn't at the same level as the first Cloud Run feature.

Slide 19

Slide 19 text

A checklist defines the expected state. It doesn't execute itself. PROBLEM 01 · DOC GAP PROBLEM 02 · HUMAN BOTTLENECK

Slide 20

Slide 20 text

03 LLM Gateway

Slide 21

Slide 21 text

20+ autonomous squads. 3-person MLOps team. PRODUCT SQUADS 01 02 03 04 05 06 07 08 09 10 11 12 13 14 15 16 17 18 19 20 MLOps +

Slide 22

Slide 22 text

20+ autonomous squads. 3-person MLOps team. PRODUCT SQUADS 01 02 03 04 05 06 07 08 09 10 11 12 13 14 15 16 17 18 19 20 MLOps + "We want to use LLMs"

Slide 23

Slide 23 text

ARCHITECTURE DECISION LLM Gateway as common entry point. One internal HTTP endpoint. Every team. Every language. Every model.

Slide 24

Slide 24 text

Risk vs Unlocking capabilities THREE STRATEGIC GOALS TO BALANCE THE ACKNOWLEDGED RISK Single Point of Failure THE STRATEGIC REASONING Observable and Controllable 01 02 03 Agility Reliability Governance VS

Slide 25

Slide 25 text

The direction was right. The rollout was incomplete. JANUARY 2026 First Gateway release FEBRUARY 5, 2026 The outage

Slide 26

Slide 26 text

February 5, 2026: The Outage Hits RESPONSE · CLAUDE.VERTEXAI.GOOGLEAPIS.COM 429 rate-limited HTTP/2 429 Too Many Requests

Slide 27

Slide 27 text

Visible LLM fallback DATADOG SCREENSHOT Cloud Run Request Error Count Total LLM Token Consumption Cloud Run Success Rate LLM Token Consumption Trend Feature A - Cloud Run/Python

Slide 28

Slide 28 text

The Cost of Zero Visibility CRITICAL QUESTIONS WE COULDN'T ANSWER ? When will the LLM recover? ? How many requests are actually failing? ? Should we wait, switch models, or turn off the feature? RESULT: 3 HOURS OF DOWNTIME IN PRODUCTION Feature B - Rails/ECS

Slide 29

Slide 29 text

Same outage. Same provider. Same company. Different structure. Not a people problem. A structure problem.

Slide 30

Slide 30 text

Three good pieces. One missing connection. The pieces existed. But they weren't connected into a single shared path. 01 PIECE 1 Checklist Defined the expected state. 02 PIECE 2 Datadog Provided evidence — in some places. 03 PIECE 3 Gateway Started — not yet shared.

Slide 31

Slide 31 text

THE LESSON Production readiness cannot stay as a checklist. It has to be built into the path teams naturally use.

Slide 32

Slide 32 text

04 What the Gateway brings ー Today

Slide 33

Slide 33 text

All LLM calls go through one common path. CALLERS ECS・Rails Job description gen Cloud Run · Python Job-flow LLM feature Batch · Python Offline / scheduled New use cases PoCs & future stacks LLM GATEWAY LiteLLM POST /v1/chat/completions MODELS PRIMARY Claude · Vertex AI FALLBACK Gemini EXPANSION Future hosted models LATER Open-weight + vLLM

Slide 34

Slide 34 text

Inside Gateway Pillar 1 Routing & fallback

Slide 35

Slide 35 text

Inside Gateway Pillar 2 Identity & tagging

Slide 36

Slide 36 text

Inside Gateway Pillar 3 Observability by Default

Slide 37

Slide 37 text

Inside Gateway Pillar 4 Governance & safety

Slide 38

Slide 38 text

01 Routing & fallback 02 Identity & tagging 03 Observability by Default 04 Governance & safety What every call gets, automatically

Slide 39

Slide 39 text

Current Gateway Architecture Overview

Slide 40

Slide 40 text

Production readiness becomes something the platform helps deliver. CHECKLIST Defines the expected state What "ready" looks like. MONITORING Proves what is happening Evidence of readiness in the moment. GATEWAY Defaults into the path teams use Brings evidence to every call.

Slide 41

Slide 41 text

The platform makes sure these don't depend on chance. STILL OWNED BY TEAMS Product UX and copy Business-specific fallback behavior Domain-aware prompt design Quality bar for their feature NO LONGER LEFT TO CHANCE ✕ No observability ✕ No cost visibility ✕ No clear fallback path ✕ No team-by-team variance

Slide 42

Slide 42 text

Four things, every team, by default. 01 Faster incident decisions 02 Cross-cloud tracing 03 Cost visibility from day one 04 Lower instrumentation burden

Slide 43

Slide 43 text

05 What's next.

Slide 44

Slide 44 text

From one external API to part of our infrastructure. 01 More models More hosted models. More open-weight models we serve ourselves with vLLM. 02 Agent workflows Tool calls, multi-step reasoning, internal tool servers. The call graph gets harder to reason about. 03 MCP-connected systems External context, internal tools, distributed planning. End-to-end tracing becomes essential.

Slide 45

Slide 45 text

Complex LLMs Demand Smarter Monitoring

Slide 46

Slide 46 text

For your own teams, regardless of stack. 01 A checklist alone is not enough. If it doesn't show up in the path teams use to ship, expect inconsistent outcomes the next time something breaks. 02 Pick one common path before you have five. Cost, fallback, and observability are far harder to retrofit than to standardize. If you have three teams today, that is the best time. 03 SDK gaps are an argument for the gateway pattern. Don't treat them as a blocker. Solve it once at the network boundary — every language benefits at the same time.

Slide 47

Slide 47 text

One outage. Two LLM Features. Two outcomes.

Slide 48

Slide 48 text

One platform. Shared by every product teams. One default. Production-ready from day one. One outcome. No more two stories inside one company. That is Day 1 Production-Ready.

Slide 49

Slide 49 text

Thank you! @tomoppi_31 Timee Corp Page