Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Beyond the hype: Practical AI Applications for SRE

Avatar for Pete Hampton Pete Hampton
September 25, 2025

Beyond the hype: Practical AI Applications for SRE

Avatar for Pete Hampton

Pete Hampton

September 25, 2025

More Decks by Pete Hampton

Other Decks in Technology

Transcript

  1. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 2 What is ClickHouse?

    ClickHouse is an open-source , columnar OLAP database Designed for blazing fast , analytics of massive volumes of data Easily scalable to any size Processes data very fast Highly efficient storage Speaks SQL fluently
  2. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 5 Autonomous RCA is

    not there yet. “Use LLMs to assist investigations, summarize findings, draft updates, and suggest next steps while engineers stay in control with a fast…” “…shorten incidents and improve documentation when paired with a fast observability stack” (ClickHouse / ClickStack) The path forward is better context and better tools, with humans in control. See: https://clickhouse.com/blog/llm-observability-challenge
  3. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 6 2 agents LogHouse

    AI + CH Assist — This talk: Additional findings - LLMs are powerful co-pilots Building effective agents requires disciplined orchestration, context and tools Demonstrate building blazing fast and reliable observability agents on ClickHouse
  4. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 8 Lightning-fast data ingestion

    Enables closing the gap between event to action Analyzing logs across trillions of Internet requests Real-time observability for 300k+ customers Scales to 10s of trillions of samples 3 million messages per second 11 million rows ingested per second Up to billions of rows ingested per second Tesla
  5. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 14 Focus clustered in

    root dashboards per team. Now, our SRE root has 160 panels
  6. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 15 We need an

    intelligent HUD for observability. Enter LogHouse AI
  7. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 17 Building LogHouse AI

    • All of our docs are open and included in the base models. • We pay zero tokens for this • Fewer tokens = better results and less cost Accuracy ˮContext Rot: How Increasing Input Tokens Impacts LLM Performanceˮ research.trychroma.com/context-rot
  8. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 18 Building LogHouse AI

    • Despite good model intelligence, generated query jitter was an issue. • With low temperature, we still experience slight variations on the same query when asking the agent multiple times Accuracy
  9. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 19 Building LogHouse AI

    Accuracy SELECT toStartOfMinute(event_time) as minute, errorCodeToName(code) as error_name, count() as occurrences FROM loghouse.export_error_log WHERE spoken_name = 'plum-qa-31' AND event_time >= now() - INTERVAL 2 HOUR GROUP BY minute, error_name ORDER BY minute DESC, occurrences DESC SELECT toStartOfHour(event_time) as hour, level, count() as log_count FROM otel.server_text_log_0 WHERE Namespace = 'ns-plum-qa-31' AND EventDate >= today() AND event_time >= now() - INTERVAL 6 HOUR AND SeverityText IN ('ERROR', 'FATAL') GROUP BY hour, level ORDER BY hour DESC Query jitter for “Check the recent errors for plum-qa-31ˮ Same query → different table (errors vs logs), different time period
  10. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 20 Building LogHouse AI

    Raw inference of queries is slow, unreliable, prone to hallucination. Much better to provide dedicated tools. getExceptionsHistogram(), getK8sReadinessEvents(), getQueryLatencies(), getCPUWaitTimeseries(), getMergesRunning(), getSelectedBytesTimeseries(), getInsertedRowsTimeseries(), getPageCacheHitRateTimeseries(), getAutoscalerEvents(), getServerVersionDistribution() Accuracy
  11. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 21 Building LogHouse AI

    Inference of SQL can also be slow. Speed Reasoning… Reading autoscaler docs… Crafting sql… Lots of inference… User must check SQL__ SQL executes… Response Why did the instance scale up, then down, then up again? Reasoning… getAutoscalerEvents(...) sql executes… response
  12. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 22 Building LogHouse AI

    The Knowledge Flywheel Engineers can make tools easily 1 Many tools are created 2 Agent gets smarter 3
  13. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 24 Building LogHouse AI

    Knowledge Engineers making LLM tools is a solved problem We solved it years ago..
  14. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 28 CH Assist •

    An experimental fork of the xata.io SRE agent db agent • Will it just work? Yes - kind of
  15. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 29 Beyond the hype

    • Not the complete picture • Some solutions require creativity and innovation • More data/tools/prompting != better outcomes • Always find a problem in a large search space ◦ 😭 invent it ◦ 󰷺 overinflate a non issue
  16. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 31 12-factor agents HumanLayer)

    1. Natural language to Tool Calls 2. Own your prompts 3. Own your context window 4. Tools are just structured outputs 5. Unify execution state and business state 6. Launch/Pause/Resume with simple APIs 7. Contact humans with tool calls 8. Own your control Flow 9. Compact Errors into Context Window 10. Small, Focused agents 11. Trigger from anywhere, meet users where they are 12. Make your agent a stateless reducer See: https://github.com/humanlayer/12-factor-agents/tree/main Discovering best practices for developing AI agents
  17. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 32 Own your Prompts

    • Treat natural language as code • Build tests and evals - better consistency Best practices • Modular and composable prompts • Version control your prompts • Test prompts like you test code • Adding monitoring/logging for prompt performance
  18. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 33 Own your Context

    Window Curate what we put into the model(s) Best Practices • Strategic summarization and distillation • YNGI: Dynamic context hydration / retrieval • Intentional inclusion & Prioritized eviction
  19. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 34 Compact errors •

    Tools fail - LLMs typically recover • Limit retries Best Practices • Compress / Summarize the error - type & cause • Limit retries • Feed into context in structured form: eg. as a special event • Escalation / fallback: human-in-loop, default behaviour, or abort
  20. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 35 Small and focused

    agents Benefits • Manageable context - smaller context windows mean better perf • Clear responsibilities • Easier testing / Better reliability • Improved debugging Drawbacks • Complex observability • Sprawl
  21. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 36 Evaluating SRE agents

    General Agent Evals (LangChain AgentEvals, AWS agent-evaluation) → measure multi-step reasoning & tool use. What’s Needed for SRE Agent Evals • Simulated incident testbeds (logs, metrics, configs, outages) • Metrics: ◦ Time to detect, diagnose, remediate ◦ Correctness & relevance of actions ◦ Safety (avoiding harmful remediation) ◦ Efficiency (token/context use, retries) • Human-in-the-loop (HITL) approval Measuring efficacy / usefulness
  22. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 37 Challenges • Prompt

    Spaghetti → scattered or inconsistent (Factor 2: Own Prompts) • Context Overload → Logs, metrics, configs flood the model. (Factor 3: Own Your Context Window) • Plan for failures → Failures loop or get ignored unless compacted into context. (Factor 9: Compact Errors) • Evals → Important, open question 👉 SRE tasks naturally push against 12-Factor best practices — careful decomposition, context control, and error management enables improvement.
  23. ©2025 CLICKHOUSE INC., CONFIDENTIAL & PROPRIETARY 44 Thank you! https://clickhouse.ai

    https://clickhouse.com/blog/llm-observability-challenge