ClickHouse is an open-source , columnar OLAP database Designed for blazing fast , analytics of massive volumes of data Easily scalable to any size Processes data very fast Highly efficient storage Speaks SQL fluently
not there yet. “Use LLMs to assist investigations, summarize findings, draft updates, and suggest next steps while engineers stay in control with a fast…” “…shorten incidents and improve documentation when paired with a fast observability stack” (ClickHouse / ClickStack) The path forward is better context and better tools, with humans in control. See: https://clickhouse.com/blog/llm-observability-challenge
AI + CH Assist — This talk: Additional findings - LLMs are powerful co-pilots Building effective agents requires disciplined orchestration, context and tools Demonstrate building blazing fast and reliable observability agents on ClickHouse
Enables closing the gap between event to action Analyzing logs across trillions of Internet requests Real-time observability for 300k+ customers Scales to 10s of trillions of samples 3 million messages per second 11 million rows ingested per second Up to billions of rows ingested per second Tesla
• All of our docs are open and included in the base models. • We pay zero tokens for this • Fewer tokens = better results and less cost Accuracy ˮContext Rot: How Increasing Input Tokens Impacts LLM Performanceˮ research.trychroma.com/context-rot
• Despite good model intelligence, generated query jitter was an issue. • With low temperature, we still experience slight variations on the same query when asking the agent multiple times Accuracy
Accuracy SELECT toStartOfMinute(event_time) as minute, errorCodeToName(code) as error_name, count() as occurrences FROM loghouse.export_error_log WHERE spoken_name = 'plum-qa-31' AND event_time >= now() - INTERVAL 2 HOUR GROUP BY minute, error_name ORDER BY minute DESC, occurrences DESC SELECT toStartOfHour(event_time) as hour, level, count() as log_count FROM otel.server_text_log_0 WHERE Namespace = 'ns-plum-qa-31' AND EventDate >= today() AND event_time >= now() - INTERVAL 6 HOUR AND SeverityText IN ('ERROR', 'FATAL') GROUP BY hour, level ORDER BY hour DESC Query jitter for “Check the recent errors for plum-qa-31ˮ Same query → different table (errors vs logs), different time period
Raw inference of queries is slow, unreliable, prone to hallucination. Much better to provide dedicated tools. getExceptionsHistogram(), getK8sReadinessEvents(), getQueryLatencies(), getCPUWaitTimeseries(), getMergesRunning(), getSelectedBytesTimeseries(), getInsertedRowsTimeseries(), getPageCacheHitRateTimeseries(), getAutoscalerEvents(), getServerVersionDistribution() Accuracy
Inference of SQL can also be slow. Speed Reasoning… Reading autoscaler docs… Crafting sql… Lots of inference… User must check SQL__ SQL executes… Response Why did the instance scale up, then down, then up again? Reasoning… getAutoscalerEvents(...) sql executes… response
• Not the complete picture • Some solutions require creativity and innovation • More data/tools/prompting != better outcomes • Always find a problem in a large search space ◦ 😭 invent it ◦ overinflate a non issue
1. Natural language to Tool Calls 2. Own your prompts 3. Own your context window 4. Tools are just structured outputs 5. Unify execution state and business state 6. Launch/Pause/Resume with simple APIs 7. Contact humans with tool calls 8. Own your control Flow 9. Compact Errors into Context Window 10. Small, Focused agents 11. Trigger from anywhere, meet users where they are 12. Make your agent a stateless reducer See: https://github.com/humanlayer/12-factor-agents/tree/main Discovering best practices for developing AI agents
• Treat natural language as code • Build tests and evals - better consistency Best practices • Modular and composable prompts • Version control your prompts • Test prompts like you test code • Adding monitoring/logging for prompt performance
Window Curate what we put into the model(s) Best Practices • Strategic summarization and distillation • YNGI: Dynamic context hydration / retrieval • Intentional inclusion & Prioritized eviction
Tools fail - LLMs typically recover • Limit retries Best Practices • Compress / Summarize the error - type & cause • Limit retries • Feed into context in structured form: eg. as a special event • Escalation / fallback: human-in-loop, default behaviour, or abort
Spaghetti → scattered or inconsistent (Factor 2: Own Prompts) • Context Overload → Logs, metrics, configs flood the model. (Factor 3: Own Your Context Window) • Plan for failures → Failures loop or get ignored unless compacted into context. (Factor 9: Compact Errors) • Evals → Important, open question 👉 SRE tasks naturally push against 12-Factor best practices — careful decomposition, context control, and error management enables improvement.