
This is a guest conversation episode of Ship It Weekly, separate from the weekly news recaps.In this Ship It Conversations episode, I talk with Mat Ryer of Grafana Labs about AI observability, production agents, evals, telemetry cost, guardrails, and what changes once AI moves beyond demos and into systems teams actually depend on.Mat is Senior Director of AI at Grafana Labs, where he focuses on how AI fits into observability and production systems.We talk about Grafana Assistant, why AI observability is not just logs, latency, and HTTP 200s, and how teams can measure whether agents are actually helping. Mat gets into evals, LLM-as-judge patterns, traces as a way to think about conversations, user feedback, tool choice, model changes, and the cost of collecting new telemetry.We also dig into UX and trust. If an AI assistant gives you a wall of text, you still have to decide whether to believe it. If it can show the graph, deep link into Grafana, apply filters, and expose the source data, that becomes a much more useful operating experience.The big takeaway: start small, enhance workflows you already have, build feedback loops, and treat production AI like something you actually have to operate.Highlights• Why AI demos are easy, but production AI is harder• Why agents need observability, evals, and guardrails• How Grafana thinks about AI Assistant and AI observability• Where LLM-as-judge patterns, traces, tool calls, and feedback fit• Why telemetry cost problems may repeat with AI workloads• Why UX matters when operators need to trust the answer• Where AI can help SRE and platform teams todayLinksMat Ryer on LinkedIn: https://www.linkedin.com/in/matryer/Mat Ryer on GitHub: https://github.com/matryerGrafana Labs: https://grafana.comGrafana Assistant: https://grafana.com/products/cloud/ai-assistant/Grafana AI Observability: https://grafana.com/docs/grafana-cloud/machine-learning/ai-observability/Grafana Adaptive Telemetry: https://grafana.com/products/cloud/adaptive-telemetry/Grafana MCP server: https://github.com/grafana/mcp-grafanaOpenTelemetry: https://opentelemetry.ioPrometheus: https://prometheus.ioGrafana Loki: https://grafana.com/docs/loki/latest/More episodes and show notes: https://shipitweekly.fmOn Call Brief: https://oncallbrief.com
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

Telstra’s Time Sync Outage, DoorDash’s 1.5M RPS Cache, Stateless MCP, GitHub and PyPI Supply Chain Delays, and Why the Quietest Infrastructure Often Has the Biggest Blast Radius

Ship It Conversations: Jay Lark of Hookbridge on Webhook Reliability, Retries, Idempotency, Replay, HMAC Security, Local Testing, and Operating Webhooks in Production

AWS CloudFormation Express Mode, Spark 4.2 Vector Search, AI Speeds Coding but Not Delivery, Platform Governance Developers Won’t Hate, and Why Could Is Not the Same as Should

GitHub API Enumeration, Grok Build CLI Data Exposure, AWS Security Hub Network Scanning, AI-Powered Patch Pressure, and Why Visibility Is Not Ownership
Free AI-powered recaps of Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.