{"id":4893,"date":"2026-08-21T12:40:16","date_gmt":"2026-08-21T12:40:16","guid":{"rendered":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/08\/21\/production-grade-ai-eval-systems-what-i-learned-putting-llms-on-call\/"},"modified":"2026-08-21T12:40:16","modified_gmt":"2026-08-21T12:40:16","slug":"production-grade-ai-eval-systems-what-i-learned-putting-llms-on-call","status":"publish","type":"post","link":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/08\/21\/production-grade-ai-eval-systems-what-i-learned-putting-llms-on-call\/","title":{"rendered":"Production-Grade AI Eval Systems. What I Learned Putting LLMs on Call"},"content":{"rendered":"<div><img data-opt-id=305064770  fetchpriority=\"high\" decoding=\"async\" width=\"770\" height=\"330\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/07\/ai_eval_system_770x3302.jpg\" class=\"attachment-large size-large wp-post-image\" alt=\"\" \/><\/div>\n<p><img data-opt-id=341082178  fetchpriority=\"high\" decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/07\/ai_eval_system_770x3302-150x150.jpg\" class=\"attachment-thumbnail size-thumbnail wp-post-image\" alt=\"\" \/><\/p>\n<p>It was around 11 p.m. on a Thursday. Our AI support agent had been live for three weeks. Latency: Green. Error rate: Green. Then a customer DM landed in Slack with a screenshot \u2014 the bot had cheerfully invented a refund policy for a product we had never sold. Made up the SKU. Made up the rules. Returned a confident answer in 1.2 seconds.<\/p>\n<p>Every SRE metric said the system was healthy. The system was lying to customers at scale, and we had no signal. That night I started building what I now call a production-grade eval system. This is the version I wish someone had written for me before that Thursday.<\/p>\n<h3>The Honest Problem<\/h3>\n<p>For 15 years, the SRE playbook worked because systems were deterministic \u2014 same input and output. LLMs break that contract. Your vendor can silently push a new model checkpoint on a Wednesday, and your agent develops a new personality. A re-indexed retrieval store can send \u2018what\u2019s your return policy?\u2019 to a marketing blog post instead of the actual policy doc. None of those registers as a 4xx.<\/p>\n<p>The result: Dashboards all green; product quietly degrading. Everything below exists to close that gap.<\/p>\n<h3>The Three Places You Must Evaluate<\/h3>\n<p>Most teams vibes-check their AI feature: A PM tries six prompts, says \u201cfeels good\u201d and ships. There\u2019s no regression suite, no quality baseline and no rubric. The right fix is to evaluate in three explicit places:<\/p>\n<p><img data-opt-id=1074693373  data-opt-src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/07\/Picture2-18.png\"  decoding=\"async\" class=\"alignnone size-full wp-image-188460\" src=\"data:image/svg+xml,%3Csvg%20viewBox%3D%220%200%20100%%20100%%22%20width%3D%22100%%22%20height%3D%22100%%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%3Crect%20width%3D%22100%%22%20height%3D%22100%%22%20fill%3D%22transparent%22%2F%3E%3C%2Fsvg%3E\" alt=\"\" width=\"624\" height=\"194\" \/><\/p>\n<p>Most teams have Phase 1, sort of. Almost nobody has Phase 2. Phase 3 is where the screenshots-in-Slack live. If I could pick only one, I\u2019d start with Phase 3 \u2014 it catches failures you didn\u2019t predict. You need all three to call yourself production-grade.<\/p>\n<h3>The Four-Layer Evaluator Stack<\/h3>\n<p>There is no single quality metric. What works is layers \u2014 cheap checks at the bottom, expensive judges at the top \u2014 with sampling upward.<\/p>\n<p><img data-opt-id=1533102967  data-opt-src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/07\/Picture3-15.png\"  decoding=\"async\" class=\"alignnone size-full wp-image-188461\" src=\"data:image/svg+xml,%3Csvg%20viewBox%3D%220%200%20100%%20100%%22%20width%3D%22100%%22%20height%3D%22100%%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%3Crect%20width%3D%22100%%22%20height%3D%22100%%22%20fill%3D%22transparent%22%2F%3E%3C%2Fsvg%3E\" alt=\"\" width=\"624\" height=\"263\" \/><\/p>\n<p>Layer 1 (deterministic code) runs on every trace \u2014 free. Layers 2\u20133 run on a sampled subset; LLM judges cost real money and will eat your budget if you let them run inline. Layer 4 (domain checks) runs with humans in the loop \u2014 slowest, most valuable.<\/p>\n<h3>The Flywheel: Where Good Evals Come From<\/h3>\n<p>Nobody hands you a good eval set. Vendors sell generic benchmarks that score 95% on your system because they test nothing customers care about. Real evals come from your own production failures.<\/p>\n<ul>\n<li>Pull production traces \u2014 bias toward anomalies, low-confidence scores and customer complaints.<\/li>\n<li>One person labels them \u2014 consistency beats coverage. Two labelers create noise that looks like signal.<\/li>\n<li>Cluster failure modes \u2014 hallucination? tone? wrong tool call? Each mode needs its own judge.<\/li>\n<li>Promote failures to regression tests \u2014 the failure can never happen silently again.<\/li>\n<\/ul>\n<p>When you lack data, generate only inputs synthetically, then run your actual application on them. Never let an LLM generate both sides \u2014 you\u2019ll get a data set that scores 99% and means nothing.<\/p>\n<h3>RAG: the Three Metrics That Matter<\/h3>\n<p>If you\u2019re running retrieval-augmented generation \u2014 and most of you are \u2014 every interesting metric is a relationship between the question, the retrieved context and the generated answer.<\/p>\n<ul>\n<li>Retrieval Tier: Context precision and recall. If this is broken, nothing downstream matters. Start here.<\/li>\n<li>Core RAG Tier: Faithfulness (model stays grounded in context) and answer relevance (model answers the actual question).<\/li>\n<li>Diagnostic Tier: Citation accuracy, noise sensitivity, per-chunk hallucination rate. Run these on investigations, not alerts.<\/li>\n<\/ul>\n<p><img data-opt-id=1761074329  data-opt-src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/07\/Picture4-6.png\"  decoding=\"async\" class=\"alignnone size-full wp-image-188462\" src=\"data:image/svg+xml,%3Csvg%20viewBox%3D%220%200%20100%%20100%%22%20width%3D%22100%%22%20height%3D%22100%%22%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%3E%3Crect%20width%3D%22100%%22%20height%3D%22100%%22%20fill%3D%22transparent%22%2F%3E%3C%2Fsvg%3E\" alt=\"\" width=\"624\" height=\"291\" \/><\/p>\n<h3>Guardrails Vs. Evaluators<\/h3>\n<p>This confusion causes more bad architecture than anything else I see in production AI today.<\/p>\n<p>Keep them in separate code paths. Draw them in separate boxes on your architecture diagram. A guardrail that takes three seconds is a broken product. An evaluator that takes three seconds on 5% of traffic is healthy.<\/p>\n<h3>The Production-Readiness Checklist<\/h3>\n<p>If someone asks me whether their AI product is ready for production, I walk through this list:<\/p>\n<ul>\n<li>Golden regression data set \u2014 in the repo, version-controlled, owned by a named human.<\/li>\n<li>CI eval gate \u2014 blocks merges when quality drops on any PR touching prompts, retrievers or model config.<\/li>\n<li>Online evals on sampled production traces \u2014 at least Layers 1\u20132 on every span.<\/li>\n<li>LLM judges validated against humans (Cohen\u2019s kappa &gt;0.7), re-validated on every model upgrade.<\/li>\n<li>Weekly annotation session with a written agenda \u2014 not random sampling.<\/li>\n<li>Pinned model versions for production model and judge model. Vendor upgrades are deploys.<\/li>\n<li>Guardrails and evaluators in separate code paths.<\/li>\n<li>Someone whose job is to turn findings into roadmap items.<\/li>\n<\/ul>\n<h3>The Mindset Shift<\/h3>\n<p><a href=\"https:\/\/devops.com\/production-grade-ai-eval-systems\/devops.com\/scaling-ai-the-right-way\">Reliability for AI<\/a> is not measured in uptime, it\u2019s measured in quality of output over time. The dashboards your SRE team already built are necessary but no longer sufficient.<\/p>\n<p>The new layer is the eval system \u2014 multi-tier, fed by an error analysis loop, wired into CI and live traffic, with calibrated human-in-the-loop judges. Build it, and AI stops being the system your team is scared to put on the on-call rotation.<\/p>\n<p><em>That Thursday night was almost a year ago. We haven\u2019t had another one like it. The dashboards are still green \u2014 and now I know what green actually means.<\/em><\/p>\n<p><a href=\"https:\/\/devops.com\/production-grade-ai-eval-systems\/\" target=\"_blank\" class=\"feedzy-rss-link-icon\">Read More<\/a><\/p>\n<p>\u200b<\/p>","protected":false},"excerpt":{"rendered":"<p>It was around 11 p.m. on a Thursday. Our AI support agent had been live for three weeks. Latency: Green. [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4894,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[5],"tags":[],"class_list":["post-4893","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops"],"_links":{"self":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/4893","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/comments?post=4893"}],"version-history":[{"count":0,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/4893\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media\/4894"}],"wp:attachment":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media?parent=4893"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/categories?post=4893"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/tags?post=4893"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}