{"id":5233,"date":"2026-10-06T10:13:54","date_gmt":"2026-10-06T10:13:54","guid":{"rendered":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/06\/availability-should-follow-the-workload-not-the-infrastructure\/"},"modified":"2026-10-06T10:13:54","modified_gmt":"2026-10-06T10:13:54","slug":"availability-should-follow-the-workload-not-the-infrastructure","status":"publish","type":"post","link":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/06\/availability-should-follow-the-workload-not-the-infrastructure\/","title":{"rendered":"Availability Should Follow the Workload, Not the Infrastructure"},"content":{"rendered":"<div><img data-opt-id=1655035192  fetchpriority=\"high\" decoding=\"async\" width=\"770\" height=\"330\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/intelligent-operations-platform-engineering-770x330-1.jpg\" class=\"attachment-large size-large wp-post-image\" alt=\"\" \/><\/div>\n<p><img data-opt-id=1669028304  fetchpriority=\"high\" decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/intelligent-operations-platform-engineering-770x330-1-150x150.jpg\" class=\"attachment-thumbnail size-thumbnail wp-post-image\" alt=\"\" \/><\/p>\n<p>I find that I cannot stop asking myself this question\u2026 How did we automate almost everything in software delivery, yet when production fails, we still throw humans at the problem?<\/p>\n<p>Think about what <a href=\"https:\/\/devops.com\/platform-engineering-vs-devops-why-this-is-the-wrong-question\/\" target=\"_blank\" rel=\"noopener\">modern platform engineering<\/a> has accomplished. We provision infrastructure with code. CI\/CD? We build pipelines that can deploy hundreds of times a day. We use GitOps to keep environments consistent. Workloads? Ours can automatically scale based on demand. Kubernetes? We can spin up an entire Kubernetes cluster in minutes. That is incredible progress.<\/p>\n<p>Then a production database fails at 2:13 a.m.<\/p>\n<p>The Slack channel lights up. Someone starts a bridge call. Someone else starts digging through a runbook that hasn\u2019t been updated in a year. And everyone waits for the one engineer who \u201cknows how this thing works.\u201d<\/p>\n<p>How is that still our operating model?<\/p>\n<p>We\u2019ve spent the last decade teaching our platforms how to deploy software. We haven\u2019t spent nearly as much time teaching them how to respond when production breaks. That\u2019s the next maturity curve for platform engineering \u2013 not more deployment automation, but intelligent operational automation.<\/p>\n<p>One of the biggest misconceptions I see? Organizations that equate modern infrastructure with resilient infrastructure. They\u2019re not the same thing, not at all. Kubernetes is fantastic at orchestrating workloads. Better than any platform before, it can restart a pod, replace a failed node, and reconcile desired state.<\/p>\n<p>But Kubernetes doesn\u2019t know whether your customers can still place an order, process a payment, or complete a transaction. That\u2019s not a knock on Kubernetes. It\u2019s simply not its job. The problem is we\u2019ve started treating orchestration as if it were a complete reliability strategy. It isn\u2019t.<\/p>\n<p>Orchestration answers the question, \u201cWhere should this workload run?\u201d<\/p>\n<p>Reliability answers a different question: \u201cHow do I keep this workload available when something inevitably fails?\u201d<\/p>\n<p>These are not the same problem. In modern environments, availability shouldn\u2019t be tied to a server, a cluster, or even a cloud. It should follow the workload.<\/p>\n<p>Here\u2019s the question every platform team should ask. When I speak with engineering teams, I tell them to forget buzzwords for a minute. Forget \u201ccloud native.\u201d Forget \u201cself-healing.\u201d Forget \u201chigh availability.\u201d Instead, ask one simple question, \u201cIf something breaks in production, what should happen next \u2013 and how much of it still depends on people?\u201d<\/p>\n<p>If the honest answer is, \u201cWell\u2026someone gets paged, we jump on a call, and then we figure it out,\u201d your platform isn\u2019t as automated as you think it is. It\u2019s not a criticism. It\u2019s an opportunity.<\/p>\n<p>The goal isn\u2019t to remove engineers from the process. It\u2019s to eliminate the routine operational decisions they\u2019ve already made hundreds of times before. Engineers should spend their time solving new problems \u2013 not repeating the same recovery steps every time a server, node, or workload fails.<\/p>\n<h3><strong>So\u2026What Should You Do Differently on Monday?<\/strong><\/h3>\n<p>Here\u2019s where I\u2019d start.<\/p>\n<p>Find your 2 a.m. processes.<\/p>\n<p>Make a list of everything your team does manually during a production incident. Not deployments. Not upgrades. Failures. If only one person knows how to recover a critical workload, that\u2019s not expertise \u2013 it\u2019s technical debt. Recovery shouldn\u2019t live inside someone\u2019s head. It should be built into the platform.<\/p>\n<p>Measure application recovery \u2013 not infrastructure recovery.<\/p>\n<p>Most teams know exactly how long it takes to deploy a release. Far fewer know how long it takes for a critical application to recover from a real failure. Your users don\u2019t care how quickly a pod restarted. They care how quickly they can get back to work. Measure what they experience. Better yet, start measuring how many recovery decisions still require human intervention. That may be the best indicator of platform maturity you have.<\/p>\n<p>Test recovery as often as you test deployments.<\/p>\n<p>Most organizations exercise their CI\/CD pipelines constantly. In a controlled way, how often do you deliberately break production to see what happens? Is the answer, \u201cAlmost never?\u201d \u2013 then I think you have found your next engineering project.<\/p>\n<p>Stop assuming Kubernetes solved everything.<\/p>\n<p>Kubernetes is one of the best orchestration platforms ever built. It\u2019s also just one layer of your platform. Stateful applications, databases, storage, networking, and application dependencies all have their own failure modes. Make sure your operational strategy accounts for them. Modern platforms aren\u2019t resilient simply because they run on Kubernetes. They\u2019re resilient because every layer \u2013 from infrastructure to the application itself \u2013 is designed to recover intelligently when something breaks.<\/p>\n<p>Eliminate decisions \u2013 not just steps.<\/p>\n<p>The goal isn\u2019t to make your outage runbook shorter. The goal is to make it unnecessary. The best platforms don\u2019t just automate tasks \u2013 they automate decisions that have already been defined through policy and testing. That\u2019s how you reduce downtime and reduce stress on your engineers at the same time.Here\u2019s a simple test. If your most experienced engineer took two weeks off tomorrow, would your recovery process work exactly the same way? If the answer is \u201cprobably,\u201d your platform still depends on tribal knowledge. Mature platforms don\u2019t just automate tasks \u2013 they automate operational decisions.<\/p>\n<h3><strong>The Next Frontier Isn\u2019t Deployment. It\u2019s Intelligent Operations<\/strong><\/h3>\n<p>For years, we\u2019ve measured platform engineering by one question: \u201cHow fast can we deploy?\u201d I think there\u2019s a better question for the decade ahead: \u201cHow intelligently can we recover?\u201d<\/p>\n<p>The answer has less to do with where an application runs than with how easily it can keep running when conditions change. Modern enterprises aren\u2019t choosing one environment over another. They\u2019re operating across on-premises infrastructure, virtual machines, Kubernetes clusters, public clouds, and edge environments \u2013 all at the same time. Reliability can no longer be tied to a particular server, cluster, or cloud. It has to follow the workload.<\/p>\n<p>The best platform teams won\u2019t be the ones that deploy the fastest. They\u2019ll be the ones whose applications remain available no matter where those workloads are running \u2013 or where they need to move next.<\/p>\n<p>That\u2019s the next evolution of platform engineering: Building platforms where operational intelligence, not infrastructure, determines availability.<\/p>\n<p><a href=\"https:\/\/devops.com\/availability-should-follow-the-workload-not-the-infrastructure\/\" target=\"_blank\" class=\"feedzy-rss-link-icon\">Read More<\/a><\/p>\n<p>\u200b<\/p>","protected":false},"excerpt":{"rendered":"<p>I find that I cannot stop asking myself this question\u2026 How did we automate almost everything in software delivery, yet [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":5234,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[5],"tags":[],"class_list":["post-5233","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops"],"_links":{"self":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5233","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/comments?post=5233"}],"version-history":[{"count":0,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5233\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media\/5234"}],"wp:attachment":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media?parent=5233"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/categories?post=5233"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/tags?post=5233"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}