{"id":5053,"date":"2026-09-11T09:15:49","date_gmt":"2026-09-11T09:15:49","guid":{"rendered":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/09\/11\/githubs-august-outages-show-growth-is-outpacing-infrastructure\/"},"modified":"2026-09-11T09:15:49","modified_gmt":"2026-09-11T09:15:49","slug":"githubs-august-outages-show-growth-is-outpacing-infrastructure","status":"publish","type":"post","link":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/09\/11\/githubs-august-outages-show-growth-is-outpacing-infrastructure\/","title":{"rendered":"GitHub\u2019s August Outages Show Growth Is Outpacing Infrastructure"},"content":{"rendered":"<div><img data-opt-id=1261231762  fetchpriority=\"high\" decoding=\"async\" width=\"770\" height=\"330\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/09\/github-august-availability-770x330-1.jpg\" class=\"attachment-large size-large wp-post-image\" alt=\"\" \/><\/div>\n<p><img data-opt-id=2054570631  fetchpriority=\"high\" decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/09\/github-august-availability-770x330-1-150x150.jpg\" class=\"attachment-thumbnail size-thumbnail wp-post-image\" alt=\"\" \/><\/p>\n<p><span>Five incidents in one month is a lot for any platform, but it\u2019s especially notable when that platform hosts most of the world\u2019s code review, CI\/CD, and now AI coding agents. GitHub\u2019s August 2026 availability report reads less like a list of unrelated hiccups and more like a pattern: the company\u2019s own infrastructure is straining to keep up with how fast usage of Actions and Copilot has grown.<\/span><\/p>\n<p><span>The report opens with a line that doubles as GitHub\u2019s operating philosophy right now: \u201cAvailability, then capacity, then features.\u201d That ordering matters, because most of August\u2019s incidents trace back to capacity limits, not bugs in new features.<\/span><\/p>\n<p><span>The first hit came August 6, when a routine deployment reduced pod capacity in one datacenter. That triggered service mesh saturation and cascading failures across multiple clusters, knocking out Actions (both hosted and self-hosted runners), Copilot\u2019s coding agent, code review, Pages builds, Dependabot, and repository migrations for roughly nine hours. GitHub\u2019s own postmortem put it bluntly: \u201cThe affected actions services were running close to their capacity and concurrency limits.\u201d The fix was a rollback, expanded capacity, and a fix for a latent bug that had let runners pick up jobs they couldn\u2019t actually run.<\/span><\/p>\n<p><span>Eleven days later, on August 17, a traffic peak exceeded datacenter load balancer limits. A service-mesh sidecar failed to scale, network flow limits got exhausted, and the shared authentication path degraded. The blast radius was wide: issues, pull requests, both APIs, Actions, Copilot, authentication, and webhooks all took a hit. Peak front-door failure rate reached 56%, and about 29,000 organizations saw errors across roughly 4.8 million requests. GitHub also flagged a compounding factor familiar to anyone who has debugged a distributed system under load: \u201cA latent client retry bug sharply amplified traffic to one internal authentication endpoint.\u201d Retry storms make bad days worse, and this was no exception.<\/span><\/p>\n<p><span>August 20 brought a different kind of failure, tied to Copilot\u2019s cloud agent. An upstream managed database had a regional outage, and a storage configuration issue slowed the failover. Tasks kept running, but status and results visibility lagged for up to 90 minutes at more than 54 organizations over nearly 11 hours.<\/span><\/p>\n<p><span>Mitch Ashley, vice president and practice lead for CIO &amp; Technology Buyers and Software Lifecycle Engineering at <a href=\"https:\/\/futurumgroup.com\/\" target=\"_blank\" rel=\"noopener\">The Futurum Group<\/a>, said that the incident points to a gap most engineering teams haven\u2019t closed yet. \u201cOrdering availability ahead of features is the right call,\u201d Ashley said. \u201cThe Copilot agent incident showed why: tasks kept running while status lagged, so teams could not tell what had finished. Engineering leaders should require agent status checks they own, outside any vendor\u2019s UI.\u201d<\/span><\/p>\n<p><span>Then came back-to-back incidents on August 26 and 27. The first was a database saturation problem tied to Actions run starts, Pages deployments, and Copilot code review, with GitHub owning the underlying issue directly: \u201cOur shared infrastructure services have not kept up with our month-over-month actions growth and peak load.\u201d At least 24 organizations saw run-start failures, and 386 organizations felt some impact before manual throttling brought things back under control. The next day, a separate issue hit only Copilot requests routed to the Kimi K3 model, with 63% of those requests failing due to a problem at the upstream model provider. Other models were unaffected, but it\u2019s a useful data point on how much of Copilot\u2019s reliability now depends on providers GitHub doesn\u2019t fully control.<\/span><\/p>\n<p><span>For platform teams, the throughline across all five incidents is the same: growth in Actions usage and AI-assisted workflows is outpacing the shared infrastructure meant to support it.<\/span><\/p>\n<p><span>Ashley sees that as more than an infrastructure problem. \u201cGitHub is bidding to become the surface agents run on, which changes what availability means,\u201d he said. \u201cWith Actions and Copilot in the delivery path, an outage there is a production incident for every team downstream.\u201d<\/span><\/p>\n<p><span>GitHub isn\u2019t ignoring the pattern. Alongside the incident writeups, the company detailed parallel infrastructure work: moving MySQL primaries to Azure, cutting roughly a million queries per second from its database load, routing a third of Actions jobs to spare capacity, and expanding pull request isolation to authenticated reads. None of that is glamorous, and none of it will show up in a product announcement. But it\u2019s the plumbing that determines whether a platform this central to software delivery can keep scaling without buckling every few weeks.<\/span><\/p>\n<p><span>That\u2019s the real takeaway for engineering leaders who depend on GitHub as critical infrastructure, not just a code host. Actions and Copilot have become part of the software delivery pipeline itself, which means an outage there isn\u2019t an inconvenience \u2014 it\u2019s a production incident for every team whose builds, deployments, and reviews run through it. Teams that haven\u2019t already should be asking what happens to their pipeline when GitHub has a bad day: whether critical deployments have a manual fallback, whether long-running agent tasks have their own status checks outside GitHub\u2019s UI, and whether retry logic on their end could make an incident worse instead of better.<\/span><\/p>\n<p><span>GitHub deserves credit for its transparency here. Five detailed postmortems in one report, each naming a specific root cause, is more candor than most vendors offer. But candor doesn\u2019t replace capacity. The next few months of availability reports will say a lot about whether GitHub\u2019s infrastructure investments are keeping pace with a platform that\u2019s becoming more central to how software gets built.<\/span><\/p>\n<p><a href=\"https:\/\/devops.com\/githubs-august-outages-show-growth-is-outpacing-infrastructure\/\" target=\"_blank\" class=\"feedzy-rss-link-icon\">Read More<\/a><\/p>\n<p>\u200b<\/p>","protected":false},"excerpt":{"rendered":"<p>Five incidents in one month is a lot for any platform, but it\u2019s especially notable when that platform hosts most [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":5054,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[5],"tags":[],"class_list":["post-5053","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops"],"_links":{"self":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5053","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/comments?post=5053"}],"version-history":[{"count":0,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5053\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media\/5054"}],"wp:attachment":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media?parent=5053"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/categories?post=5053"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/tags?post=5053"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}