Building Git infrastructure for agent-scale development

Every day on GitHub, millions of developers build the products their customers rely on, contribute to open source, and pursue personal projects. GitHub’s architecture has changed steadily over the years to support that work and the growing demands of the developers and organizations who depend on it.

Agentic software development is driving the next architectural shift. With developers and agents working concurrently in repositories that receive millions of commits a day, these workloads demand a different Git architecture. We’re rebuilding GitHub’s Git infrastructure to support them. This post explores the demands shaping that work and the design principles behind it.

Today’s highest-volume workloads show the scale we’re building for. The gap between a typical repository and the busiest ones is wider than most people expect. Here’s a rough picture of the monthly repository activity distribution on GitHub from August 2026:

Repository activity climbs sharply at the far end of the distribution. The busiest repository on GitHub saw roughly a billion requests in August.

Beyond these highest-volume workloads, total Git activity on GitHub is also growing rapidly: between September 2025 and August 2026, it increased to more than 2x its previous level, from 218.2 billion events per month to 473.3 billion.

In September alone, developers and agents made 7.38 billion commits on GitHub, more than five times as many as a year earlier.

The repositories at the top of this curve show what agentic development looks like at its leading edge: large engineering teams running busy CI pipelines alongside growing fleets of agents. Supporting these teams means building Git infrastructure for sustained, concurrent reads and writes at a scale few repositories reach today. We’re investing deeply in Git infrastructure to meet the demands of agentic software development and give teams a foundation built for their most ambitious workloads.

Building for this scale means addressing several architectural challenges:

  • Commit turnaround becomes a bottleneck per agent. An agent in a tight loop commits or checkpoints after nearly every action. Its speed is bounded by how fast a single push completes, so latency that a human would never notice becomes the limiting factor.
  • Write throughput demand is increasing by orders of magnitude. Pushes grew 4.9x year over year, from 0.69 billion to 3.35 billion per month. Thousands of agents working on their own branches in one repository produce a sustained write rate that converges on a single point in our architecture.
  • Merges contend on one reference. Trunk-based development, release trains, and merge queues funnel all that work onto a single ref that has to absorb every merge. Pull request merges on GitHub grew to nearly 4x their volume a year ago.
  • Each push multiplies into thousands of reads. For example, CI and code scanning clone or fetch the same branch tip thousands of times per minute, and that fan-out has to be cheap. GitHub Actions alone ran 3.26 billion times in September, more than 4x as many as a year ago.
  • Operations on a repository must continue to be fast. To keep them fast, we continually compact repository data and clean up objects that are no longer needed. Every new write adds to that work, and the cost compounds as volume climbs.

This is why fast clones only solve part of the problem. Reads are relatively easy to scale: add caches, add replicas, and serve the same bytes to more clients. Scaling reads is essential, but these workloads require more than that. Writes are way harder. Every push has to be stored durably and made visible consistently before the next agent or CI job can build on it.

Where today’s architecture meets new demands

The current architecture has served developers well for years. Every repository is stored by Spokes, which keeps a full copy on the local disks of several fileservers, five by default. Those fast local disks let Git operations read native repository data with low latency, and the extra copies provide redundancy while spreading reads across fileservers. When a push updates a reference, a three-phase commit protocol uses a quorum to ensure that CI, the web UI, and API clients see a consistent repository state. That pairing serves a billion repositories today.

However, the mechanism we use for durability is the same one we use for scale. The copies on disk are the source of truth, so adding read capacity means adding another durable replica. Every replica participates in every write, so a push is only as fast as the slowest replica in its set. The net effect: adding replicas to absorb read load makes writes slower.

For most repositories, this tradeoff works well. At the highest activity levels, it becomes a ceiling: adding read replicas adds overhead to writes, losing a replica reduces read capacity, and losing quorum stops writes entirely. To meet agent-first demands, we need to separate durability from scale without losing what teams rely on today.

Built for the busiest, better for everyone

We’re rebuilding the infrastructure while GitHub keeps running. There’s no maintenance window where the world’s code stops moving, and no version of this work where we ask people to change how they build software while we do it.

We’re building for the most demanding workloads on GitHub: an enterprise shipping under strict regulatory requirements, a team landing a change across a repo that builds an operating system, and an organization running thousands of agents against a single codebase. Engineering for that scale raises the floor for everyone. The maintainer reviewing contributions from volunteers across time zones and the student opening their first pull request get the same faster, more resilient foundation.

The new architecture also must preserve the controls teams already operate on. A maintainer needs branch protections and required reviews so an unreviewed change never reaches the default branch. A security team needs audit logs and repository visibility to investigate a suspicious access event. An on-call engineer needs dependable automation and enough observability to understand why a deployment failed.

For the platform to keep serving everyone here while it scales for the busiest workloads, these are our guiding principles:

  • Build on the workflows developers already trust. Teams rely on workflows like branching, review, merge, and history to build, ship, and govern software at scale. Our new infrastructure is designed to support those same workflows at much higher volumes of activity.
  • Put reliability first. We measure every decision against the reliability that developers and organizations require. Confidence in the platform is what lets an engineering organization build automation, ship on a schedule, meet compliance obligations, and understand the software it produces. This work will meaningfully improve throughput and scale, and those gains extend a foundation of trust that’s already there.
  • Keep people in control of their code. If the system isn’t helping the people and organizations who use it, and isn’t under their control, it isn’t worth building. As agents take on more of the work, the people who own the code can still review, understand, and approve it.

The approach

We are building a new GitHub architecture that can scale much more effectively. Our approach centers around core distributed systems design tenets, applied to the concurrency and scale of agentic software development. Our goal is to continue the forward momentum for open-source communities and enterprises around the world who have built their projects with Git and GitHub, while preserving and adapting the features and controls around it to meet the new needs of the agentic era.

Minimize coordination

A repository that receives many pushes must accept and publish updates quickly. Coordination is valuable when it protects correctness, but too much of it limits write throughput and can turn a busy repository into a bottleneck. Our current architecture is tightly coupled in places it doesn’t need to be, which limits our ability to scale across reads and writes without tough trade-offs. We’re redesigning the system to preserve the coordination that Git semantics require and let everything else proceed independently.

  • Coordinate only what needs agreement. The part of a push that truly needs agreement is the reference update itself. Storing the underlying objects, validating object connectivity, and secret scanning are much more work, but most of it can happen in parallel to other writes. That shrinks the critical path of a push to the small step that needs coordination, so the rest of the work no longer delays the acknowledgment.
  • Move maintenance off the serving path. Compaction and garbage collection are among the heaviest work a repository does, and today they run on the same hosts that answer live Git requests. In the new architecture, separate workers handle maintenance directly against durable storage. A busy repository can be optimized continuously in the background without slowing pushes and fetches.

Decouple storage from compute

Today, complete repository copies on local disks serve both as durable storage and as the layer that answers Git requests. Separating the two lets us scale each one independently.

  • Scale reads without adding durable copies. In the new architecture, read capacity comes from lightweight workers that cache data to serve requests. The authoritative copy of the repository lives in a durable storage layer underneath. That way, the platform can absorb large read spikes from CI fan-out, agent fleets, and large clones without adding work to every push.
  • Let each layer do one job. Authoritative repository data lives in Azure Blob Storage, which already provides durability and replication at Azure scale. The compute layer is optimized for throughput at the lowest latency.
  • Recover faster from failures. When storage and compute are coupled, losing a host reduces both capacity and durability, and recovery means rebuilding a full repository copy. When they’re separate, losing a compute worker is closer to a cache miss: a replacement worker can start serving requests right away and fill its cache from durable storage as traffic arrives.
  • Match capacity to demand. Compute workers can be added or removed as traffic changes instead of provisioning for peak load in advance. A repository going through a burst of activity, like a release or a new agent fleet coming online, can get extra capacity for the burst. Once it passes, that capacity goes away.

Together, these tenets allow us to support higher throughput and more concurrent work without abandoning the reliability and controls that our users need.

What comes next

We’re building an architecture designed to provide the highest throughput and reliability available. Reads and writes scale independently, and the system recovers gracefully from failures. In internal benchmarks, it has delivered up to 35 times higher write throughput, with read capacity that scales on its own to meet demand.

As automated development increases the frequency and concurrency of software change, GitHub will evolve its foundations without trading away the governance and control that teams rely on. We’re already putting that foundation in place. In the next post in this series, we’ll dive deeper into our future architecture and the journey that led us there.

The post Building Git infrastructure for agent-scale development appeared first on The GitHub Blog.

Scroll to Top