Every team considering AI-assisted development asks the same question: does it actually work on a real codebase, or just on demos? The .NET team at Microsoft decided to find out the hard way — by pointing GitHub's Copilot Coding Agent (CCA) at dotnet/runtime, one of the largest and most complex .NET repositories, and letting it run for ten months.

The results, published in a detailed retrospective on the .NET Blog, offer something rare in the AI hype cycle: actual production data. Not cherry-picked demos. Not contrived benchmarks. Real pull requests, real reviews, real reverts. Here's what the numbers reveal, and what your team can learn from them.

The headline numbers

Between May 2025 and March 2026, CCA generated 878 pull requests against dotnet/runtime. Of those, 535 were merged — a 67.9% success rate. The merged PRs represented roughly 95,000 lines added and 31,000 lines removed.

For context, human Microsoft contributors achieved an 87.1% merge rate, and community contributors reached 79.7%. CCA's lower rate isn't necessarily a quality signal — the agent tackled more experimental and exploratory work that humans might not have prioritised.

The more telling metric is the revert rate. Of those 535 merged PRs, only 3 were reverted (0.6%). Human-authored PRs during the same period had a 0.8% revert rate. The code that made it through review was solid.

What the agent actually did

Not all tasks are created equal, and CCA's success varied dramatically by category:

Category Success Rate
Removal and cleanup 84.7%
Testing 75.6%
Refactoring 69.7%
Bug fixes 69.4%
Documentation 68.1%
Feature development 64.5%
Performance optimisation 54.5%

The pattern is intuitive. Mechanical, well-scoped work like removing dead code or adding test coverage played to the agent's strengths. Performance optimisation — which requires deep understanding of runtime behaviour, cache hierarchies, and platform-specific nuances — was the weakest category.

The file composition of merged PRs tells a similar story: 65.7% of added lines were test code, 29.6% were production code, and 4.7% were documentation or project files. CCA was predominantly a testing and cleanup machine.

The instruction file that changed everything

Early in the experiment, CCA's success rate was a dismal 38.1%. Builds failed because the agent didn't know the right incantations. Tests broke because it asserted on exact exception messages that change across localisations. PRs were rejected because they violated architectural conventions that aren't written down anywhere obvious.

The turning point was investing in a thorough copilot-instructions.md file. After documenting build processes, testing conventions, and architectural boundaries, the success rate jumped to 69%.

The dotnet/runtime instruction file covers:

// TIP

If you're adopting CCA on your own repositories, treat copilot-instructions.md with the same rigour as production code. The .NET team's experience shows this single file has more impact on agent success than any other factor.

.github/copilot-instructions.md
# Build Requirements

Any code you commit MUST compile, and new and existing tests
related to the change MUST pass.

Complete a baseline build BEFORE making any code changes.
Skipping this causes "missing testhost" and "shared framework"
errors that waste time.

# Testing Conventions

- Prefer [Theory] with multiple data sources over duplicate [Fact] methods
- Do NOT assert on exact exception messages (they vary by localisation)
- Extend existing test files rather than creating new ones

Where CCA excelled

Certain areas of the runtime proved remarkably well-suited to AI-assisted development:

These namespaces share common traits: clear patterns, comprehensive existing test coverage, and tasks that tend toward mechanical transformation rather than architectural judgement. When the agent could pattern-match against existing code, it performed exceptionally well.

The sweet spot for PR size was 1-50 lines, with a 76-80% success rate. But interestingly, the largest PRs (1,001+ lines) also showed a strong 72% success rate — provided they were well-scoped mechanical tasks. A single PR that removed obsolete NET9_0 preprocessor constants across 112 files is a perfect example: massive in scope, trivial in judgement.

The issue backlog effect

One of the most compelling findings was CCA's impact on long-neglected issues. Of the 464 issues the agent addressed, 20% had been open for two or more years. The mean issue age was 382 days.

These weren't critical bugs sitting unresolved. They were the kind of improvements that every team knows should happen but never reach the priority queue: adding missing test coverage, cleaning up deprecated patterns, fixing minor inconsistencies. CCA turned these perpetual backlog items into merged PRs.

// NOTE

The .NET team reported that a single afternoon — dubbed "The Birthday Party Experiment" — yielded 22 assigned issues with over 20 resulting PRs, including a thread-safety fix in System.Text.Json and debugging assistance with regex engines.

Where CCA struggled

Native code: Running exclusively on Linux, CCA couldn't effectively handle Windows-specific code, cross-architecture variations, or C++ code that involved memory management and platform-specific preprocessor guards. C++ was consistently more error-prone than C#.

Architectural decisions: Tasks requiring broad codebase familiarity, design pattern selection, or understanding downstream implications across platforms showed markedly lower success rates. The agent could refactor code within a well-defined boundary, but it couldn't reason about how a change in one subsystem would ripple through the build matrix.

Vague requirements: The team found a direct correlation between issue clarity and PR quality. Issues with clear reproduction steps, expected behaviour descriptions, and defined scope produced significantly better first attempts. Vague assignments produced vague results — no surprise, but worth quantifying.

The review bottleneck nobody expected

Here's the finding that should concern every team planning large-scale AI adoption: CCA created a review bottleneck.

One contributor generated 9 complex PRs in a single session, each requiring 5-9 hours of expert review. The asymmetry is stark — generating PRs is now dramatically faster than reviewing them. The .NET team found that code review, not code generation, became the constraining factor in their workflow.

This isn't a theoretical concern. It's a capacity planning problem. If your team adopts CCA without also investing in review tooling, mentoring, or process changes, you'll shift the bottleneck without relieving the pressure.

// WARNING

Adopting AI code generation without scaling your review capacity is like adding more lanes to a motorway that ends in a single-lane roundabout. The traffic just backs up somewhere else.

How other .NET repositories compared

The .NET team ran CCA across seven repositories, merging 1,885 of 2,963 total PRs (68.6%):

Repository Success Rate
dotnet/extensions 79.7%
modelcontextprotocol/csharp-sdk 77.3%
dotnet/roslyn 74.7%
dotnet/aspnetcore 71.2%
dotnet/efcore 69.2%
dotnet/runtime 67.9%
microsoft/aspire 64.8%

Newer codebases with modern patterns (extensions, MCP SDK) performed better. Older, more constrained codebases with legacy patterns showed lower success rates. The lesson: the cleaner and more consistent your codebase, the more value you'll extract from AI agents.

Common pitfalls

Skipping the instruction file. CCA at 38% success isn't useful. CCA at 69% is. The difference is documentation. If you're trialling CCA on a repository without a thorough copilot-instructions.md, you're measuring the wrong thing.

Assigning vague issues. "Improve performance of X" is not an actionable issue for a human, and it's worse for an agent. "Reduce allocations in HttpClient.SendAsync by pooling HttpRequestMessage instances" gives the agent something concrete to work with.

Ignoring the review cost. Every AI-generated PR still needs human review. If your team is already at capacity for reviews, adding CCA will make things worse before it makes them better. Budget review time explicitly.

Expecting architectural reasoning. CCA excels at executing well-scoped tasks within established patterns. It does not excel at deciding whether a task should be done, choosing between competing approaches, or anticipating side effects across subsystem boundaries. Keep humans in the loop for judgement calls.

Treating early failures as final verdicts. The .NET team's success rate nearly doubled from May to October. Most of that improvement came from better instructions and task selection, not from model improvements. Give the process time to mature.

Summary