Every team considering AI-assisted development asks the same question: does it actually work on a real codebase, or just on demos? The .NET team at Microsoft decided to find out the hard way — by pointing GitHub's Copilot Coding Agent (CCA) at dotnet/runtime, one of the largest and most complex .NET repositories, and letting it run for ten months.
The results, published in a detailed retrospective on the .NET Blog, offer something rare in the AI hype cycle: actual production data. Not cherry-picked demos. Not contrived benchmarks. Real pull requests, real reviews, real reverts. Here's what the numbers reveal, and what your team can learn from them.
The headline numbers
Between May 2025 and March 2026, CCA generated 878 pull requests against dotnet/runtime. Of those, 535 were merged — a 67.9% success rate. The merged PRs represented roughly 95,000 lines added and 31,000 lines removed.
For context, human Microsoft contributors achieved an 87.1% merge rate, and community contributors reached 79.7%. CCA's lower rate isn't necessarily a quality signal — the agent tackled more experimental and exploratory work that humans might not have prioritised.
The more telling metric is the revert rate. Of those 535 merged PRs, only 3 were reverted (0.6%). Human-authored PRs during the same period had a 0.8% revert rate. The code that made it through review was solid.
What the agent actually did
Not all tasks are created equal, and CCA's success varied dramatically by category:
| Category | Success Rate |
|---|---|
| Removal and cleanup | 84.7% |
| Testing | 75.6% |
| Refactoring | 69.7% |
| Bug fixes | 69.4% |
| Documentation | 68.1% |
| Feature development | 64.5% |
| Performance optimisation | 54.5% |
The pattern is intuitive. Mechanical, well-scoped work like removing dead code or adding test coverage played to the agent's strengths. Performance optimisation — which requires deep understanding of runtime behaviour, cache hierarchies, and platform-specific nuances — was the weakest category.
The file composition of merged PRs tells a similar story: 65.7% of added lines were test code, 29.6% were production code, and 4.7% were documentation or project files. CCA was predominantly a testing and cleanup machine.
The instruction file that changed everything
Early in the experiment, CCA's success rate was a dismal 38.1%. Builds failed because the agent didn't know the right incantations. Tests broke because it asserted on exact exception messages that change across localisations. PRs were rejected because they violated architectural conventions that aren't written down anywhere obvious.
The turning point was investing in a thorough copilot-instructions.md file. After documenting build processes, testing conventions, and architectural boundaries, the success rate jumped to 69%.
The dotnet/runtime instruction file covers:
- Build commands: Inner-loop vs. outer-loop builds, with component-specific commands for CoreCLR, Mono, Libraries, and WASM. The agent must complete a baseline build on
mainbefore making any changes. - Code style: File-scoped namespaces, pattern matching over equality operators,
is nullover== null, and no regression comments citing GitHub issues unless asked. - Testing patterns: Prefer extending existing test files over creating new ones, use
[Theory]with data sources instead of duplicative[Fact]methods, and never assert on exact exception messages. - Platform awareness: The agent runs on Linux only, which constrains what it can build and test.
// TIP
If you're adopting CCA on your own repositories, treat copilot-instructions.md with the same rigour as production code. The .NET team's experience shows this single file has more impact on agent success than any other factor.
# Build Requirements
Any code you commit MUST compile, and new and existing tests
related to the change MUST pass.
Complete a baseline build BEFORE making any code changes.
Skipping this causes "missing testhost" and "shared framework"
errors that waste time.
# Testing Conventions
- Prefer [Theory] with multiple data sources over duplicate [Fact] methods
- Do NOT assert on exact exception messages (they vary by localisation)
- Extend existing test files rather than creating new ones
Where CCA excelled
Certain areas of the runtime proved remarkably well-suited to AI-assisted development:
- System.Runtime.InteropServices: 90% success rate
- System.Reflection: 87.5%
- System.Net.Http: 86.7%
- System.Text.RegularExpressions: 82.6%
These namespaces share common traits: clear patterns, comprehensive existing test coverage, and tasks that tend toward mechanical transformation rather than architectural judgement. When the agent could pattern-match against existing code, it performed exceptionally well.
The sweet spot for PR size was 1-50 lines, with a 76-80% success rate. But interestingly, the largest PRs (1,001+ lines) also showed a strong 72% success rate — provided they were well-scoped mechanical tasks. A single PR that removed obsolete NET9_0 preprocessor constants across 112 files is a perfect example: massive in scope, trivial in judgement.
The issue backlog effect
One of the most compelling findings was CCA's impact on long-neglected issues. Of the 464 issues the agent addressed, 20% had been open for two or more years. The mean issue age was 382 days.
These weren't critical bugs sitting unresolved. They were the kind of improvements that every team knows should happen but never reach the priority queue: adding missing test coverage, cleaning up deprecated patterns, fixing minor inconsistencies. CCA turned these perpetual backlog items into merged PRs.
// NOTE
The .NET team reported that a single afternoon — dubbed "The Birthday Party Experiment" — yielded 22 assigned issues with over 20 resulting PRs, including a thread-safety fix in System.Text.Json and debugging assistance with regex engines.
Where CCA struggled
Native code: Running exclusively on Linux, CCA couldn't effectively handle Windows-specific code, cross-architecture variations, or C++ code that involved memory management and platform-specific preprocessor guards. C++ was consistently more error-prone than C#.
Architectural decisions: Tasks requiring broad codebase familiarity, design pattern selection, or understanding downstream implications across platforms showed markedly lower success rates. The agent could refactor code within a well-defined boundary, but it couldn't reason about how a change in one subsystem would ripple through the build matrix.
Vague requirements: The team found a direct correlation between issue clarity and PR quality. Issues with clear reproduction steps, expected behaviour descriptions, and defined scope produced significantly better first attempts. Vague assignments produced vague results — no surprise, but worth quantifying.
The review bottleneck nobody expected
Here's the finding that should concern every team planning large-scale AI adoption: CCA created a review bottleneck.
One contributor generated 9 complex PRs in a single session, each requiring 5-9 hours of expert review. The asymmetry is stark — generating PRs is now dramatically faster than reviewing them. The .NET team found that code review, not code generation, became the constraining factor in their workflow.
This isn't a theoretical concern. It's a capacity planning problem. If your team adopts CCA without also investing in review tooling, mentoring, or process changes, you'll shift the bottleneck without relieving the pressure.
// WARNING
Adopting AI code generation without scaling your review capacity is like adding more lanes to a motorway that ends in a single-lane roundabout. The traffic just backs up somewhere else.
How other .NET repositories compared
The .NET team ran CCA across seven repositories, merging 1,885 of 2,963 total PRs (68.6%):
| Repository | Success Rate |
|---|---|
| dotnet/extensions | 79.7% |
| modelcontextprotocol/csharp-sdk | 77.3% |
| dotnet/roslyn | 74.7% |
| dotnet/aspnetcore | 71.2% |
| dotnet/efcore | 69.2% |
| dotnet/runtime | 67.9% |
| microsoft/aspire | 64.8% |
Newer codebases with modern patterns (extensions, MCP SDK) performed better. Older, more constrained codebases with legacy patterns showed lower success rates. The lesson: the cleaner and more consistent your codebase, the more value you'll extract from AI agents.
Common pitfalls
Skipping the instruction file. CCA at 38% success isn't useful. CCA at 69% is. The difference is documentation. If you're trialling CCA on a repository without a thorough copilot-instructions.md, you're measuring the wrong thing.
Assigning vague issues. "Improve performance of X" is not an actionable issue for a human, and it's worse for an agent. "Reduce allocations in HttpClient.SendAsync by pooling HttpRequestMessage instances" gives the agent something concrete to work with.
Ignoring the review cost. Every AI-generated PR still needs human review. If your team is already at capacity for reviews, adding CCA will make things worse before it makes them better. Budget review time explicitly.
Expecting architectural reasoning. CCA excels at executing well-scoped tasks within established patterns. It does not excel at deciding whether a task should be done, choosing between competing approaches, or anticipating side effects across subsystem boundaries. Keep humans in the loop for judgement calls.
Treating early failures as final verdicts. The .NET team's success rate nearly doubled from May to October. Most of that improvement came from better instructions and task selection, not from model improvements. Give the process time to mature.
Summary
- CCA merged 535 of 878 PRs (67.9%) in
dotnet/runtimeover ten months, with a 0.6% revert rate — comparable to human-authored code - Success rates improved from 41.7% to 72.1% as the team invested in documentation and task scoping
- The agent was strongest at cleanup (84.7%), testing (75.6%), and refactoring (69.7%), and weakest at performance optimisation (54.5%)
- A thorough
copilot-instructions.mdfile was the single biggest factor in improving outcomes, lifting success from 38.1% to 69% - CCA cleared 20% of its issues from a backlog of two or more years, proving its value for work that never reaches human priority queues
- Code review, not code generation, became the bottleneck — teams must plan for this capacity shift
- Cleaner, more consistent codebases with modern patterns yield higher AI agent success rates