Add .NET refactoring skill - #873
Conversation
Skill Coverage Report
Uncovered:
|
|
/evaluate |
Skill Validation Results
[1]
Model: claude-opus-4.6 | Judge: claude-opus-4.6 🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
There was a problem hiding this comment.
Pull request overview
This PR adds two new .NET-focused skills to the plugins/dotnet plugin—one for behavior-preserving C# refactoring workflows and one for .NET-specific breaking-change/compatibility guardrails—along with eval coverage and fixtures that exercise the key hazards (public API baselines, multi-targeting #if, generated/partial code, and friend assemblies).
Changes:
- Added
csharp-refactoringskill guidance plus an operation catalog reference. - Added
dotnet-breaking-changesskill guidance plus focused reference docs for the four “hidden surfaces”. - Added eval suites + identical lightweight Billing fixture solutions under
tests/dotnet/for both skills; updated CODEOWNERS and dotnet plugin README.
Show a summary per file
| File | Description |
|---|---|
| tests/dotnet/dotnet-breaking-changes/tests/Billing.Tests/BillingTests.cs | Adds fixture tests used by the breaking-change eval scenarios. |
| tests/dotnet/dotnet-breaking-changes/tests/Billing.Tests/Billing.Tests.csproj | Adds an xUnit test project targeting net10.0 for the breaking-change fixture solution. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/PublicAPI.Shipped.txt | Adds a public API baseline file used by eval prompts/assertions. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/Pricing.cs | Adds pricing/tax types used for refactor/breaking-change exercises. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/PlatformInfo.cs | Adds multi-targeting #if example surface for breaking-change checks. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/OrderProcessor.cs | Adds intentionally-refactorable implementation used by scenarios. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/Coupons.g.cs.template | Adds generator-input template to simulate generated/partial hazards. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/Coupons.cs | Adds hand-authored partial type paired with the generated template. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/Billing.csproj | Adds multi-targeted library project and build-time generation hook. |
| tests/dotnet/dotnet-breaking-changes/src/Billing/AppSettingsHelper.cs | Adds duplicated parsing helpers for public-API consolidation scenarios. |
| tests/dotnet/dotnet-breaking-changes/Fixture.sln | Adds the breaking-change fixture solution container. |
| tests/dotnet/dotnet-breaking-changes/eval.yaml | Adds breaking-change eval scenarios and build/test gates. |
| tests/dotnet/csharp-refactoring/tests/Billing.Tests/BillingTests.cs | Adds fixture tests used by the refactoring eval scenarios. |
| tests/dotnet/csharp-refactoring/tests/Billing.Tests/Billing.Tests.csproj | Adds an xUnit test project targeting net10.0 for the refactoring fixture solution. |
| tests/dotnet/csharp-refactoring/src/Billing/PublicAPI.Shipped.txt | Adds a public API baseline file used by refactoring prompts/assertions. |
| tests/dotnet/csharp-refactoring/src/Billing/Pricing.cs | Adds pricing/tax types used for refactoring exercises. |
| tests/dotnet/csharp-refactoring/src/Billing/PlatformInfo.cs | Adds multi-targeting #if example surface for refactoring checks. |
| tests/dotnet/csharp-refactoring/src/Billing/OrderProcessor.cs | Adds intentionally-refactorable implementation used by scenarios. |
| tests/dotnet/csharp-refactoring/src/Billing/Coupons.g.cs.template | Adds generator-input template to simulate generated/partial hazards. |
| tests/dotnet/csharp-refactoring/src/Billing/Coupons.cs | Adds hand-authored partial type paired with the generated template. |
| tests/dotnet/csharp-refactoring/src/Billing/Billing.csproj | Adds multi-targeted library project and build-time generation hook. |
| tests/dotnet/csharp-refactoring/src/Billing/AppSettingsHelper.cs | Adds duplicated parsing helpers used in refactoring scenarios. |
| tests/dotnet/csharp-refactoring/Fixture.sln | Adds the refactoring fixture solution container. |
| tests/dotnet/csharp-refactoring/eval.yaml | Adds refactoring eval scenarios and build/test gates. |
| plugins/dotnet/skills/dotnet-breaking-changes/SKILL.md | Introduces the breaking-change skill and when/when-not-to-use guidance. |
| plugins/dotnet/skills/dotnet-breaking-changes/references/source-generation.md | Adds detailed guidance for generated/partial code hazards. |
| plugins/dotnet/skills/dotnet-breaking-changes/references/public-api.md | Adds detailed guidance for public API gating and compatibility decisions. |
| plugins/dotnet/skills/dotnet-breaking-changes/references/multi-targeting.md | Adds detailed guidance for multi-targeting and #if branch correctness. |
| plugins/dotnet/skills/dotnet-breaking-changes/references/internals-visible-to.md | Adds detailed guidance for friend-assembly/internal-surface hazards. |
| plugins/dotnet/skills/csharp-refactoring/SKILL.md | Introduces the refactoring skill and a stepwise safety contract. |
| plugins/dotnet/skills/csharp-refactoring/references/operation-catalog.md | Adds a more detailed refactoring operation taxonomy reference. |
| plugins/dotnet/README.md | Updates the dotnet plugin skill list to include the new skills. |
| eng/known-domains.txt | Adds github.com/dotnet/skills to known domains. |
| .github/CODEOWNERS | Adds ownership entries for the new skills and their tests. |
Copilot's findings
Tip
Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
- Files reviewed: 34/34 changed files
- Comments generated: 0
|
/evaluate |
|
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
Skill Validation Results
[1]
Model: claude-opus-4.6 | Judge: claude-opus-4.6 🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
JanKrivanek
left a comment
There was a problem hiding this comment.
Since the skills have nontrivial size - the main concern now is overlap between the two skills + cross-loading cost. csharp-refactoring instructs the agent to also load dotnet-breaking-changes, and the two share material ("stop and escalate," public API, forwarders, partial/generated). The split is defensible (breaking-changes is meant to apply to features/fixes too, not just refactors), but loading two long skills for one refactor is a real context cost that the evals don't yet justify.
Trim study for PR #873: csharp-refactoring-trim is a ~30% shorter rewrite of csharp-refactoring (drops When-to-use/Inputs/Outputs boilerplate + verbose LSP launch prose, compresses the dbc-overlapping hazards section) while preserving every rubric-rewarded behavior. Adds a 'trimmed' experiment arm and bumps runs to 3 so CI generates baseline/skilled/trimmed trajectories for local cross-family re-judging. Eval-only scratch branch; not part of the PR. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3c9f9823-7f2f-4f7d-9d1b-f2b9e7a20c60
|
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
Robustness probe for the csharp-refactoring trim: all prior CI trajectories were executed by claude-opus-4.6. Flip executor to gpt-5.5 (GPT family) and CI judge to claude-opus-4.8 so the (executor, judge) pair stays cross-family. Tests whether trimmed-vs-current non-inferiority holds under a different model family. Eval-branch only; not on PR #873. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 3c9f9823-7f2f-4f7d-9d1b-f2b9e7a20c60
Closes the 'refactoring-only endpoint too narrow' risk flagged by the
cross-family rubber-duck. Adds two boundary stimuli never used in trim tuning:
9. Separate a smuggled behavior change from a requested rename (rename +
smuggled 10%->12% discount change) -> should separate/flag, not bundle.
10. Recognize a behavior-changing simplification is not a refactor (remove
free-shipping-over-\ rule under a 'cleanup' framing) -> should not
silently change observable behavior.
Executor reverted to claude-opus-4.6 / judge gpt-5.5 (re-judge) to match the
stringent Confirm A family where boundary behavior looked weakest.
Eval-branch only; not on PR #873.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 3c9f9823-7f2f-4f7d-9d1b-f2b9e7a20c60
|
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
|
@JanKrivanek — fair concern, thanks. I took the size + cross-loading cost seriously and did two things: trimmed both skills and ran a breadth eval to confirm the trims don't cost quality. Size (SKILL.md bodies; the
Combined the two bodies drop ~40% (20.8k→12.4k chars), which directly cuts the "load two long skills for one refactor" cost you flagged. I removed duplicated prose — the Eval — does the trim hurt? 5 executor families (opus-4.8, gpt-5.5, sonnet-4.6, haiku-4.5, mai-flash), a 2-family judge ensemble (gpt-5.5 + claude-opus-4.8), blinded arm order, n=5, with a negative-control scenario per skill:
Honest caveats: single self-authored fixture suite, small scenario counts (6 for dbc), and CIs are bootstrapped over trials — so treat the sub-significant deltas as suggestive rather than proof of broad generalization. I'm not claiming formal non-inferiority, only that the smaller versions hold up at least as well as the larger ones across every family I could measure. On the split: I kept it deliberately (breaking-changes is meant to apply to features/fixes, not only refactors), but the shared material is now much thinner, so cross-loading is meaningfully cheaper. Re-triggering |
|
/evaluate |
|
❌ Evaluation failed. View workflow run |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
- Updated the skill description to clarify usage and restrictions for refactoring requests. - Added new test cases for preserving serialized contracts in CustomerProfile. - Introduced CustomerProfile class to support new test scenarios.
…aming and compatibility
…ctoring skill documentation and tests
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…elines and evaluation scenarios
…ntation for clarity and completeness
…ude ordinary tests
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Clarify mixed-request classification, add a real shipped-wrapper caller, and strengthen deterministic fixture and delegation guards. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
Restore the focused one-operation workflow and remove broad worktree restore advice after integrating current main. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The missing advertised skill, incorrect experiment label, weak executable-test checks, and stale ownership entries must be addressed.
Get a fresh assessment by requesting another Copilot review.
Review tier: Lite
Findings: 1
Move target discovery into a tested PowerShell script so GitHub Actions can parse and dispatch the workflow. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
📊 Skill and Agent Evaluation Results2 model/target results across 1 target and 2 models — ✅ 2 improved, ➖ 0 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
✅ Improved — csharp-refactoring (claude-sonnet-5)Why: Net win +58.8% (11W/5T/1L over 17 preference-eligible stimulus vote(s), sign test p=0.003), mean preference +38.9% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=17; 11W/5T/1L; d=12; p=0.003; net +58.8%; 1 dormancy excluded Warnings: Activation: isolated 12/17; plugin 15/17 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 18 paired runs (11W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. |
Correct the judge-comparison label and remove stale dotnet-breaking-changes registry entries left by the rebase. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 8c45529b-2515-483d-9e51-e6c0b7cb6852
📊 Skill and Agent Evaluation Results2 model/target results across 1 target and 2 models — ✅ 2 improved, ➖ 0 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
✅ Improved — csharp-refactoring (claude-sonnet-5)Why: Net win +64.7% (12W/4T/1L over 17 preference-eligible stimulus vote(s), sign test p=0.002), mean preference +34.4% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=17; 12W/4T/1L; d=13; p=0.002; net +64.7%; 1 dormancy excluded Warnings: Activation: isolated 14/17; plugin 13/17 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 18 paired runs (12W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — csharp-refactoring (gpt-5.6-luna)Why: Net win +35.3% (7W/9T/1L over 17 preference-eligible stimulus vote(s), sign test p=0.035), mean preference +23.3% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Review overfit evidence. State: Gate evidence: n=17; 7W/9T/1L; d=8; p=0.035; net +35.3%; 1 dormancy excluded Overfit: Moderate (score 0.22) Repeated-run reliability (not used by the gate): 18 paired runs (7W/10T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. |
|
✅ Evaluation passed for |
📊 Skill and Agent Evaluation Results6 model/target results across 1 target and 6 models — ✅ 4 improved, ➖ 2 not proven improved, Measurement identity: evaluated commit Measurement health: 6 expected / 6 observed / 6 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
➖ Not proven improved — csharp-refactoring (gpt-5.3-codex)Why: Net win +17.6% (7W/6T/4L over 17 preference-eligible stimulus vote(s), sign test p=0.274), mean preference +6.7% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.274 > 0.05) Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=17; 7W/6T/4L; d=11; p=0.274; net +17.6%; 1 dormancy excluded Warnings: Activation: isolated 7/17; plugin 8/17; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 2 failed runs Overfit: Low (score 0.19) Repeated-run reliability (not used by the gate): 18 paired runs (7W/7T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — csharp-refactoring (mai-code-1.1-flash)Why: Net win +11.8% (4W/11T/2L over 17 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=17; 4W/11T/2L; d=6; p=0.344; net +11.8%; 1 dormancy excluded Warnings: Activation: isolated 2/17; plugin 0/17 Overfit: Moderate (score 0.31) Repeated-run reliability (not used by the gate): 18 paired runs (5W/11T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — csharp-refactoring (claude-haiku-4.5)Why: Net win +29.4% (5W/12T/0L over 17 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +15.6% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=17; 5W/12T/0L; d=5; p=0.031; net +29.4%; 1 dormancy excluded Warnings: Activation: isolated 12/17; plugin 14/17 Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 18 paired runs (5W/12T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — csharp-refactoring (claude-opus-4.8)Why: Net win +58.8% (10W/7T/0L over 17 preference-eligible stimulus vote(s), sign test p=0.001), mean preference +38.9% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=17; 10W/7T/0L; d=10; p=0.001; net +58.8%; 1 dormancy excluded Warnings: Activation: isolated 14/17; plugin 16/17 Overfit: Moderate (score 0.50) Repeated-run reliability (not used by the gate): 18 paired runs (10W/8T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — csharp-refactoring (claude-sonnet-5)Why: Net win +82.4% (14W/3T/0L over 17 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +47.8% across 18 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=17; 14W/3T/0L; d=14; p=0.000; net +82.4%; 1 dormancy excluded Warnings: Activation: isolated 15/17; plugin 13/17 Overfit: Moderate (score 0.48) Repeated-run reliability (not used by the gate): 18 paired runs (14W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 1 result are in Full Results. 🔍 Full Results - all metrics and investigation details
|


Summary
Refactoring is a common operation for .NET developers, and a dedicated skill can help agents make these changes in the way .NET teams expect: behavior-preserving, incremental, and validated with build/test gates.
This PR adds two related skills:
csharp-refactoring: guides safe C#/.NET refactoring work such as rename, move, extract, inline, split, consolidate/de-duplicate, and modernization while preserving behavior.dotnet-breaking-changes: provides the compatibility guardrails needed when a refactor touches public API, multi-targeting, source-generated/partial code, orInternalsVisibleTosurfaces.Why these skills
csharp-refactoringfocuses on the refactoring workflow itself: establish a green baseline, choose one named operation, find true references, prefer semantics-aware edits, and re-gate after each step. Its reference file expands the operation catalog and maps common refactoring operations to the kinds of changes .NET teams regularly make.dotnet-breaking-changescovers the .NET-specific surfaces where a change can compile and pass tests while still breaking downstream consumers. Its reference files explain how to inspect and handle:PublicAPI.*, ApiCompat, package validation)#ifbranchesInternalsVisibleToTogether, the skills let an agent keep refactoring focused on behavior preservation while still checking the .NET compatibility surfaces that matter in real repos.
Validation