GitHub is enhancing its AI coding assistant with a built-in second opinion. The latest experimental feature for GitHub Copilot CLI, dubbed 'Rubber Duck', leverages a different AI model family to scrutinize the primary agent's plans and outputs.
This approach aims to mitigate the inherent biases of a single model reviewing its own work. By introducing an independent reviewer, GitHub seeks to catch critical errors that might otherwise compound through the development process.
Catching Confident Mistakes
Traditional AI coding agents follow a linear process: assess, plan, implement, test, and iterate. However, foundational decisions made early on, particularly during the planning phase, can embed inefficiencies or errors. A model reviewing itself is still constrained by its own training data and blind spots.
Rubber Duck acts as a focused review agent, powered by a complementary AI family. For instance, when a Claude model orchestrates the task, Rubber Duck might employ GPT-5.4 for review. This cross-family perspective is designed to surface overlooked details, questionable assumptions, and potential edge cases.
Performance Gains on Complex Tasks
Evaluations using the SWE-Bench Pro benchmark, which comprises difficult real-world coding problems, demonstrate significant improvements. Claude Sonnet paired with Rubber Duck (GPT-5.4) narrowed the performance gap with the more powerful Claude Opus model by 74.7%.
