If AI coding is so good … where are the performance numbers? – Pivot to AI
►https://pivot-to-ai.com/2026/01/13/if-ai-coding-is-so-good-where-are-the-performance-numbers
(...)
We have one public study of AI coding performance that applied any reasonable methodology to the coding itself. That’s the METR study from July last year.
METR got 16 experienced open-source developers with “moderate AI experience.” The devs fixed real bug reports in their own projects. They used either Cursor with Claude Code or no AI help, at random.
METR actually screen-recorded and timed the work. The devs said they’d worked 20% faster — but they’d actually been slowed down by 19%. [blog post; paper, PDF]
►https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study
And that’s why you have to measure. Vibe self-reports are wrong. You must measure.
(...)
In fact, IEEE Spectrum ran a story where an ardent vibe coder notices exactly that: “AI Coding Assistants Are Getting Worse: Newer models are more prone to silent but deadly failure modes.” [IEEE]
– A task that might have taken five hours assisted by AI, and perhaps 10 hours without it, is now more commonly taking seven or eight hours, or even longer. It’s reached the point where I am sometimes going back and using older versions of large language models.
(...)
En rapport aussi avec :
▻https://seenthis.net/messages/1153910
AI Coding Degrades : Silent Failures Emerge - IEEE Spectrum
►https://spectrum.ieee.org/ai-coding-degrades