AI moves too fast for traditional academic pace. During peer review alone, another generation of models comes along and might obsolete your findings. Though, in this case, I am not sure that Github Copilot specifically has improved that much in the meantime.
Setting aside issues with the sample size and response rate, the current progress within the field shows how difficult it is to study the impact of AI on productivity.
Even if the survey was conducted one year later, I would not find it useful to make inferences about the use of AI tools in September 2026.
What's very weird is that this study was published by two people working professionally at Okta. Yes, the auth tech company.
This is the title they chose for the study:
Beyond the Hype: The Efficiency-Throughput Gap with GitHub Copilot
How can you possibly have a title like that when the study was done in 2024? I'm guessing even their own engineers at Okta would roll the eyes at this study.
Copilot is so shoddily written (all LLMs are, but Copilot's is especially terrible). It's Microsoft's sloppy "we're here too!" attempt that continues to represent the benchmark for the bottom of the barrel.
GitHub Copilot is an incredibly limited tool compared to any real harness, and I'm so tired about all of these studies that claim a lack of efficiency while at the same time doing everything in their power to shoot themselves in the foot.
Also, it's insane to me that this whole article has been written without specifying anywhere the models being used.
Like, of course you're not gonna be productive if all you have is Sonnet 4.6???
> we found no immediate increase in key engineering metrics such as monthly pull requests and lines of code
> To establish a before-Copilot baseline, we used data from April, May, and June 2024. After-Copilot data was represented by the period of September, October, and November 2024
> GitHub Copilot usage varied significantly among engineers, the tool demonstrably fostered positive changes in perceived engineer value, reduced time spent on various engineering activities, and boosted motivation and perceived skills
> subsequent monitoring of PRs and LOC for participants from December 2024 to May 2025 showed no statistical improvements
> the implications of more advanced capabilities, such as retrieval-augmented generation (RAG) over enterprise codebases or deeper engineering workflow integrations, warrant separate investigation
I think it is just out of date, habits have changed as well. I have not seen much gain personally at that period except in the last 12 months. Also, models not named, token counts not shown. Not to mention it was the older autocomplete + chat that were in use, these days it is much more advanced with RAG, cli use, MCPs, etc.
Note the timelines of this study:
4/2024: Baseline metrics
8/2024: Participants given Github Copilot licenses
11/2024: Conducted surveys, 261 people invited and 97 responded in survey
9/2026: Study published
AI moves too fast for traditional academic pace. During peer review alone, another generation of models comes along and might obsolete your findings. Though, in this case, I am not sure that Github Copilot specifically has improved that much in the meantime.
Setting aside issues with the sample size and response rate, the current progress within the field shows how difficult it is to study the impact of AI on productivity.
Even if the survey was conducted one year later, I would not find it useful to make inferences about the use of AI tools in September 2026.
What's very weird is that this study was published by two people working professionally at Okta. Yes, the auth tech company.
This is the title they chose for the study:
How can you possibly have a title like that when the study was done in 2024? I'm guessing even their own engineers at Okta would roll the eyes at this study.Copilot is so shoddily written (all LLMs are, but Copilot's is especially terrible). It's Microsoft's sloppy "we're here too!" attempt that continues to represent the benchmark for the bottom of the barrel.
GitHub Copilot is an incredibly limited tool compared to any real harness, and I'm so tired about all of these studies that claim a lack of efficiency while at the same time doing everything in their power to shoot themselves in the foot.
Also, it's insane to me that this whole article has been written without specifying anywhere the models being used.
Like, of course you're not gonna be productive if all you have is Sonnet 4.6???
I think you’re missing that it’s not only about coding. That’s just one aspect of the work.
Claude Cowork is also a better harness than GitHub Copilot. Many such cases. I never mentioned coding being the only goal.
Considering the timeline, it would have been Sonnet 3.5
> we found no immediate increase in key engineering metrics such as monthly pull requests and lines of code
> To establish a before-Copilot baseline, we used data from April, May, and June 2024. After-Copilot data was represented by the period of September, October, and November 2024
> GitHub Copilot usage varied significantly among engineers, the tool demonstrably fostered positive changes in perceived engineer value, reduced time spent on various engineering activities, and boosted motivation and perceived skills
> subsequent monitoring of PRs and LOC for participants from December 2024 to May 2025 showed no statistical improvements
> the implications of more advanced capabilities, such as retrieval-augmented generation (RAG) over enterprise codebases or deeper engineering workflow integrations, warrant separate investigation
I think it is just out of date, habits have changed as well. I have not seen much gain personally at that period except in the last 12 months. Also, models not named, token counts not shown. Not to mention it was the older autocomplete + chat that were in use, these days it is much more advanced with RAG, cli use, MCPs, etc.
[dead]