webtrail Brian · My field notebook for trails and the web
Two women operating the ENIAC computer around 1946: one stands at the left-hand control panel while the other adjusts a bank of switches and patch cables on the machine, in a room lined with equipment racks.
4 September 2026 · 7 stops

The Psychology of Software Development in the AI Era

You have a strong opinion about whether AI has made you a better developer. So does everyone on your team, and almost none of it rests on evidence. Research into AI and the psychology of software development is younger than the tools it studies, and it is already more interesting than either the vendor decks or the backlash.
What the field does not offer is a verdict. The findings are mixed, in places contested, and the measurement itself is still being revised while the tools move underneath it. Several of these papers are preprints or under review, so read the collection as early evidence rather than settled science. What makes it worth your afternoon anyway is the questions being asked. Not only how fast you ship, but what the work now costs you in attention, judgement, learning and morale, and how closely your own sense of a good session tracks what actually happened in it.
The pages collected here are the primary sources, not write-ups of them. Each names its own limits, and those limits are usually the most useful paragraph on the page.
This post references third-party websites for informational purposes only. webtrail does not host, own, or claim any rights over the content of the linked sites. All screenshots are used for illustrative purposes and link back to their original source.

The stops

7 sites, each opened and read
01
metr.org
METR blog page titled 'Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity', with a summary paragraph reporting that developers take 19% longer when using AI tools, a byline crediting Joel Becker, Nate Rush, Beth Barnes and David Rein dated July 10 2025, and a yellow warning callout below reading 'These results are out of date' beside a link to results current as of early 2026.

AI and Developer Psychology

METR's Randomized Trial of AI Coding Tools on Real Issues

ai coding assistants developer productivity randomized controlled trial research

METR is an AI research nonprofit, and this July 2025 study measures something usually argued from anecdote: whether AI tools actually speed up experienced developers working in their own repositories.

The design is what gives it weight. Sixteen experienced open-source developers contributed 246 real issues from projects they already maintain, and each issue was randomly assigned to permit or forbid early-2025 AI tools. Developers took 19% longer on the issues where AI was allowed. The more uncomfortable finding sits right beside it: they had forecast a 24% speedup going in, and after doing the work still believed they had been sped up by 20%. The slowdown was real and invisible to the people living it.

The page is built for reading rather than skimming, with sections for motivation, methodology, the core result chart and a factor analysis working through candidate explanations. Joel Becker, Nate Rush, Beth Barnes and David Rein are credited, and arXiv and citation links sit at the top. Be clear about the scope before you quote it: sixteen developers on mature codebases is one setting, and METR presents the number as a snapshot rather than a verdict on AI coding generally.

Treat it as the landmark early result, not the current one. A yellow callout at the top of the page states plainly that the results are out of date and points to a February 2026 continuation measuring late-2025 models.

02
metr.org
METR blog post headed 'We are Changing our Developer Productivity Experiment Design', contributed by Joel Becker, Nate Rush, Tom Cunningham, David Rein and Khalid Mahamud and dated February 24 2026, with opening paragraphs recapping the earlier 20% slowdown finding and explaining why the new experiment's data is an unreliable signal, beside a sidebar table of contents and a newsletter subscribe box

AI and Developer Psychology

METR on Why Its AI Productivity Experiment Stopped Working

ai coding assistants developer productivity research methodology selection bias

METR is a research nonprofit that studies AI systems, and this February 2026 post is a rare thing: an institute reporting that its own measurement stopped working, before anyone forced the issue.

The setup is short. METR's earlier trial, run on data from February to June 2025, found that experienced open-source developers were about 20% slower when they used AI tools. To track how that changes over time, METR started a second experiment in August 2025 with a larger pool of developers and the latest tools. The update is that this new data gives an unreliable signal of the current productivity effect.

Three reasons are named, and all of them push the estimate the same way. Developers increasingly chose not to take part because they did not want to work without AI. The pay rate was cut from $150 an hour to $50, which changes who signs up. And time-on-task measurement is unreliable for the developers running several AI agents concurrently. METR's own read is that developers are likely more sped up in early 2026 than the 2025 figure implies, but that their data is only very weak evidence for the size of that increase, so the task-level design has to change.

The trade-off is that you get no new number here. If you came for a fresh productivity estimate, this will frustrate you. What you get is an honest account of how an experiment gets overtaken by the thing it measures.

03
arxiv.org
Screenshot of an arXiv abstract page under the breadcrumb 'Computer Science > Software Engineering', showing the paper title 'Same Scrutiny, More Time: Eye Tracking Insights into Reviewing LLM-Labelled Code', its five authors, the full abstract, a comments line noting acceptance at ASE 2026, and an 'Access Paper' sidebar linking to PDF, HTML and TeX source.

AI and Developer Psychology

Eye-Tracking Study of LLM-Labelled Code Review on arXiv

ai generated code code review developer trust eye tracking

arXiv hosts the preprint of "Same Scrutiny, More Time", an eye-tracking study of what software engineers actually do when the code in front of them is labelled as LLM-generated. Most AI-policy arguments settle that by assertion; this one measures it.

The abstract page sets out the design plainly. Ranim Khojah and four co-authors ran a Wizard-of-Oz experiment in which participants reviewed code explicitly labelled as LLM-generated, pairing eye-tracking with exit interviews so observed behaviour could be set against what people said they were doing, then analysed with Bayesian methods alongside qualitative coding. The headline result is narrower than the usual framing, and worth getting right: review thoroughness did not change, but reviewers spent more time fixating on the labelled code. The label moves attention, not rigour. The authors read that as a gap between reviewers' intentions and their actual reviewing behaviour.

The rest is about strategy. Participants adapted how they handled labelled code, judging it against specific criteria such as logical correctness, or using the originating prompt to steer the review, which is why the paper argues the prompt deserves treating as a software artifact in its own right.

One limit before you click through: this is the abstract page, so you get the abstract, the submission history and links to PDF, experimental HTML and TeX source, while sample size and effect sizes stay in the paper. It is a preprint, accepted at the 41st IEEE/ACM Automated Software Engineering conference for 2026.

04
arxiv.org
arXiv abstract page under the breadcrumb “Computer Science > Human-Computer Interaction”, titled “Using Biometrics to Understand AI-Assisted Coding Performance and its Perception”, submitted 19 May 2026, with eight author names, the full abstract describing the EEG, eye-tracking, electrodermal and NASA-TLX measurements, a comments line reading “Stage 2 RR under review at EMSE” with the Stage 1 protocol on OSF, and an “Access Paper” sidebar of PDF, HTML and TeX links.

AI and Developer Psychology

Biometrics of AI-Assisted Coding Workload, on arXiv

ai pair programming cognitive load eeg nasa tlx

arXiv hosts the preprint of a multisite study that wires developers up to sensors while they code with and without an AI assistant, and it earns catalog space because it measures what most AI-coding claims only assert.

You get the whole design on the abstract page, not a summary of it. Researchers at universities in Bari, Italy and Copenhagen, Denmark ran a within-subjects crossover study, recording electroencephalography, eye-tracking, electrodermal activity and heart-rate variability alongside a rubric-based performance score and self-reported workload across the six NASA Task Load Index dimensions. Under AI assistance the EEG theta/alpha ratio was lower on the first task and the gaze blink rate higher on the second — both, the authors write, consistent with reduced cognitive engagement when generative effort is offloaded to the model. Those are correlates, not evidence that anyone thought less, and the pattern did not differ between undergraduate and graduate students.

The conclusion is the part worth carrying into your own arguments: AI-assisted programming looks like a cognitively distinct activity, not a faster version of solo coding.

The honest limit is its status. The comments line marks it a Stage 2 registered report still under review at Empirical Software Engineering, with the Stage 1 protocol archived on OSF. The design was reviewed before any data was collected, which is more than most studies here can say, but the findings themselves have not finished peer review.

05
arxiv.org
arXiv abstract page under Computer Science > Software Engineering, headed "Cognitive Biases in LLM-Assisted Software Development", submitted 12 January 2026 by six authors, showing the full abstract and a metadata block listing 13 pages, 6 figures, 7 tables and a related ACM DOI.

AI and Developer Psychology

Cognitive Biases in LLM-Assisted Development on arXiv

cognitive bias decision making llm assisted development

This arXiv preprint is a rare attempt to measure how developers think while working with a language model, rather than how fast they ship. Xinyi Zhou and five co-authors start from the claim that coding with an LLM turns programming from a solution-generative activity into a solution-evaluative one, then go looking for what that shift does to judgement.

The method is mixed. Observational sessions with 14 student and professional developers are followed by surveys of 22 more. From a systematic analysis of 90 cognitive biases specific to developer-LLM interaction, the authors build a taxonomy of 15 bias categories and have cognitive psychologists validate it. Two numbers carry the paper: 48.8% of programmer actions were coded as biased, and developer-LLM interactions accounted for 56.4% of those biased actions. It ends with practices for developers and mitigation suggestions for people building LLM tooling.

Know what you are getting before you cite it. The sample is small, 36 people across both arms, so read the percentages as what this study observed rather than a rate you should expect on your own team. The paper also never names automation bias as a construct; deferring to model output sits inside the wider taxonomy rather than being isolated and measured. arXiv lists a related ACM DOI beside the preprint, so a publisher version exists, but what you open here is the 13-page preprint with its 6 figures and 7 tables.

06
anthropic.com
Anthropic research page tagged Alignment, headed "How AI assistance impacts the formation of coding skills", dated Jan 29 2026, with a "Read the paper" button above a large illustration of printed sheet music covered in a grid of bright red dots.

AI and Developer Psychology

How AI Assistance Shapes Coding Skills, from Anthropic

deskilling junior developers learning to code skill formation

Anthropic published a randomized controlled trial in January 2026 that asks what most tooling debates skip: not whether an AI assistant makes you faster today, but whether you still understand the code afterwards.

Judy Hanwen Shen and Alex Tamkin gave 52 mostly junior software engineers the same job — learn Trio, a Python library built around asynchronous programming — with or without AI assistance. On a later comprehension quiz the AI group averaged 50% against 67% for the hand-coding group, a gap the page puts at nearly two letter grades (Cohen's d=0.738, p=0.01). The largest difference showed up on debugging questions. The AI group finished about two minutes faster, and that difference was not statistically significant.

The more useful finding is that how you used the assistant changed the outcome. Engineers who scored well asked follow-up questions, requested explanations and posed conceptual questions rather than only generating code; the weaker scores clustered around heavy delegation of both code generation and debugging.

Worth stating plainly: this is Anthropic's own research into the category of tool it sells, and the finding runs against its commercial interest. The limitation is scope — one library, one sitting, 52 people — so read it as a signal about how skills form, not a verdict on your daily work.

07
arxiv.org
arXiv abstract page under the breadcrumb Computer Science, Software Engineering, headed From Gains to Strains: Modeling Developer Burnout with GenAI Adoption by Zixuan Feng, Sadia Afroz and Anita Sarma, with the full abstract about a survey of 442 developers analysed with PLS-SEM, a CC BY licence badge, and a submission history listing version 1 from October 2025 and version 2 from January 2026.

AI and Developer Psychology

From Gains to Strains: Developer Burnout Study on arXiv

developer burnout generative ai job demands wellbeing

arXiv hosts the preprint "From Gains to Strains," in which Zixuan Feng, Sadia Afroz and Anita Sarma ask what generative AI adoption does to developers' well-being rather than to their output. It earns catalog space because it puts survey numbers on a question most write-ups handle with anecdote.

The abstract page gives you the whole argument before you open the PDF. The authors take the Job Demands-Resources model as their analytic lens, run a concurrent embedded mixed-methods design, and survey 442 developers across a range of organisations, roles and experience levels. Partial Least Squares Structural Equation Modeling and regression carry the quantitative side, and a qualitative pass over open-ended responses puts those numbers in context. The headline result is genuinely two-sided: GenAI adoption heightens burnout by increasing job demands, while job resources and developers' own positive perceptions of GenAI mitigate that effect — which leads the authors to reframe adoption as an opportunity rather than a straightforward harm.

Two limits are worth knowing before you cite it. This is a preprint, first posted in October 2025 and revised in January 2026, with no acceptance notice on the page, so it has not visibly cleared peer review. And it is self-reported survey data analysed with structural equation modelling, so the relationships are correlational; read "heightens" as a modelled association rather than a demonstrated cause. The page carries a CC BY licence, PDF and experimental HTML versions and the usual citation exports, so reading the ten pages yourself costs nothing.

End of trail

← All trails