In a writing experiment, professionals with ChatGPT finished 40% faster. In Denmark, users of AI chatbots saved on average 2.8% of their work hours. Among 244 consultants, three ways of working with AI emerged, and they built different skills. In a support centre, the newest agents gained most, and those who followed the suggestions learned from them; students who leaned on an open chatbot did worse once it was gone.
Management Review · Second series · November 2026 · No. 51
Generative AI at work: what the studies show
Forty per cent faster in a writing test, 2.8% of work hours saved in Denmark, three ways consultants worked with AI, what support agents learned from it, what to count besides speed, and a card for one trial.
- No.
- 51
- Pages
- 10
- Sources
- 8
- Topics
- AI
Management Review · No. 51
The figures of the issue
The charts of the printed pages, with their sources.
Source: Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025
The whole text Read the issue as text For reading on a small screen, searching or a screen reader. The same words, without the page design.
In this issue
Few tools have reached offices as fast as AI chatbots, and few have been studied so quickly. The results differ widely. This issue reads them side by side: what each study measured, on which task, and what a manager can take from it for a team.
In a writing experiment, professionals with ChatGPT finished 40% faster. In Denmark, users of AI chatbots saved on average 2.8% of their work hours. Among 244 consultants, three ways of working with AI emerged, and they built different skills. In a support centre, the newest agents gained most, and those who followed the suggestions learned from them; students who leaned on an open chatbot did worse once it was gone.
Stiven Janaqi, Editor
Cover story
For t y per cent faster, on paper
In an experiment published in 2023, Shakked Noy and Whitney Zhang of MIT gave 453 college-educated professionals two writing tasks from their own jobs, such as a cover letter for a grant or an email about a restructuring. Half, chosen at random, could use ChatGPT for the second. Experienced people from the same occupations graded the work without knowing who had used it.
- 40% less time on the task, on average
- 18% better grades from the evaluators
- 68% handed in ChatGPT's first draft unedited (working paper)
Those who had done worse on the first task gained most, so the gap between workers narrowed. The authors also named the limits: the tasks needed no knowledge of a real company or customer, the instructions were explicit, and no one checked the facts in the texts.
Our reading
The experiment timed a draft, not a decision. Where a text has to be right about a customer, the time saved counts only after someone has checked it.
A preregistered online experiment with 20–30-minute tasks, published in Science in July 2023. The 68% comes from the working paper of March 2023 (444 participants), not from the published article; the limits are as the authors described them to MIT News.
Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Shakked Noy & Whitney Zhang, MIT, 2023; Zach Winn, MIT News, 2023
The numbers
From the test t o the working week
Anders Humlum and Emilie Vestergaard surveyed about 25,000 workers in Denmark, in 11 occupations exposed to AI chatbots, and linked the answers to registers of pay and hours. Time saved by users, as a share of their work hours, 2024:
Time saved as a share of work hours, with and without the employer's encouragement: Marketing: Not encouraged 4.6%, Encouraged 6.8%; Developers: Not encouraged 3.9%, Encouraged 6.5%; Accountants: Not encouraged 0.9%, Encouraged 2.2%; Teachers: Not encouraged 0.6%, Encouraged 1.0%.
Across all 11 occupations, users saved 2.8% of their hours, and 80% put the time into other tasks. In no occupation did pay or recorded hours change significantly.
Our reading
An experiment measures a task the tool suits. A working week is mostly other tasks.
Self-reported savings, turned by the authors into a share of work hours; 4 of the 11 occupations. Working paper of May 2025.
Source: Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025
The model
Three ways t o work with AI
244 junior consultants at Boston Consulting Group solved the same problem with GPT-4: which of a fictional company's three brands should get the investment, in a memo of at most 500 words. Steven Randazzo, Hila Lifshitz and colleagues read all 4,975 exchanges with the AI and held 237 interviews.
- Cyborgs · 60%. work with AI at every step, in constant back and forth; gained skill with AI and kept their domain knowledge
- Centaurs · 14%. decide what to do and how, and give AI chosen parts; most accurate answers; deepened their domain knowledge
- Self-au
t omat ors · 27%. hand the task over in one or two prompts; fast and polished, but shallow; built neither skill
Two questions separate the three: who chooses what has to be done, and who decides how it is done. Of the self-automators, 44% accepted the AI's text without any change; the rest made only surface edits.
Our reading
“A human in the loop” can mean three different things. For a manager the question is who still does the thinking, and what each way teaches.
At the time, the task lay outside what GPT-4 did well, beyond the “jagged frontier” of issue 10. Shares as rounded by the authors (101% in total); the names in Albanian and German are ours. Working paper of 2025.
Sources: Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick et al., Harvard Business School, 2023
What the research says
Who gains, who learns
Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed 5,179 support agents serving a large software firm as an AI assistant was rolled out. It suggested replies during each chat; the agents stayed responsible and could ignore it.
- 2 = 6 months: agents with two months and AI did as well as those with over six months without it
- 38% of the suggestions were followed, on average
Agents who followed the most suggestions gained close to 25%, those who followed the fewest about 10%. When the system was down, close followers still worked faster than before: they had learned. Before the AI, coaching meant a short weekly session with a manager.
It can go the other way. Nearly 1,000 high-school students in Turkey who practised maths with an open GPT-4 did 48% better in practice, but 17% worse in the exam without it than students who never had it. A version that gave hints instead of answers avoided the loss.
Our reading
The same tool can coach or replace. Whether people can still do the work when it is switched off has to be tested, not assumed.
One firm, text chats only; the link between following and gains may partly reflect who chose to follow. The published version (QJE, 2025) reports the same patterns. Bastani et al. studied students, not workers.
Sources: Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)
More in the essay: How I built Front Desk Control
How it is measured
Count more than speed
Speed is the easiest thing to measure. In the studies of this issue, the results also hinged on checking, on who gained, on where the time went and on what people could do without the tool. Five things to count in a trial:
- Time, with the checking. From the request to the finished, checked result.
- Quali
t y, judged blind. By someone who knows the work and not which version used AI. - New and experienced, apa
r t. An average can hide a gain for one and a loss for the other. - Where the saved time goes. Name the task it moves to, or it disappears.
- A day without the
t ool. Can the person still do the task alone?
Hypothe tical example, a month of AI draf ts for cust omer emails
- Time: 9 min per email before, 6 min now with checking
- Blind check: errors in 2 of 40 before, in 5 of 40 now
- By experience: new staff −4 min, experienced −1 min
Faster, but the errors rose: the trial is not finished. The numbers are invented.
The five checks are the editors' choice from the studies in this issue; the example is the editors'.
Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)
Tool of the issue
The AI trial card
One task, one month, the same check as before. Write down what you will count before the trial starts, so the result cannot be read to fit a wish. Keep confidential data out of tools the company has not approved.
- The task what, how often, who does it today
- Before the trial two weeks without AI: time per task and errors
- Who tries it new and experienced people, counted apart
- What
t o check facts, figures, names, tone, data that must not leave the team - Who does the thinking what the person decides, what the AI does
- Result and decision time with checking, errors, a day without the tool; keep, change or stop
A practice proposed by the editors, after Noy & Zhang, Randazzo et al. and Brynjolfsson et al.; the day without the tool follows the outages in the support centre.
Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023
Open the tool: Sigma & Control Chart
Sources and method
Every figure has a source.
The figures in this issue come from the sources below. The year shows how recent each one is.
- Shakked Noy & Whitney Zhang, Science 381(6654), “Experimental evidence on the productivity effects of generative artificial intelligence”, 2023 (via Abstract at Stanford SCALE). https://doi.org/10.1126/science.adh2586
- Shakked Noy & Whitney Zhang, MIT, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, 10 March 2023)”, 2023. https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1_0.pdf
- Zach Winn, MIT News, “Study finds ChatGPT boosts worker productivity for some writing tasks”, 2023. https://economics.mit.edu/news/study-finds-chatgpt-boosts-worker-productivity-some-writing-tasks
- Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, “Large Language Models, Small Labor Market Effects (Working Paper 2025-56, May 2025)”, 2025. https://bfi.uchicago.edu/wp-content/uploads/2025/04/BFI_WP_2025-56-3.pdf
- Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, “Cyborgs, Centaurs and Self-Automators: The Three Modes of Human-GenAI Knowledge Work and Their Implications for Skilling and the Future of Expertise”, 2025. https://www.hbs.edu/ris/Publication%20Files/26-036_e7d0e59a-904c-49f1-b610-56eb2bdfe6f9.pdf
- Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick et al., Harvard Business School, “Navigating the Jagged Technological Frontier”, 2023. https://papers.ssrn.com/abstract=4573321
- Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, “Generative AI at Work”, 2023. https://www.nber.org/papers/w31161
- Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), “Generative AI without guardrails can harm learning: Evidence from high school mathematics”, 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton). https://doi.org/10.1073/pnas.2422633122
Edit orial me thod
Each figure was checked for its year, its publisher and what exactly it measures. Where the publisher's page could not be opened, the figure was checked against independent summaries and is marked “via”. The editors' interpretation is marked “Our reading”. Figures that could not be confirmed are not in the issue.
Management Review · Monthly edition
