Management Review · Second series · November 2026 · No. 51

Generative AI at work: what the studies show

Forty per cent faster in a writing test, 2.8% of work hours saved in Denmark, three ways consultants worked with AI, what support agents learned from it, what to count besides speed, and a card for one trial.

No.
51
Pages
10
Sources
8
Topics
AI
Stiven CatalystSecond series · November 2026
ManagementReview

Management without theatre.

AI

Generative AI at work:what the studies show

Forty per cent faster in a writing test, 2.8% of work hours saved in Denmark, three ways consultants worked with AI, what support agents learned from it, what to count besides speed, and a card for one trial.

No.51

2.8%

of their work hours: what users of AI chatbots in 11 occupations in Denmark saved on average in 2024.Humlum & Vestergaard, 2025

Inside

  1. Cover storyForty per cent faster, on paperPage 03
  2. The modelThree ways to work with AIPage 05
  3. Tool of the issueThe AI trial cardPage 08

stivencatalyst.com

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst2

No. 51 · AI

In this issue

Few tools have reached offices as fast as AI chatbots, and few have been studied so quickly. The results differ widely. This issue reads them side by side: what each study measured, on which task, and what a manager can take from it for a team.

In a writing experiment, professionals with ChatGPT finished 40% faster. In Denmark, users of AI chatbots saved on average 2.8% of their work hours. Among 244 consultants, three ways of working with AI emerged, and they built different skills. In a support centre, the newest agents gained most, and those who followed the suggestions learned from them; students who leaned on an open chatbot did worse once it was gone.

  1. 03Cover storyForty per cent faster, on paper
  2. 04The numbersFrom the test to the working week
  3. 05The modelThree ways to work with AI
  4. 06What the research saysWho gains, who learns
  5. 07How it is measuredCount more than speed
  6. 08Tool of the issueThe AI trial card
  7. 09SourcesSources and method

How to read this issue

Figure

Every figure has its source and year at the foot of its page.

Our reading

Where the editors interpret rather than the research, it says so.

Practice

The steps and the card are proposals to try, not research results.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst3

Cover story

Forty per cent faster,on paper

In an experiment published in 2023, Shakked Noy and Whitney Zhang of MIT gave 453 college-educated professionals two writing tasks from their own jobs, such as a cover letter for a grant or an email about a restructuring. Half, chosen at random, could use ChatGPT for the second. Experienced people from the same occupations graded the work without knowing who had used it.

40%less time on the task, on average
18%better grades from the evaluators
68%handed in ChatGPT's first draft unedited (working paper)

Those who had done worse on the first task gained most, so the gap between workers narrowed. The authors also named the limits: the tasks needed no knowledge of a real company or customer, the instructions were explicit, and no one checked the facts in the texts.

Our reading

The experiment timed a draft, not a decision. Where a text has to be right about a customer, the time saved counts only after someone has checked it.

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Shakked Noy & Whitney Zhang, MIT, 2023; Zach Winn, MIT News, 2023

A preregistered online experiment with 20–30-minute tasks, published in Science in July 2023. The 68% comes from the working paper of March 2023 (444 participants), not from the published article; the limits are as the authors described them to MIT News.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst4

The numbers

From the testto the working week

Anders Humlum and Emilie Vestergaard surveyed about 25,000 workers in Denmark, in 11 occupations exposed to AI chatbots, and linked the answers to registers of pay and hours. Time saved by users, as a share of their work hours, 2024:

Time saved as a share of work hours, with and without the employer's encouragement

Not encouragedEncouragedMarketing4.6%6.8%Developers3.9%6.5%Accountants0.9%2.2%Teachers0.6%1.0%

Across all 11 occupations, users saved 2.8% of their hours, and 80% put the time into other tasks. In no occupation did pay or recorded hours change significantly.

Our reading

An experiment measures a task the tool suits. A working week is mostly other tasks.

Source: Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025

Self-reported savings, turned by the authors into a share of work hours; 4 of the 11 occupations. Working paper of May 2025.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst5

The model

Three waysto work with AI

244 junior consultants at Boston Consulting Group solved the same problem with GPT-4: which of a fictional company's three brands should get the investment, in a memo of at most 500 words. Steven Randazzo, Hila Lifshitz and colleagues read all 4,975 exchanges with the AI and held 237 interviews.

01

Cyborgs · 60%

work with AI at every step, in constant back and forth; gained skill with AI and kept their domain knowledge

02

Centaurs · 14%

decide what to do and how, and give AI chosen parts; most accurate answers; deepened their domain knowledge

03

Self-automators · 27%

hand the task over in one or two prompts; fast and polished, but shallow; built neither skill

Two questions separate the three: who chooses what has to be done, and who decides how it is done. Of the self-automators, 44% accepted the AI's text without any change; the rest made only surface edits.

Our reading

“A human in the loop” can mean three different things. For a manager the question is who still does the thinking, and what each way teaches.

Sources: Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick et al., Harvard Business School, 2023

At the time, the task lay outside what GPT-4 did well, beyond the “jagged frontier” of issue 10. Shares as rounded by the authors (101% in total); the names in Albanian and German are ours. Working paper of 2025.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst6

What the research says

Who gains,who learns

Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed 5,179 support agents serving a large software firm as an AI assistant was rolled out. It suggested replies during each chat; the agents stayed responsible and could ignore it.

2 = 6months: agents with two months and AI did as well as those with over six months without it
38%of the suggestions were followed, on average

Agents who followed the most suggestions gained close to 25%, those who followed the fewest about 10%. When the system was down, close followers still worked faster than before: they had learned. Before the AI, coaching meant a short weekly session with a manager.

It can go the other way. Nearly 1,000 high-school students in Turkey who practised maths with an open GPT-4 did 48% better in practice, but 17% worse in the exam without it than students who never had it. A version that gave hints instead of answers avoided the loss.

Our reading

The same tool can coach or replace. Whether people can still do the work when it is switched off has to be tested, not assumed.

Sources: Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)

One firm, text chats only; the link between following and gains may partly reflect who chose to follow. The published version (QJE, 2025) reports the same patterns. Bastani et al. studied students, not workers.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst7

How it is measured

Count morethan speed

Speed is the easiest thing to measure. In the studies of this issue, the results also hinged on checking, on who gained, on where the time went and on what people could do without the tool. Five things to count in a trial:

  1. Time, with the checking

    From the request to the finished, checked result.

  2. Quality, judged blind

    By someone who knows the work and not which version used AI.

  3. New and experienced, apart

    An average can hide a gain for one and a loss for the other.

  4. Where the saved time goes

    Name the task it moves to, or it disappears.

  5. A day without the tool

    Can the person still do the task alone?

Hypothetical example, a month of AI drafts for customer emails
Time
9 min per email before, 6 min now with checking
Blind check
errors in 2 of 40 before, in 5 of 40 now
By experience
new staff −4 min, experienced −1 min

Faster, but the errors rose: the trial is not finished. The numbers are invented.

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)

The five checks are the editors' choice from the studies in this issue; the example is the editors'.

Management Review · No. 51 · November 2026AI
stivencatalyst.comStiven Catalyst8

Tool of the issue

The AItrial card

One task, one month, the same check as before. Write down what you will count before the trial starts, so the result cannot be read to fit a wish. Keep confidential data out of tools the company has not approved.

  1. 01The taskwhat, how often, who does it today

  2. 02Before the trialtwo weeks without AI: time per task and errors

  3. 03Who tries itnew and experienced people, counted apart

  4. 04What to checkfacts, figures, names, tone, data that must not leave the team

  5. 05Who does the thinkingwhat the person decides, what the AI does

  6. 06Result and decisiontime with checking, errors, a day without the tool; keep, change or stop

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023

A practice proposed by the editors, after Noy & Zhang, Randazzo et al. and Brynjolfsson et al.; the day without the tool follows the outages in the support centre.

Management Review · No. 51 · November 2026Sources
stivencatalyst.comStiven Catalyst9

Sources and method

Every figurehas a source.

The figures in this issue come from the sources below. The year shows how recent each one is.

  1. Experimental evidence on the productivity effects of generative artificial intelligenceShakked Noy & Whitney Zhang, Science 381(6654), 2023 · via Abstract at Stanford SCALEhttps://doi.org/10.1126/science.adh2586
  2. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, 10 March 2023)Shakked Noy & Whitney Zhang, MIT, 2023https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1_0.pdf
  3. Study finds ChatGPT boosts worker productivity for some writing tasksZach Winn, MIT News, 2023https://economics.mit.edu/news/study-finds-chatgpt-boosts-worker-productivity-some-writing-tasks
  4. Large Language Models, Small Labor Market Effects (Working Paper 2025-56, May 2025)Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025https://bfi.uchicago.edu/wp-content/uploads/2025/04/BFI_WP_2025-56-3.pdf
  5. Cyborgs, Centaurs and Self-Automators: The Three Modes of Human-GenAI Knowledge Work and Their Implications for Skilling and the Future of ExpertiseSteven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025https://www.hbs.edu/ris/Publication%20Files/26-036_e7d0e59a-904c-49f1-b610-56eb2bdfe6f9.pdf
  6. Navigating the Jagged Technological FrontierFabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick et al., Harvard Business School, 2023https://papers.ssrn.com/abstract=4573321
  7. Generative AI at WorkErik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023https://www.nber.org/papers/w31161
  8. Generative AI without guardrails can harm learning: Evidence from high school mathematicsHamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 · via Abstract at IDEAS/RePEc and Knowledge at Whartonhttps://doi.org/10.1073/pnas.2422633122
Editorial method

Each figure was checked for its year, its publisher and what exactly it measures. Where the publisher's page could not be opened, the figure was checked against independent summaries and is marked “via”. The editors' interpretation is marked “Our reading”. Figures that could not be confirmed are not in the issue.

ManagementReview

Management without theatre.

Every issue, one management question, checked against the best research.

All issues

stivencatalyst.com/magazine/management-review.html

Management Review · No. 51 · November 2026 · Stiven Catalyst

Management Review · No. 51

The figures of the issue

The charts of the printed pages, with their sources.

The numbersTime saved as a share of work hours, with and without the employer's encouragement
Not encouragedEncouragedMarketing4.6%6.8%Developers3.9%6.5%Accountants0.9%2.2%Teachers0.6%1.0%
Not encouragedEncouragedMarketing4.6%6.8%Developers3.9%6.5%Accountants0.9%2.2%Teachers0.6%1.0%

Source: Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025

The whole text Read the issue as text For reading on a small screen, searching or a screen reader. The same words, without the page design.

In this issue

Few tools have reached offices as fast as AI chatbots, and few have been studied so quickly. The results differ widely. This issue reads them side by side: what each study measured, on which task, and what a manager can take from it for a team.

In a writing experiment, professionals with ChatGPT finished 40% faster. In Denmark, users of AI chatbots saved on average 2.8% of their work hours. Among 244 consultants, three ways of working with AI emerged, and they built different skills. In a support centre, the newest agents gained most, and those who followed the suggestions learned from them; students who leaned on an open chatbot did worse once it was gone.

Stiven Janaqi, Editor

Cover story

Forty per cent faster, on paper

In an experiment published in 2023, Shakked Noy and Whitney Zhang of MIT gave 453 college-educated professionals two writing tasks from their own jobs, such as a cover letter for a grant or an email about a restructuring. Half, chosen at random, could use ChatGPT for the second. Experienced people from the same occupations graded the work without knowing who had used it.

  • 40% less time on the task, on average
  • 18% better grades from the evaluators
  • 68% handed in ChatGPT's first draft unedited (working paper)

Those who had done worse on the first task gained most, so the gap between workers narrowed. The authors also named the limits: the tasks needed no knowledge of a real company or customer, the instructions were explicit, and no one checked the facts in the texts.

Our reading

The experiment timed a draft, not a decision. Where a text has to be right about a customer, the time saved counts only after someone has checked it.

A preregistered online experiment with 20–30-minute tasks, published in Science in July 2023. The 68% comes from the working paper of March 2023 (444 participants), not from the published article; the limits are as the authors described them to MIT News.

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Shakked Noy & Whitney Zhang, MIT, 2023; Zach Winn, MIT News, 2023

The numbers

From the test to the working week

Anders Humlum and Emilie Vestergaard surveyed about 25,000 workers in Denmark, in 11 occupations exposed to AI chatbots, and linked the answers to registers of pay and hours. Time saved by users, as a share of their work hours, 2024:

Time saved as a share of work hours, with and without the employer's encouragement: Marketing: Not encouraged 4.6%, Encouraged 6.8%; Developers: Not encouraged 3.9%, Encouraged 6.5%; Accountants: Not encouraged 0.9%, Encouraged 2.2%; Teachers: Not encouraged 0.6%, Encouraged 1.0%.

Across all 11 occupations, users saved 2.8% of their hours, and 80% put the time into other tasks. In no occupation did pay or recorded hours change significantly.

Our reading

An experiment measures a task the tool suits. A working week is mostly other tasks.

Self-reported savings, turned by the authors into a share of work hours; 4 of the 11 occupations. Working paper of May 2025.

Source: Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025

The model

Three ways to work with AI

244 junior consultants at Boston Consulting Group solved the same problem with GPT-4: which of a fictional company's three brands should get the investment, in a memo of at most 500 words. Steven Randazzo, Hila Lifshitz and colleagues read all 4,975 exchanges with the AI and held 237 interviews.

  • Cyborgs · 60%. work with AI at every step, in constant back and forth; gained skill with AI and kept their domain knowledge
  • Centaurs · 14%. decide what to do and how, and give AI chosen parts; most accurate answers; deepened their domain knowledge
  • Self-automators · 27%. hand the task over in one or two prompts; fast and polished, but shallow; built neither skill

Two questions separate the three: who chooses what has to be done, and who decides how it is done. Of the self-automators, 44% accepted the AI's text without any change; the rest made only surface edits.

Our reading

“A human in the loop” can mean three different things. For a manager the question is who still does the thinking, and what each way teaches.

At the time, the task lay outside what GPT-4 did well, beyond the “jagged frontier” of issue 10. Shares as rounded by the authors (101% in total); the names in Albanian and German are ours. Working paper of 2025.

Sources: Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick et al., Harvard Business School, 2023

What the research says

Who gains, who learns

Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed 5,179 support agents serving a large software firm as an AI assistant was rolled out. It suggested replies during each chat; the agents stayed responsible and could ignore it.

  • 2 = 6 months: agents with two months and AI did as well as those with over six months without it
  • 38% of the suggestions were followed, on average

Agents who followed the most suggestions gained close to 25%, those who followed the fewest about 10%. When the system was down, close followers still worked faster than before: they had learned. Before the AI, coaching meant a short weekly session with a manager.

It can go the other way. Nearly 1,000 high-school students in Turkey who practised maths with an open GPT-4 did 48% better in practice, but 17% worse in the exam without it than students who never had it. A version that gave hints instead of answers avoided the loss.

Our reading

The same tool can coach or replace. Whether people can still do the work when it is switched off has to be tested, not assumed.

One firm, text chats only; the link between following and gains may partly reflect who chose to follow. The published version (QJE, 2025) reports the same patterns. Bastani et al. studied students, not workers.

Sources: Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)

More in the essay: How I built Front Desk Control

How it is measured

Count more than speed

Speed is the easiest thing to measure. In the studies of this issue, the results also hinged on checking, on who gained, on where the time went and on what people could do without the tool. Five things to count in a trial:

  • Time, with the checking. From the request to the finished, checked result.
  • Quality, judged blind. By someone who knows the work and not which version used AI.
  • New and experienced, apart. An average can hide a gain for one and a loss for the other.
  • Where the saved time goes. Name the task it moves to, or it disappears.
  • A day without the tool. Can the person still do the task alone?

Hypothetical example, a month of AI drafts for customer emails

  • Time: 9 min per email before, 6 min now with checking
  • Blind check: errors in 2 of 40 before, in 5 of 40 now
  • By experience: new staff −4 min, experienced −1 min

Faster, but the errors rose: the trial is not finished. The numbers are invented.

The five checks are the editors' choice from the studies in this issue; the example is the editors'.

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023; Anders Humlum & Emilie Vestergaard, Becker Friedman Institute, University of Chicago, 2025; Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı & Rei Mariman, PNAS 122(26), 2025 (via Abstract at IDEAS/RePEc and Knowledge at Wharton)

Tool of the issue

The AI trial card

One task, one month, the same check as before. Write down what you will count before the trial starts, so the result cannot be read to fit a wish. Keep confidential data out of tools the company has not approved.

  1. The task what, how often, who does it today
  2. Before the trial two weeks without AI: time per task and errors
  3. Who tries it new and experienced people, counted apart
  4. What to check facts, figures, names, tone, data that must not leave the team
  5. Who does the thinking what the person decides, what the AI does
  6. Result and decision time with checking, errors, a day without the tool; keep, change or stop

A practice proposed by the editors, after Noy & Zhang, Randazzo et al. and Brynjolfsson et al.; the day without the tool follows the outages in the support centre.

Sources: Shakked Noy & Whitney Zhang, Science 381(6654), 2023 (via Abstract at Stanford SCALE); Steven Randazzo, Hila Lifshitz, Katherine C. Kellogg, Fabrizio Dell'Acqua, Ethan Mollick, François Candelon & Karim R. Lakhani, Harvard Business School Working Paper 26-036, 2025; Erik Brynjolfsson, Danielle Li & Lindsey Raymond, NBER, 2023

Open the tool: Sigma & Control Chart

Sources and method

Every figure has a source.

The figures in this issue come from the sources below. The year shows how recent each one is.

Editorial method

Each figure was checked for its year, its publisher and what exactly it measures. Where the publisher's page could not be opened, the figure was checked against independent summaries and is marked “via”. The editors' interpretation is marked “Our reading”. Figures that could not be confirmed are not in the issue.

Management Review · Monthly edition

Read another issue

All issues