Does AI Actually Save Time? Count the Verification Tax
A Google, DeepMind and MIT study found AI saves scientists about 6.9 hours a week, but 46% of them spend over a quarter of that checking its work. How to measure the same trade in your business.

Yes, but less than the headline suggests. A September 2026 study by Google, Google DeepMind and MIT FutureTech found scientists save about 6.9 hours a week with AI. Of those who save time, 89% spend more than a tenth of it checking the output, and 46% spend more than a quarter. The authors call it a "verification tax."
What did the AI in Science study find?
The paper, "AI in Science: Early Insights," was published in September 2026 by researchers from Google, Google DeepMind, MIT FutureTech and several universities. It combines three sources: a sample of 15 million Gemini interactions, an inventory of more than 2,600 specialised scientific AI models, and a survey of 637 scientists in the United States and United Kingdom, run by More in Common between July 27 and August 11, 2026.
The survey is where the time numbers come from. Just under three quarters of respondents said AI saves them time on net in a working week, against 6% who said it costs them time. The average saving was about 6.9 hours a week. Almost 47% use some form of AI daily and another 31% weekly.
Respondents said they reinvest the saved time mainly in more research output (just under 30%), physical lab work and data collection (about 21%) and harder problems (about 19%). About 18% took it as better work life balance or fewer hours.
The authors are careful about the sample. Results are unweighted, and they note it may not be fully representative of all scientists.
What is the verification tax?
It is the share of AI's time saving that goes straight back into checking what the AI produced.
The survey asked respondents who save time with AI what percentage of that saved time is spent verifying, debugging or fact checking the output. The answers:
- 89% spend more than 10% of their saving on verification
- 46% spend more than 25% of it
So for a large group, a saving of seven hours becomes five and a quarter hours or less once checking is counted. The authors say the tax is particularly high in the life sciences, and they attribute it to the high value science places on correct results. They also flag that the link between heavier AI use and a higher verification tax is only marginally statistically significant, depending on the model specification, while other effects they measured are much more robust.
The paper draws the comparison directly: as with software engineering, people who adopt AI tools may be pushed toward auditing, debugging and quality screening.
Why does this matter outside science?
Because the same arithmetic applies to any job where a wrong answer costs more than a slow one. A contract summary, a customer refund decision, a financial figure in a board pack and a piece of production code all share the property that made verification expensive for scientists: you cannot ship it until someone has checked it.
The study also found a second effect that small businesses will recognise. About 44% of respondents said their main bottleneck had moved downstream over the past two years, into lab execution, validation or manuscript writing, and 41% said their backlog of untested hypotheses had grown. In business terms: AI makes drafts, ideas and code cheap, and the queue moves to whatever still needs a person, usually review, approval and the physical or customer facing step.
If your team adopts AI for drafting and nobody plans for review capacity, the likely outcome is more work in progress, not more finished work.
Are newer models reducing the need to check?
Vendors say so, and some of the numbers are large, but they are the vendors' own tests.
When OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, VentureBeat reported that Sol made about 50% fewer mistakes than GPT-5.6 Sol on OpenAI's internal factual error evaluations. OpenAI also reported a coding deception test where Sol's rate fell to 1.3% from 10.4%, and a test of whether the model admits a tool is broken where Sol's failure rate fell to 5.4% from 77.8%.
Those are real improvements if they hold on your work, and fewer errors should shrink the verification tax. They do not remove it. Half as many mistakes is still some mistakes, and you only find out which outputs contain them by checking.
How should a business measure whether AI is saving time?
Measure the whole loop, not the drafting step. Code4U suggests a simple two week comparison for any task you are thinking of handing to AI:
- Pick one repeatable task, such as answering a category of support ticket, writing product descriptions, or reviewing a type of pull request.
- Time the old way from start to accepted result, including the review a manager or colleague already does.
- Time the AI way the same way: prompt, generation, reading the output, fixing it, and the final review. The fixing and reviewing is the verification tax.
- Count the errors that got through, not just the time. A faster process that lets one wrong refund or broken deploy through a week may not be a saving.
- Decide the checking level by consequence. Low stakes, easy to reverse outputs can be spot checked. Anything that touches money, customers or production should be checked every time.
If the net saving is still clearly positive after step 3, you have a real case. If it is close to zero, the task may need a better tool, a narrower prompt, or no AI at all.
Code4U wrote about the related problem of supervising agents that run unattended in what breaks when AI agents run for weeks. The verification tax is the everyday, human scale version of the same issue.
FAQ
How much time does AI actually save at work?
In the Google, DeepMind and MIT survey of 637 US and UK scientists, the average saving was about 6.9 hours a week, and just under three quarters reported a net saving. Of those saving time, 46% spend more than a quarter of it checking AI output. The sample is scientists, not office workers, so treat it as an indication and measure your own tasks.
What does verification tax mean in AI?
Verification tax is the portion of time saved by AI that is then spent verifying, debugging or fact checking its output. The term comes from the September 2026 "AI in Science: Early Insights" paper by Google, Google DeepMind and MIT FutureTech, which found 89% of scientists who save time spend over a tenth of it on verification.
Can I skip checking AI output if the model is more accurate?
Not for anything with real consequences. OpenAI reports GPT-6 Sol makes about half as many factual mistakes as GPT-5.6 Sol on its internal tests, which should reduce checking time. But fewer mistakes is not zero, and you cannot tell which outputs are wrong without looking. Match the checking level to the cost of an error.
If you want help working out where AI genuinely saves your team time, that is part of the AI consulting Code4U offers.

