Copilot and AI app

Two British government departments ran a Microsoft Copilot trial at almost exactly the same time, on almost exactly the same software, and published almost opposite conclusions.

HMRC came away positive. The Department for Business and Trade came away unconvinced enough to write it down in public. Both reports are on GOV.UK, both are honest, and as far as I can tell nobody selling Copilot training in this country has mentioned either of them.

That gap is the interesting bit, so here is what they actually found.

HMRC gave three thousand people a licence, and most of them wanted to keep it

HMRC ran its phase three trial across the autumn of 2024 and published the evaluation on GOV.UK, under the title Evaluating the impact of Microsoft Copilot in HMRC.

Three thousand licences went out at random. Of the people who got one, 83% actually used it, which is a far better take-up than most enterprise software manages in its first quarter. Satisfaction landed at 7.1 out of ten across 1,364 survey responses which aint half bad.

Staff reported saving 2-3% of their working week, roughly sixty minutes, and HMRC then knocked about 20% off that figure themselves to allow for non-users and the usual optimism in survey answers. I like a report that argues against its own headline.

The number that stuck with me is 61%. That is the share who said they would be disappointed to lose their licence at the end of the trial. People do not say that about software they are tolerating.

The Department for Business and Trade found the opposite and said so plainly

DBT handed out a thousand licences over roughly the same period and published its Microsoft 365 Copilot evaluation as a PDF on GOV.UK.

Across 63 working days, the average user took 72 actions with it. That works out to 1.14 actions a day. Once a day, someone opened the most talked about product in enterprise software, asked it one thing, and closed it again.

DBT also ran observed task exercises, eleven of them in early 2025, which is where it gets uncomfortable. Eleven is a small number and I will come back to that. People using Copilot built a PowerPoint deck about seven minutes faster than people without it, and scored significantly lower on both quality and accuracy. On spreadsheet analysis, they were four and a half minutes slower, and again less accurate.

Faster and worse on slides. Slower and worse on numbers. Just over a fifth of respondents, 22%, spotted the tool inventing information during the pilot.

The conclusion DBT reached, in its own words, was that there was no evidence the time savings had led to improved productivity.

Same product, same year, opposite verdicts

HMRC DBT
Licences 3,000 1,000
Take up 83% used it 1.14 actions per user per day
Satisfaction 7.1 out of 10 Not reported that way
Would miss it 61% Not reported that way
Verdict Worth continuing No evidence of improved productivity

Two departments of the same government, running the same software, in the same months, reaching findings that point in opposite directions. Microsoft did not ship two different products. So the difference sits somewhere else.

Why has nobody in the Copilot market mentioned these?

Because almost everyone writing about Copilot in the UK sells Copilot licences, and one of these two reports is very bad for that conversation.

Search for Copilot evidence and you get vendor case studies, partner blogs and Microsoft’s own research, all of which arrive at the answer you would expect. The DBT evaluation is a thousand licences of publicly funded, independently written, deliberately unflattering evidence, sitting on GOV.UK where anyone can download it, and it barely appears anywhere.

We do not resell licences, which makes this an easy piece for us to write and an awkward one for a Microsoft partner. That is worth keeping in mind whenever you read anything about this product, including this.

There is also a genuine reason the reports get ignored, which is that neither is convenient. HMRC’s positive result comes wrapped in caveats and a self imposed 20% haircut. DBT’s negative result comes from a pilot where hardly anybody used the thing, which invites the obvious rebuttal. Nuance travels badly.

The difference is what people were taught to do with it

Read the two reports next to each other and one thing separates them more than anything else. HMRC’s evaluation names, as an area needing attention, the need for more practical and tailored training and time to learn how to use Copilot effectively. Its own authors identified the gap.

DBT’s usage data shows what happens when that gap goes unfilled. An average of one action a day is the signature of people who tried it, got a mediocre answer, and quietly went back to doing the job the way they always did.

That is not a criticism of DBT, who deserve a great deal of credit for publishing an unflattering result at all. It is a description of what an unsupported rollout looks like from the inside, and I have watched the same pattern in agencies and marketing teams that never went near Whitehall.

Buying licences is a procurement decision. Getting value out of them is a training and habit decision, and the second one rarely gets a budget line.

Roger Hurni, who has run the same Phoenix agency for nearly twenty-seven years, told me on the podcast about the enterprise AI he built for his own firm when he should have rented one off the shelf. His line was that you have to work out whether it is cheaper to rent than to build, and that it was not his core competency. That conversation with Roger is mostly about judgment beating speed, which is more or less what DBT measured on PowerPoint.

What should a marketing or comms team take from this?

Three things, and none of them is “do not bother”.

Expect roughly an hour a week per person, honestly measured, rather than the transformation the sales deck promises. HMRC’s carefully hedged 2 to 3% is the most credible number published in Britain so far, and it is a perfectly good return on thirteen or so pounds a month.

Assume the tool is faster and worse until somebody is checking the work. DBT measured that directly on PowerPoint. Speed without a checking habit produces more output of lower quality, which for a comms team is an actively bad trade.

Watch usage rather than sentiment. If your people are averaging around one action a day three months in, the rollout has failed quietly and nobody will tell you, because nobody wants to be the person who says they cannot get the AI to work.

One more caveat before anyone quotes either of these at a board meeting. Neither department proved anything universal. Both trials ran in late 2024, on a version of the product that has moved on since, with self reported time savings on one side and a small observed task exercise on the other. The observed task findings in particular rest on eleven sessions, which is enough to raise a question and nowhere near enough to answer one. Anyone quoting either report as settled evidence, in either direction, is overreaching.

What the pair of them do establish is that outcomes vary enormously between organisations running identical software, and that the variable is people rather than product. That is a more useful finding than a clean win would have been.

If your own rollout is closer to the DBT picture than the HMRC one, the fix is not another licence tier. A day of Microsoft Copilot training for marketing teams is built around exactly that problem, working on your own campaign documents rather than a demo.

HMRC wrote the sales pitch for training into its own evaluation, which I doubt was the intention. Worth reading both reports before anyone signs the renewal.

[author-box]