The Copilot default: Why limiting Government staff to one AI tool is a missed opportunity

Microsoft Copilot Suite graphic

Ethan Mollick, the Wharton professor whose writing on AI adoption is widely read across the sector, uses the term “secret cyborgs” for what happens when organisations restrict AI access too tightly. People quietly use AI to do their jobs better, but don’t tell anyone because they’re worried about getting in trouble.

This is almost certainly happening in NZ government agencies right now. Public servants who use ChatGPT or Claude at home know what these tools can do. They know the free tier of Copilot doesn’t compare. When the gap between what’s available and what’s allowed gets too wide, people find workarounds. Those workarounds happen in silence, on personal devices, without governance, without shared learning, and without any of the enterprise protections agencies are trying to maintain.

From a security perspective, that’s worse than giving people properly managed access to a range of tools.

There’s a common belief that because Copilot sits within the Microsoft ecosystem, it’s inherently safer than alternatives. The recent Corrections story shows where that thinking breaks down.

Corrections staff were found using Copilot to draft Extended Supervision Order reports containing personal information, in breach of policy. The tool they were using was the free Copilot Chat bundled with M365: a standalone chat function, not integrated into agency systems, with no enterprise-grade data protections. Being “approved” didn’t prevent misuse. Being inside the Microsoft ecosystem didn’t make the data handling appropriate for sensitive casework.

The paid enterprise tiers of ChatGPT, Claude and Gemini all offer data processing agreements, commitments not to train on inputs, and audit capabilities. In several cases they match or exceed what the free Copilot tier provides.

There’s also a misunderstanding about what “security risk” looks like for most government work. The worry tends to be about someone putting classified information into ChatGPT. But the vast majority of public service work isn’t classified. It’s drafting policy advice, writing briefings, preparing reports, analysing publicly available data, planning projects. For this kind of work, paid enterprise tiers offer solid protections.

NZ Department of Corrections sign

Where real sensitivity exists, tighter controls make sense. No one should be putting names, NHI numbers, or case details into a general-purpose AI tool without proper agreements. But treating all government work as if it carries the same sensitivity as an Extended Supervision Order report locks most staff out of tools that could genuinely help with routine tasks.

A better approach would classify the work, not just the tool.

  1. What can safely go into an enterprise AI?
  2. What needs anonymisation first?
  3. What shouldn’t go anywhere near these systems?

The proof-of-concept trap

Teams wanting to trial AI tools are increasingly being asked to build a proof-of-concept business case to justify the cost. For a major platform investment, that makes sense. For giving a small team access to a paid AI subscription to test whether it helps with literature reviews or policy analysis, the overhead of building the case often outweighs the experiment itself.

Meanwhile, agencies invest in Copilot licences through existing Microsoft contracts without the same scrutiny, because it’s bundled.

Scaling AI subscriptions across thousands of public servants adds up. The question worth asking is whether defaulting to one bundled tool is the best use of what agencies are already spending.

What most people aren’t using AI for yet

Most public servants with AI access are using it for small tasks. Tidying an email. Summarising a meeting. Reformatting text.

There’s a lot more these tools can do. A policy analyst could talk through a new project for ten minutes, upload a project management template, and come away with a draft plan, timeline, stakeholder brief, and talking points. A team lead could upload a messy spreadsheet and get a structured analysis back in minutes. A service designer could run a research scan across international evidence on a policy question and get a synthesised summary with sources, replacing days of desk research.

People are doing this in the private sector and in their own time. But if your only work tool is a free-tier chatbot, you don’t discover what’s possible. You don’t build judgement about when AI output is reliable and when it needs checking. You don’t learn which tool works best for which task.

Mollick’s consistent finding is that the gap between casual and power users isn’t about prompting skill. It’s about knowing what features exist and applying them to real work. You don’t develop that knowledge on a locked-down free tier.

Parts of government are already showing what happens when people have the space to experiment. DocRef, a collaboration between the Government Chief Digital Office and Syncopate Lab, has turned NZ legislation into structured, machine-readable data with unique identifiers for every element, enabling version comparison and integration with AI systems. The team recently used AI in drafting the new NZ API Standard for GCDO and were transparent about how they did it and what risks they mitigated. That work came from trusting people with domain knowledge to choose the right tools for the job.

What we’d like to see

We’re not arguing for a free-for-all. Data classification, consent and appropriate use policies all matter. But the current default is producing worse outcomes than a more considered approach would.

The GCDO’s Responsible AI Guidance for the Public Service (February 2025) doesn’t require agencies to lock down to one vendor. It asks them to assess risks, apply procurement and security standards and make informed choices, weighing cost, functionality, privacy and security across providers. The Algorithm Impact Assessment toolkit on data.govt.nz can support that process. Most of the permission structures agencies need are already in place. The gap is in using them.

There are practical things different people in the system can do. IT and security leads can run assessments across multiple providers rather than defaulting to what’s bundled.

Managers can set aside regular time, even an hour a fortnight, for teams to experiment with AI on real, low-stakes work and share what they find. A standing show-and-tell or a shared channel can turn quiet workarounds into shared knowledge.

Agency leaders can take up the GCDO’s AI masterclasses, set aside small budgets for paid tiers, and make it clear that experimenting within the guidance is encouraged. And for practitioners: try the free tiers of different tools to understand the differences, follow sources like Mollick’s One Useful Thing to stay current, and when you find something useful, share it. The “secret cyborg” problem only exists when there’s no safe space to talk about what you’re learning.

The GCDO is running courses and the guidance frameworks exist. What will separate the agencies that build real AI capability from those that don’t is whether they actually use the flexibility they already have.


Written by Jonnie Haddon
GM Government Innovation