Buying Microsoft 365 Copilot licences takes an afternoon. Getting a pilot group to trust what Copilot tells them takes months, and most of that time has nothing to do with the model. When enterprise pilots stall, it is almost always for one of four reasons: the agent can see things it shouldn’t, it reads the wrong version of the truth, nobody can prove it’s accurate, or nobody owns adoption once the launch email goes out.

This post walks through each of the four, and the governance work that fixes it, in the order we tackle it on a Copilot Studio engagement.

TL;DR

  • Copilot respects your existing permissions. That is the good news and the bad news: it will surface anything a user can already reach, including content that was overshared years ago.
  • Fix access before the agent goes live. Map who can see what in the pilot sources, tighten the worst of it, and use the controls Microsoft already provides.
  • Answer quality is a content problem first. Taxonomy, metadata, and a clear “approved version” rule do more for accuracy than prompt tuning.
  • Build an evaluation set with your subject-matter experts before launch, and score every change against it. Check citations, not just answers.
  • Treat adoption as a workstream with an owner, champions, and a benefit baseline set in discovery.

Why does Copilot surface documents people shouldn’t see?

Because it was never supposed to decide who sees what. Microsoft 365 Copilot works within the permissions each user already has. If a user can open a file, Copilot can use it to answer that user’s question.

In most large tenants, permissions have drifted for years. Sites shared with “everyone except external users” for a one-off project. Folders shared by link and forgotten. Teams that outlived the reorg that created them. Before Copilot, that sprawl was mostly invisible, because nobody went looking. A copilot goes looking on every question.

So the first workstream on any engagement is a permission review of the sources in scope. We inventory the sites, libraries, and Teams the pilot will draw on, flag broad or anonymous sharing, and agree with the owners what should change. Microsoft provides tooling for much of this, including SharePoint Advanced Management reports on oversharing, Restricted SharePoint Search to limit which sites Copilot draws on during a pilot, and Microsoft Purview sensitivity labels and data loss prevention policies. The tools matter less than the decision: someone has to own “who should see this” for every source before the agent reads it.

Why are the answers accurate in the demo and wrong in the pilot?

Because the demo used five clean documents and the pilot uses five thousand real ones.

A knowledge agent is only as good as what it retrieves. Real repositories hold drafts next to finals, last year’s plan next to this year’s, and three copies of the same tracker with different numbers. The model will faithfully summarise whichever one it finds first.

The fix is unglamorous content governance:

  • A taxonomy and metadata model so documents carry their type, owner, status, and period.
  • An “approved version” rule so the agent prefers final, current, owner-approved content, and says so when it can’t find any.
  • Quality gates on sources that flag stale or duplicated content for the owner to fix. The consultant flags it, the business remediates it.

This is where most of the accuracy comes from. Prompt tuning helps at the margin. Clean sources change the result.

How do you prove a Copilot agent is accurate?

With an evaluation set, built before launch and kept after it.

Sit down with two or three subject-matter experts and write the questions the pilot users will actually ask, with the correct answer and the document that answer should come from. A few hundred well-chosen questions is usually enough to start. Then score the agent against them: is the answer correct, is the citation real and relevant, and does the agent decline when the answer isn’t in the sources?

That last check matters. An agent that confidently answers questions outside its sources is worse than one that says “I couldn’t find that.” We also include a few questions the user should not be able to get answered, to confirm the permission work held.

Every configuration change, new source, or model update gets re-scored against the same set. That turns “it seems better” into a number your sponsor can use at a go/no-go gate. We go deeper on the method in AI evals: beyond vibes-based QA.

Why do Copilot pilots lose users after launch?

Because a pilot without an owner decays. The first week has curiosity. By week six, people have gone back to the habits that worked before, unless the agent has become part of how a recurring task gets done.

Adoption needs the same planning as the build:

  • A benefit baseline set in discovery. How long does the weekly brief take today? How many questions does the team field? Without a before, there is no after.
  • Champions in the pilot group who get early access, give feedback, and show colleagues what works.
  • Role-based training on the two or three tasks the agent does well, not a tour of every feature.
  • A named owner who reviews usage and the evaluation score every fortnight.

How should an enterprise Copilot engagement be phased?

In short phases with a decision at the end of each, so nobody commits to the whole programme on faith. A shape that works for one business function:

PhaseTypical lengthWhat you have at the gate
Discover and prepare~3 weeksConfirmed use case, source inventory, permission map, success metrics, firm quote
Knowledge agent~6 weeksCitation-backed agent, evaluation set, UAT with pilot users
Insight agent~4 weeksRecurring briefs and KPI context across plans and action logs
Automate and govern~4 weeksFlows with human review, monitoring and evals live, champions trained

Each gate is a real decision: continue, adjust, or stop with working software. That structure is what lets a sponsor say yes to the first phase without betting the budget on the last one.

FAQ

Does Microsoft 365 Copilot ignore our permissions? No. It works within each user’s existing permissions. The risk is that those permissions are broader than anyone intended, so the governance work is about fixing access, not about Copilot bypassing it.

Do we need Copilot Studio, or is Microsoft 365 Copilot enough? Microsoft 365 Copilot covers general assistance in the apps. Copilot Studio is how you build agents scoped to a specific function, knowledge base, and set of actions. Most enterprise pilots with a defined use case need both.

How big should the evaluation set be? Start with a few hundred questions that represent real usage, written with subject-matter experts. Coverage of the important question types matters more than raw volume, and the set should grow as users find new edge cases.

Should we fine-tune a model on our documents? Rarely for this kind of work. Grounding a model on well-governed sources, with citations, is usually more accurate, easier to update, and simpler to govern than a custom-trained model.

Who should own the agent after launch? A named product owner in the business function, supported by IT for the platform. The owner reviews usage and evaluation results on a regular cadence and decides what changes next.

Planning a Copilot pilot and want a second opinion on scope, governance, or phasing? See how we run enterprise AI engagements, or book a discovery workshop and leave with a plan you can take to your sponsor.