A method for assessing content exposure before a Microsoft 365 Copilot rollout.
If we enable Copilot today, what will it find — and should it be finding it?
This is a framework, not a tool. It describes what to assess, why each thing matters specifically for Copilot, and how to turn findings into a decision a leadership team can act on. Apply it with your own scripts, a vendor product, or by hand.
It is deliberately complementary to configuration-readiness tooling, including Microsoft's own. Configuration readiness asks whether the controls exist. This asks what is currently reachable.
Oversharing in Microsoft 365 is not new. Most tenants carry a decade of it: a link shared with "anyone" for a meeting in 2019, a Team created as public because that was the default, a guest account from a supplier relationship that ended two reorganisations ago.
None of it was urgent, because finding any of it required knowing it existed. The permissions were wrong, but the discovery cost was high enough to act as an accidental control.
Copilot removes the discovery cost.
Copilot answers questions using everything the signed-in user can already access. It does not grant new permissions — every vendor says this, and it is true. But it is also the least interesting fact about the problem. Copilot converts latent permission sprawl into an efficient retrieval interface over that sprawl. A file nobody could find is functionally different from a file surfaced in response to a plain-English question, even though the ACL is identical.
So the pre-rollout question is not "are our permissions correct" — they never have been. It is: which of our permission failures become materially exploitable the moment retrieval gets cheap?
That distinction runs through the whole framework.
Most assessment tooling produces a single composite score. That is the wrong shape for this decision, because it lets serious findings be averaged away by good hygiene elsewhere.
This framework separates findings into two classes that behave differently:
Data is reachable right now by someone who should not reach it, and Copilot will surface it.
Active exposure is disqualifying, not deductible. A tenant with anonymous links to content is not "mostly ready" because its labelling is excellent. The presence of a single class of active exposure should cap the outcome regardless of how the rest of the estate scores.
Controls that should exist and do not — but which do not, by themselves, put data in front of the wrong person today.
Governance gaps are real findings and they belong in the remediation queue. They are also survivable. A tenant can go live with governance gaps and a plan; it should not go live with active exposure and a plan.
Scoring implication: governance gaps deduct. Active exposure caps. Those are different operations and conflating them is the most common flaw in readiness reporting.
Four domains. Each maps to published CIS Microsoft 365 Foundations Benchmark controls, so findings are defensible to an auditor rather than resting on one consultant's opinion.
| Domain | What you are looking for | Why it matters for Copilot |
|---|---|---|
| Sharing surface | Anonymous links, organisation-wide links, link expiry policy | Anonymous links have no authentication gate. Copilot cannot security-trim what is already public |
| External identity | Guest accounts, guest licensing, stale guest access, access review coverage | A licensed guest can query your search index today. An unlicensed one shapes what your own users retrieve |
| Collaboration topology | Public Teams, ownerless Teams and sites, inheritance breaks | Public membership widens the retrieval surface for every internal user at once |
| Classification & control | Sensitivity label coverage, DLP policy scope, audit log availability | Labels are what let you constrain Copilot later. Without coverage you have no lever to pull |
Detailed guidance for each: docs/assessment-domains.md
Two rules that matter more than they sound.
Read-only. A readiness assessment that modifies the tenant is not an assessment. Every check should be achievable with read scopes, and the report should be safe to run during business hours without a change ticket.
Mark what you could not determine. Some controls genuinely cannot be evaluated from read-only data. The honest answer is MANUAL, carried through to the report as an explicit gap, not silently scored as a pass or a fail. A framework that always produces a number is hiding something.
Three states are not enough. Use four: PASS, PARTIAL, FAIL, MANUAL.
Leadership does not need a score. It needs a verdict it can act on, and the reasoning behind it.
| Verdict | Meaning |
|---|---|
| Go | No active exposure. Governance gaps, if any, are tracked with owners and dates |
| Conditional | No active exposure, but gaps significant enough that rollout should be staged or scoped to a pilot population |
| No-Go | Active exposure present. Remediate before enabling, not alongside |
A verdict is only useful if it is falsifiable. Every one should be traceable to the specific findings that produced it, and reversible by fixing them — if a reader cannot see what would change the answer, you have produced an opinion, not an assessment.
Full model: docs/verdict-model.md
The remediation sequence is not the same as the severity ranking, and getting this wrong wastes the first month.
- Cut active exposure — anonymous and organisation-wide links first. Fast, high impact, visible
- Close the external door — deprovision stale guests, then constrain what remains
- Fix topology — public Teams to private, assign owners to ownerless workspaces
- Then classify — labelling is the slowest and most political work, and it does not reduce today's exposure
Classification is where most programmes start, because it feels strategic. It is the right long-term investment and the wrong first move: it takes months, requires business buy-in, and does not close a single open link.
Sequence: docs/remediation-sequence.md
docs/worked-example.md— start here. A synthetic assessment end to end: findings, the arguments they caused, verdict, and what changed it three weeks laterCHECKLIST.md— practical assessment checklist, run it against your own tenantdocs/— reasoning behind each domain, the verdict model, remediation order
The framework is deliberately implementation-agnostic. The weightings and thresholds you apply are a calibration decision that depends on your sector, your regulatory position and your risk appetite — this document tells you what to weigh and why, not what numbers to use.
Microsoft publishes m365-copilot-automated-readiness-assessment — an open-source, API-driven tool covering seven service areas: M365 licensing, Entra, Defender, Purview, Power Platform, Copilot Studio and Agent 365.
It is good, it is free, and you should run it. It answers a question this framework does not:
Is the tenant configured for Copilot? Licences assigned, MFA covered, DLP policies present, Defender active, conditional access in place.
This framework answers a different one:
What will Copilot actually surface from our content, and should it?
The two do not overlap much. Microsoft's tool assesses configuration; this assesses content exposure. A tenant can pass every configuration check — labels published, DLP configured, MFA universal — and still have four anonymous links pointing at the board pack. Configuration readiness tells you the controls exist. It does not tell you what is currently reachable.
Specifically, the areas below sit outside the service areas that tool covers:
| Not covered by configuration assessment | Why it matters |
|---|---|
| SharePoint permissions and inheritance breaks | Where the actual reachability lives |
| Anonymous and organisation-wide sharing links | No authentication gate, cannot be security-trimmed |
| Teams membership and public Teams | Content in every internal user's retrieval scope |
| OneDrive sharing | Individually shared content, rarely governed |
| Guest access at content level | Licensing determines whether a guest can query the index |
Run both. Configuration assessment first — it is faster and it catches prerequisites. Then this, because a tenant that is configured correctly and exposed anyway is the common case, not the exception.
This covers retrieval exposure — what Copilot can surface from existing content. It does not cover prompt injection, agent permissioning, plugin and connector governance, or data residency. Those are real problems and they are different problems.
It assumes Microsoft 365 Copilot over SharePoint Online, OneDrive and Teams. The reasoning largely transfers to other retrieval-augmented assistants over enterprise content; the control mappings do not.
Disagreement is welcome, particularly on the exposure-versus-governance split — that boundary is the part of this most worth arguing about. Open an issue with the reasoning.
Charlie Delmotte — hybrid cloud and security architect, Geneva. I run Microsoft 365 security and identity architecture across a portfolio of tenants, and this framework is the shape my own assessments have settled into after doing them several times and getting the order wrong first.
The method is published. The tooling and calibration behind it are not, which is the honest position rather than a coy one: what is here is what I think is true about the problem, and it should be useful whether or not you ever talk to me.
Corrections and disagreement welcome — see CONTRIBUTING.md.
CC BY 4.0. Use it, adapt it, build on it commercially. Attribution required.