Practical working checklist. Read-only throughout — nothing here modifies a tenant.
Record each item as PASS / PARTIAL / FAIL / MANUAL, with evidence. An item without evidence is an opinion.
Class each finding as active exposure (caps the verdict) or governance gap (deducts). See docs/verdict-model.md.
- Read scopes only — confirm no write permission was consented
- Note the CIS Microsoft 365 Foundations Benchmark version you are assessing against
- Record tenant name, date, and who ran it — the trend line depends on consistency
- Agree in advance who receives the report and at what classification
- Anonymous ("Anyone") links — count, age, target locations → likely active exposure
- Organisation-wide links — count, and whether any target sensitive locations → active exposure at volume
- Tenant-level external sharing setting
- Site-level sharing settings where they diverge from tenant default
- Link expiry configured and enforced
- Inheritance breaks — count and concentration
Key question: is there anything reachable without authentication?
- Guest account inventory — count, creation date, last sign-in, sponsoring domain
- Any guest holding a Copilot licence in this tenant → active exposure
- Stale guests — no sign-in 90+ days
- Guests whose sponsor has left the organisation
- Access reviews — do they exist, do they cover guests, do they complete
- External sharing domain allow/block list in force
Key question: can any external identity query our search index today?
- Public Teams — count, and what they contain
- Ownerless Teams and sites
- Sites with a single owner who has left
- Nested group membership granting unintended breadth
- Drift between intended and actual site permission structure
Key question: how much content sits in every internal user's retrieval scope?
- Sensitivity label coverage — percentage, and coverage where it matters
- Label policy scope — which users, which locations
- DLP policies — coverage of SharePoint, OneDrive, Teams
- Unified audit log enabled, with retention matched to obligations
- Whether labels are applied correctly → usually
MANUAL
Key question: if we needed to exclude content from Copilot tomorrow, could we?
- Every finding classed as active exposure or governance gap
- Every finding traceable to specific evidence — site, link, account, policy
- Count of
MANUALitems stated explicitly in the report - Verdict follows from the findings by a stated rule, not by judgement call
- Remediation queue ordered by exposure reduction per unit of time, not by severity
- A reader can answer: what exactly would change this verdict?
Two audiences, two documents. Do not try to serve both with one.
Technical — full findings, evidence, control mappings, remediation queue with priorities.
Leadership — the verdict, the reasoning in plain language, what it costs to change it, and a date. No control IDs, no counts they cannot act on.
If the leadership version cannot be read in three minutes and acted on, it will not be.
- Same scope, same benchmark version, same calibration
- Note anything that changed in method — a trend built on a moved goalpost is worse than no trend
- Track direction, not just the current number
- Watch for regression: new links, new guests, new public Teams