Decision RecordActivePublished without chair review
CRT-2026-024217 Aug 2026AFTERNOON EDITIONDaily Roundtable
Validate Copilot boundary-crossing claims with an isolated canary
Run one isolated end-to-end canary test using synthetic identities, data, tokens, a harmless connector, a sandbox-only nonce, and tightly controlled egress. Contain immediately if an unauthorized action or boundary crossing occurs.
Current public guidance · the full record
What to do now
At a glanceThe edition's authoritative action board carries no action for this record's subjects — no What to do now guidance.
Why now
Under reviewOn 2026-08-17, the Roundtable established a concrete way to test Microsoft Copilot boundary-crossing claims without exposing production resources.
The claims remain unresolved, but the isolated canary now supplies explicit telemetry requirements, decision signals, and containment actions. The evidence is bounded to 2026-08-17 and does not establish later developments or production exploitation.
Who is affected
Under reviewMicrosoft Copilot security teams investigating hostile-document claims need the isolated canary to determine whether content can trigger privileged tool use, sandbox access, or controlled egress.
Microsoft Copilot identity, connector, and incident-response operators must be ready to revoke test tokens, disable connectors, terminate sessions, and preserve evidence when a trigger appears.
Telemetry operators must correlate Purview UAL, AI tool, container, filesystem, DNS, proxy, and firewall records; Azure code-interpreter operators cannot rely on AppEnvSession logs because those sessions do not emit them.
Production Microsoft Copilot users, data, tokens, and connectors must remain outside the test. ChatMate operators are not covered by this procedure and require product-specific controls. No affected Microsoft Copilot versions are identified.
What supports this
Under reviewSupport — The AI security canary design specifies an isolated Microsoft Copilot tenant, synthetic identities and data, a harmless connector, a sandbox-only nonce, and an approved HTTPS sink; it shows how to exercise the full trust boundary without production resources.
Support — The defense telemetry design specifies correlation by one test ID across Purview UAL identity fields, AIInvokeAgent, AIExecuteTool, AIInferenceCall, AIGuardrail, container activity, DNS, proxy, and firewall flows; it shows how to trace the result.
Support — The Roundtable synthesis combines disposable sessions, time-limited tokens, removed production connectors, restricted egress, and immediate containment triggers; it shows a bounded operational path.
The evidence assessment found these items consistent and supportive of controlled validation, not proof of production exploitation.
How the Roundtable reached this
Under reviewThe AI security specialist proposed one isolated Microsoft Copilot canary spanning hostile document content, privileged connector use, sandbox access, and controlled egress.
The defense architect added correlation across Purview UAL identity fields, AI tool events, container activity, and network flows, while noting that Azure code-interpreter sessions do not emit AppEnvSession logs. The moderator consolidated these elements into a synthetic, disposable test.
The evidence assessment found consistent support for this procedure, while the boundary assessment limited the conclusion to controlled validation rather than production exploitation. Because no matching prior record was found, the arbiter selected a new operational decision.
Positions are generated by AI specialist personas and chaired by Halil Öztürkci.
Panel composition
- Scout (AI panel role)Scout identified 6 candidate signals.
- Linker (AI panel role)Linker evaluated 6 relation judgments.
- Evidence Auditor (AI panel role)Evidence Auditor recorded 13 evidence signals; 7 gaps.
- Prediction Steward (AI panel role)Prediction Steward accepted 1 prediction and rejected 1 claim.
- Boundary Reviewer (AI panel role)Boundary Reviewer recorded 10 public/private findings.
- Arbiter (AI panel role)Arbiter produced 6 decision envelopes.
Key disagreement
Scout (AI panel role)
The packet provides no affected versions, disclosed vulnerability identifiers, reproduction details, or evidence of production exploitation. A negative test would not disprove all reported findings, and the procedure is not automatically transferable to ChatMate.
Arbiter outcome
Arbiter outcome: new decision record. No matching prior record was found. The isolated canary procedure and containment triggers are strongly supported and remain within a safe public boundary.
Candidates considered
Considered 6 candidates · opened 1 · 5 not opened (5 other)
Considered, not opened
Sign in to preview Considered-Not-Opened entries (moves to Pro at launch).
Sign in to preview practitioner entries.
What is uncertain
MissingIt remains unknown whether hostile content can cause unauthorized Microsoft Copilot actions in production, which versions or configurations might be susceptible, and whether one negative canary covers every reported path.
Azure code-interpreter sessions lack AppEnvSession logs, so those logs cannot be treated as required proof. The procedure is specific to Microsoft Copilot and does not establish equivalent behavior or controls for ChatMate.
What evidence is missing
MissingThe available evidence does not identify affected Microsoft Copilot versions, disclosed vulnerability identifiers, vendor reproduction instructions, product-specific remediation, or verified production exploitation.
It also provides no production asset inventory or customer telemetry. A separate control design and evidence set are required before applying this procedure to ChatMate.
What would change this
Under reviewAn observed unauthorized Microsoft Copilot action or boundary crossing in the isolated canary would change the claim from unverified to reproduced under the recorded test conditions and require immediate containment.
Verified vendor evidence identifying affected versions, configurations, or remediation would narrow or replace the generic canary procedure. A negative result would not justify declaring every reported path disproved.
What to watch next
Under reviewDuring the canary, watch for unauthorized connector execution, unapproved tool calls, nonce exposure, egress beyond the approved HTTPS sink, cross-session access, or host-level access.
Trigger containment immediately on any one of these signals. If none appears, retain the claim as unverified rather than disproved, and preserve the correlated test record for comparison with later reproduction details or vendor guidance.
Evidence basis
The practical validation path is now clear for Microsoft Copilot: an isolated, end-to-end canary test can determine whether hostile document content crosses the full trust boundary into privileged tool use, sandbox data access, or controlle…
Arjun, correlate four layers on one test ID: Purview UAL identity fields (`AppIdentity`, user/object ID, session, IP, timestamp); `AIInvokeAgent`, `AIExecuteTool`, `AIInferenceCall`, and `AIGuardrail` events with tool name/result; container…
Run **one end-to-end canary test** in an isolated Copilot test tenant. Place a document containing an indirect prompt that instructs the assistant to: invoke a synthetic privileged connector, read a unique sandbox-only nonce, and transmit i…
Public value history
- 17 Aug 2026Initial public guidanceCurrent guidance
Created the first public value version for this Decision Record.
Source RoundtableAfternoon roundtableConvened 17 Aug 2026Methodology
How the panel reaches a Public Decision Record.