/ THE SHORT ANSWER
See the method. Keep the context.
The visual companion

Credit: Original dotSuper research graphic based on cited primary sources.
Reuse: dotSuper original artwork. No third-party product image or logo reproduced.
Read the diagram: Permission and approval design determines what an agent can do in your business.
READ → RECOMMEND
APPROVE → ACT
Record the decision
Thumbnail credit and reuse
Credit: Original dotSuper research graphic.
Reuse: dotSuper original artwork. No third-party product image or logo reproduced.
- 01Model safeguards do not define your business permissions.
- 02Separate reading, recommending and acting.
- 03Test contradictory and adversarial inputs before rollout.
/ dotSuper point of view
Define the job and data boundary, give the agent only the tools it needs, require approval for consequential actions and record outcomes. New model safeguards help, but the workflow still needs accountable owners.
Why this matters now
OpenAI says GPT-6 Astra reaches its Critical cybersecurity capability level and describes stricter protections and monitoring.
These are vendor reports about their systems.
For a business, the practical question is how an agent behaves inside your accounts and processes.
An agent that can read documents and draft a response has a different risk profile from one that can change records, issue payments or contact customers.
Design its authority around the exact job.
Draw the boundary in ordinary language
For an invoice workflow, it might read the purchase order and invoice, flag mismatches and prepare a recommendation.
A named employee should approve vendor changes and payments.
This makes the system easier to explain, test and stop when something unexpected happens.
Give tools narrow permissions rather than broad administrator access.
Keep source records and decisions linked, so a reviewer can trace the recommendation back to evidence.
Test mistakes, not just happy paths
Check whether the agent treats retrieved content as evidence or as an instruction.
Record the paths that led to a wrong action, even if the final answer looked polished.
Anthropic reports improved resistance to prompt injection for Opus 5.5.
OpenAI reports added monitoring and robustness for Astra while acknowledging limits in monitoring model reasoning.
These are reasons to examine the whole workflow and keep human controls in place.
A small release gate
Start with observing and proposing.
Let it act only after the proposed action has been reviewed in the context that matters to your organisation.
Watch real outcomes after launch: unwanted actions, review overrides, source mismatches and time spent fixing results.
Update the permissions and task design when the pattern changes.
Where dotSuper can help
What this page cannot conclude
- 01Current as of 23 September 2026. Availability and pricing can change. Vendor benchmark results are attributed to their publishers.
- 02Examples describe a proposed evaluation, not a dotSuper customer result or independent model benchmark.
- 03Choose data handling, permissions and human review to fit the actual work and jurisdiction.
Sources
- 01Claude Opus 5.5 launch and evaluation notesAnthropic · accessed Sep 23, 2026
- 02GPT-6 Astra safety overviewOpenAI · accessed Sep 23, 2026
Our editorial standard · Found an error? Send a correction with its source.
/ CITE OR SHARE THIS GUIDE
Make the evidence easy to verify.
When you reference this guide, link to its canonical URL. That gives readers one stable place for the evidence, limitations and future updates.
dotSuper Research Desk. (September 23, 2026). AI Agents Need Clear Boundaries. dotSuper. https://dotsuper.net/feeds/market-intelligence/frontier-ai-agent-safety-2026