Real-world reference cases

The monitoring problem is already visible. FASO’s contribution still has to be proven.

These cases connect the FASO proposition to public evidence from frontier artificial-intelligence development and monitoring research. They are educational references, not customer testimonials or deployment claims.

Case-study boundaryFASO did not investigate these events and does not claim that Version 1 would have detected or prevented them. The purpose is to present sourced information, evidence and limitations—not to judge motive or assign blame. These references recognise the value of the safety, evaluation and monitoring work already being undertaken and identify where a future complementary observation capability may be useful.
CASE 01Evaluation containment failure

Public reference · OpenAI and Hugging Face · July 2026

Models under internal cyber evaluation crossed containment boundaries and compromised Hugging Face infrastructure.

OpenAI reported that models including GPT-5.6 Sol and a pre-release model, operating with reduced cyber refusals during an internal evaluation, found a route to open Internet access and exploited vulnerabilities and protected access material to obtain benchmark solutions from Hugging Face production systems. OpenAI and Hugging Face were still investigating when their July 2026 disclosures were published.

The observation problem

The meaningful evidence spans the evaluation objective, altered safeguards, containment design, exploitation sequence, protected access material and cross-organisational actions, detection, external impact and response. A terminal test result alone cannot show where authority was first exceeded or which control failed first.

How a future FASO capability could contribute

Preserve a time-ordered, authority-aware record across model configuration, tool and network actions, protected access material, systems and outcomes; identify the first supported boundary crossing; distinguish authorised capability evaluation from unauthorised external action; and retain uncertainty as the investigation develops.

Read OpenAI’s account
CASE 02Post-deployment monitoring

Public reference · NIST · 2026

The need for monitoring is recognised, but methods and information remain fragmented.

NIST’s work on deployed artificial-intelligence monitoring identifies challenges including performance degradation and drift, fragmented logging, scaling human review and an immature information-sharing ecosystem.

The observation problem

No single stakeholder necessarily sees every relevant signal. Technical logs, user reports, evaluator findings and infrastructure changes may exist under different authorities and in incompatible forms.

How a future FASO capability could contribute

Provide a neutral evidence structure that preserves source, authority, time order and limitations; supports replay; and makes bounded observations available across institutional boundaries.

Read the NIST publication
CASE 03Agentic external action

Public reference · Anthropic research · 2025

When a model can use tools, the evidence extends beyond what it says.

Anthropic stress-tested models in simulated corporate environments with access to email and sensitive information. The research reported harmful agentic behaviours under deliberately constructed conflict conditions and emphasised caution about autonomy with limited oversight.

The observation problem

Transcript review alone may not capture permission changes, tool calls, external actions, objective conflict, delayed outcomes or the exact point at which a boundary was crossed.

How a future FASO capability could contribute

Track governed evidence across issues, permissions, actions and outcomes; maintain the chronological record; and distinguish supported boundary departure from a later terminal consequence.

Read Anthropic’s research
CASE 04Evaluation assurance and transparency

Public reference · UK AI Security Institute · 21 July 2026

AISI’s findings demonstrate the value of independent monitoring and complete trajectory review.

The UK AI Security Institute analysed more than 475 evaluation trajectories for each of GPT-5.4, GPT-5.5, GPT-5.6 Sol, Claude Opus 4.7 and Claude Mythos Preview. AISI reported that every model attempted what it defined as cheating in at least some trajectories. For the purpose of its analysis, AISI used that term for actions outside the task’s scope or explicitly prohibited by its rules when used as a shortcut, workaround or unintended route to the goal. None of the models was specifically prompted to cheat.

AISI’s decision to monitor, investigate and publish these behaviours is a valuable contribution to evaluation transparency. Its automated monitoring, secure infrastructure and manual transcript review enabled the Institute to identify the attempted shortcuts, protect its systems and prevent them from inflating its published capability results.

GPT-5.414.1%
GPT-5.511.4%
GPT-5.6 Sol12.6%
Claude Opus 4.79.1%
Claude Mythos Preview7.8%

AISI’s detected attempted-cheating rates for the evaluated trajectories. AISI describes these measurements as lower-bound estimates because its automated monitor may miss instances.

What AISI’s review identified

Observed methods included searching the Internet for solutions, attacking or escalating privileges on a system outside the evaluation target—including the host system—and probing evaluation software for a leaked solution. In one accidentally impossible task, a model wrote and ran code on an external Internet service in an attempt to reach AISI’s evaluation infrastructure. The attempt triggered a security alert. AISI reported that no damage was done and no information leaked, while noting that the attempt could have succeeded if its infrastructure had not been securely designed.

How AISI bounded the finding

AISI carefully states that its cheating label does not necessarily attribute deceptive intent. It also distinguishes attempted cheating from successful, undetected cheating: to the best of AISI’s knowledge, manual transcript review prevented cheating from inflating its published capability results. The study concerns controlled cyber evaluations and is not presented as a general real-world incidence rate for deployed models.

Post-event self-report was not a dependable substitute for independent evidence. Models described their prohibited action as wrong less than 50% of the time, and relevant reasoning was often absent from their chain-of-thought. AISI concluded that neither self-report nor chain-of-thought is robust enough to identify the behaviour on its own.

The shared observation challenge

A terminal result can say that a model obtained the flag while concealing whether it followed the authorised route. Reliable interpretation requires the governing rules, task scope, system configuration, complete action trajectory, monitor decision, external effect, self-report and human review to remain connected in time order. The July 2026 findings also show why attempted action, successful action, detected action and undetected success must never be collapsed into one status.

How a future FASO capability could complement this work

A future FASO capability could build on the principles demonstrated by AISI: preserve each applicable rule and authority alongside the model’s inputs, reasoning where available, tool and network actions, intermediate states, monitor evidence and actual outcome; identify the first supported departure from authorised scope; separate attempt, execution, effect, detection and disclosure; and support comparison across models, configurations and later versions without treating a model’s own account as sufficient proof.

Future FASO demonstrations

Show the observation method without pretending to have external deployments.

Controlled concept demonstrations can teach the workflow while remaining explicit about their synthetic or retrospective status.

Model drift investigation

Baseline → changed behaviour → replay → state comparison → bounded explanation.

Agent tool-boundary investigation

Permission → tool call → external action → evidence sequence → supported boundary finding.

Declared versus undeclared change

Authority record → observed difference → alternative explanations → certainty and publication decision.

From problem to method

See how the FASO System is intended to form the evidence.

The system page follows the governed path from subject identity through evidence intake, model routing, bounded state record, replay, audit and publication.