FASO-WH-002 · Working hypothesis

Testing can reveal change. Under feedback, it may also become part of the change.

Project FASO proposes that continual evaluation must not automatically be treated as neutral observation. Where test-derived information persists, the evaluation process may alter later system behaviour or the validity of what later tests appear to measure.

Working hypothesisFASO-WH-002 · Version 1.0 · 27 July 2026This publication proposes a testable system-level explanation. It is not a completed FASO finding and has not been independently validated.

Abstract

Can continual evaluation become part of the causal history of the system it is intended to measure?

Project FASO proposes that continual artificial-intelligence evaluation can cease to be neutral observation when evaluation exposure, evaluation data, results or resulting decisions enter a persistent technical, operational or institutional feedback channel.

Under those conditions, evaluation may change later system behaviour, reduce the continuing validity of the measurement, or alter the environment against which the system is assessed. The hypothesis does not state that testing automatically causes drift. It states that observational neutrality must be evidenced whenever a feedback route exists.

Complete system boundary

The evaluated system is larger than the model parameters.

For this hypothesis, an artificial-intelligence system includes every persistent component capable of affecting later behaviour or its interpretation. That may include the model, system instructions, conversation state, persistent memory, retrieval stores, tool state, caches, safety controls, configuration, training and feedback queues, human remediation, release decisions and the operational environment.

A test is not demonstrably observational until the evaluator can show whether evaluation-derived information entered any persistent component of the system.

A genuinely frozen and stateless model, reset to the same baseline after each independent test and isolated from later development, cannot acquire persistent parameter drift merely because a question was asked. It may exhibit stochastic variation or temporary context effects, but those must not be misclassified as persistent evaluation-induced change.

Established background

Adaptive evaluation, contamination, evaluation awareness and feedback effects are recognised problems.

Research on adaptive data analysis shows that reusing holdout information while selecting later hypotheses or changes can undermine generalisation. Natural-language-processing research documents benchmark contamination when evaluation material enters training. Separate work shows that frontier models can sometimes recognise evaluation conditions and infer what an evaluation is testing.

Performative-prediction research establishes a broader causal point: predictions and the decisions made from them can change the future distribution being measured. NIST’s work on deployed-system monitoring identifies performance drift, error propagation and human–artificial-intelligence feedback loops among continuing monitoring challenges.

These bodies of work establish relevant components. They do not by themselves validate the complete FASO hypothesis or its proposed system-wide causal record.

Project FASO hypothesis formation

The question arose from sustained system testing—not from a completed causal experiment.

During extended FASO Test Rig development and repeated governed campaigns, Project FASO considered whether continual testing should always be treated as passive measurement. Every detected fault can lead to diagnosis, repair, configuration change, new evidence and retesting. Unless those stages remain separated, the test process can become an undocumented source of later system change.

This is a founder-led research proposition formed from practical system work and established research. Project FASO has not yet produced controlled evidence showing that continual testing caused drift in an artificial-intelligence system.

Working hypothesis

Where evaluation exposure or evaluation results enter a persistent technical, operational or institutional feedback channel, continual evaluation can change the system being measured, producing behavioural drift, measurement-validity drift or both.

The hypothesis treats each test as a potential causal event. It does not presume that causation occurred. It requires a time-ordered record capable of showing whether the test was isolated, temporarily elicited different behaviour, caused persistent adaptation, altered the validity of later measurement or changed the surrounding environment.

Five-effect separation

“Drift” must not collapse materially different relationships.

Detection effect

The evaluation reveals change that already existed. The test is an observation and is not shown to be the cause.

Elicitation effect

The test prompt, context or environment temporarily produces different behaviour without proving persistent post-test change.

Adaptation effect

Evaluation-derived information persists through memory, learning, configuration, remediation, training or another retained system state.

Measurement-validity drift

The system becomes increasingly adapted to the repeated measurement, so performance on that measurement becomes less representative of general behaviour.

Performative effect

Evaluation results alter human decisions, deployment, user behaviour or the future input distribution, changing the environment later measured.

Evaluation-to-system feedback

Every persistent route must be identified before neutrality can be claimed.

Evaluation exposureResult or detected faultHuman or automated interpretationMemory, remediation or learningChanged system stateLater evaluationChanged behaviour or measurement
Technical channel

Persistent state

Memory, retrieval data, caches, model adaptation, safety controls, tool state or configuration retain evaluation-derived information.

Development channel

Repair and optimisation

Prompts, policies, code, training data or model parameters change in response to test results.

Measurement channel

Repeated exposure

The same benchmark, task form or recognisable evaluation environment is encountered repeatedly.

Institutional channel

Decision feedback

Scores influence release, deployment, access, funding, user behaviour or the future evidence population.

Proposed evaluation-exposure dose

Record how much evaluative influence the system has received.

Project FASO proposes evaluation-exposure dose as a candidate measurement of the cumulative repetition, recognisability, persistence and corrective influence of evaluation experienced by a system version.

  • Number, identity and timing of prior test exposures.
  • Whether evaluation content, scores or failure explanations were disclosed.
  • Whether interactions persisted in memory, retrieval, training or feedback systems.
  • Whether developers or automated processes changed prompts, policies, code, configuration or model state.
  • Whether later system versions inherited evaluation-derived changes.
  • Which protected observations remain genuinely unexposed.

The term is proposed for testing; no accepted unit, weighting or general threshold is claimed in this publication.

Testable predictions

The hypothesis predicts different outcomes according to feedback and persistence.

  1. Isolated baselineA frozen, stateless and reset-after-each-test condition should show no persistent evaluation-caused change beyond predetermined stochastic variation.
  2. Persistent-state divergenceA memory-enabled or adapting condition should diverge from an otherwise identical unexposed baseline when evaluation-derived information is retained.
  3. Repeated-test advantagePerformance on repeatedly exposed tests should improve faster than performance on protected unseen tests if measurement-validity drift occurs.
  4. Evaluation recognitionRecognisable evaluation conditions may elicit behaviour that differs from matched operational conditions, without that difference alone proving persistent adaptation.
  5. Feedback-strength relationshipGreater evaluation-exposure dose should predict greater tested-versus-unexposed divergence where the proposed causal channels operate.
  6. Isolation controlAir-gapped evidence, immutable baselines, protected holdouts and prohibited evaluation-to-training feedback should reduce or eliminate evaluation-induced divergence.

Falsification and competing explanations

The hypothesis must distinguish causation from discovery and ordinary variation.

The hypothesis would be weakened if tested and untested system twins do not diverge under persistent feedback conditions, if measured divergence remains within predetermined stochastic bounds, or if the same effects occur when all evaluation-derived feedback is demonstrably isolated.

Analysis must separately consider pre-existing drift, model non-determinism, ordinary software updates, changing user populations, context effects, evaluator error, test-set construction, provider-side changes, unrelated environmental change and regression caused by repair rather than testing itself.

Proposed FASO controls

Observation, diagnosis, intervention and retesting must remain separate governed stages.

Evaluation Isolation Principle: evaluation evidence must not alter the evaluated baseline until the observation has been completed, sealed and made reproducible under its recorded authority.

A complete zero-write proof should cover every persistent component capable of affecting later behaviour—not only source files. The record should include model identity and parameters where accessible, instructions, memory, retrieval, tools, configuration, caches, feedback and training queues, human remediation, release decisions and protected holdouts.

The strongest controlled design would preserve two identical starting systems: one exposed to the continual evaluation campaign and one held unexposed. Both would later receive an unseen protected assessment, with every persistent state difference recorded.

Relationship to FASO-WH-001

The hypotheses are connected but must not be collapsed.

FASO-WH-001 concerns possible loss or mutation of governing accuracy-first control during extended work. FASO-WH-002 concerns possible system and measurement change caused by evaluation-derived feedback.

A future investigation may test interaction between them: additional testing can increase context, corrections and intermediate objectives; instruction attenuation may contribute to apparent failure; repair then changes the system; and repeated familiar testing may become less representative. This possible compound loop is not established by either working hypothesis.

Selected primary research

Established work to which the system-level hypothesis must remain connected.

Project FASO has not identified primary work in this selected review that directly combines the five-effect separation, evaluation-exposure dose, complete technical-and-institutional feedback map, system-wide zero-write proof and tested-versus-unexposed counterfactual record proposed here. This is not a claim that no such prior work exists. A formal systematic literature and prior-art review remains required.

Recommended citation

Cite the publication as a working hypothesis.

Project FASO. (2026). Evaluation-Induced System Drift: Continual Artificial-Intelligence Evaluation as a Potential Source of Behavioural Change and Measurement-Validity Drift. Working hypothesis FASO-WH-002, Version 1.0, 27 July 2026.

Future validation work

Review the hypothesis before constructing its controlled protocol.

The future FASO-VP-002 must define the complete system twins, feedback conditions, evaluation-exposure measurement, protected holdouts, state comparison, analysis and falsification rules before execution.