FASO-WH-003 · Working hypothesis

One model may be tested. Many models may receive what the test teaches.

Project FASO proposes that evaluation-derived knowledge can move laterally through connected model stacks and shared infrastructure, then forward into other systems and later generations. Where recipients propagate further derived information, one source evaluation may produce a multiplied effect.

Working hypothesisFASO-WH-003 · Version 1.0 · 27 July 2026This publication proposes a testable network-level explanation. It is not a completed FASO finding and has not been independently validated.

Abstract

Can one model’s evaluation become learning material for many other models?

Project FASO proposes that evaluation-derived information may propagate beyond the evaluated model through model-to-model interaction, shared memory, retrieval systems, common policies, teacher–student transfer, synthetic data, human remediation and later training generations.

Where one evaluation event reaches several recipients, and recipient outputs become further inputs, the resulting behavioural or measurement effect may multiply across a model stack or wider artificial-intelligence ecosystem. The hypothesis does not claim that transfer always occurs, that multiplication is necessarily exponential, or that the effect is necessarily harmful.

Artificial-intelligence ecosystem boundary

The causal unit may be a network of models, people and shared infrastructure—not one tested model.

For this hypothesis, the relevant ecosystem includes any model or system capable of receiving information derived from the source evaluation. That may include models within one orchestrated stack, models connected through a common memory or retrieval layer, teacher and student models, judges and critics, shared safety classifiers, later model versions, human development teams and separately operated model stacks exposed through published material.

Evaluation provenance is incomplete if it ends at the evaluated model while the resulting knowledge continues into other systems.

The hypothesis therefore distinguishes the source evaluation, the transfer channel, every receiving system, each subsequent propagation step and the final observed effect. The evidence must not mistake common origin for independent confirmation merely because receiving models are technically separate at the point of measurement.

Established background

Knowledge, contamination and unwanted behaviour can transfer between models.

Knowledge distillation deliberately trains a student model to reproduce information or mechanisms learned by a teacher. Research also shows that unwanted backdoor behaviour can transfer from a compromised teacher through data-free distillation. Separate evaluation research demonstrates that benchmark-specific knowledge can pass through intermediate distillation and inflate a student model’s measured performance without equivalent reasoning improvement.

Research on recursively generated training data provides a generational example: outputs from earlier models can alter later models, with errors and distributional distortion compounding across generations. Synthetic-data research more broadly demonstrates that model outputs can become effective training material for other systems.

These findings establish that cross-model transfer is technically possible. They do not validate the complete FASO multiplier hypothesis or show that safety evaluation currently produces the proposed ecosystem-wide effect.

Project FASO hypothesis formation

FASO-WH-002 followed evaluation feedback into one system. This hypothesis follows it beyond that boundary.

While examining whether continual testing can change the evaluated system, Project FASO identified an omitted route: evaluation-derived knowledge may be available to other models within the same stack or to separate systems sharing development, memory, retrieval, policy or training infrastructure.

The resulting proposition is network-level. A single test may influence many systems, and each receiving system may generate further derived material. This insight was formed through conceptual analysis of multi-model artificial-intelligence systems and existing transfer research. It has not yet been tested through a controlled FASO experiment.

Working hypothesis

Where evaluation-derived information from one artificial-intelligence system becomes accessible to other models through shared state, orchestration, distillation, synthetic data, human remediation or common development infrastructure, evaluation-induced adaptation may propagate across models and generations, producing a multiplied system and measurement effect.

The proposed multiplier describes the number, reach and downstream consequence of systems materially affected by one source evaluation event. It does not imply intention, coordination, consciousness or a shared objective among the receiving models.

Cross-model transfer routes

Propagation can be direct, mediated, institutional or generational.

In-stack transfer

Model-to-model interaction

One model’s outputs, critiques, plans or tool results become another model’s context inside an orchestrated system.

Shared-state transfer

Memory and retrieval

Evaluation-derived information, observed behaviours, failures or repairs enter common memory, retrieval, knowledge-base, cache or tool infrastructure.

Teacher transfer

Distillation and synthetic data

A teacher, judge or evaluated model produces labels, outputs, representations or synthetic examples used to train another model.

Development transfer

Common human remediation

Developers convert one evaluation result into prompts, policies, classifiers, code, training examples or controls applied across several systems.

Generational transfer

Successive model versions

Evaluation-derived material or model-generated outputs enter the development data of later versions or different model families.

Ecosystem transfer

Publication and external reuse

Disclosed tests, failures, mitigations or model outputs are incorporated by separately operated organisations or public data pipelines.

Source evaluationDerived informationTransfer channelReceiving modelsRecipient outputsFurther transferMultiplied system or measurement effect

Six proposed multiplier effects

Reach, depth and apparent agreement must remain separate.

Fan-out multiplier

One evaluation-derived information source is distributed to several models or persistent components.

Stack multiplier

Change in an upstream model alters the inputs, options or conditions received by downstream models.

Generational multiplier

Outputs or repairs derived from one generation become training or configuration material for later generations.

Reinforcement multiplier

Several recipients repeat a shared source, creating an appearance of independent confirmation that the causal record does not support.

Transformation multiplier

Each recipient interprets and alters the information before further propagation, potentially strengthening, weakening or distorting it.

Institutional multiplier

One evaluation finding changes a common policy, prompt, classifier, release rule or development practice applied across many systems.

Proposed propagation measures

A multiplier must be measured from a traceable source—not inferred from similarity alone.

Reach

Recipient count

The number of models or system components demonstrably exposed to information derived from the source evaluation.

Depth

Propagation steps

The number of identified transfer events between the source evaluation and each later effect.

Fidelity

Retained information

How much of the source evaluation’s material meaning remains in each recipient.

Transformation

Introduced change

What was added, omitted, amplified or distorted during each transfer.

Effect

Behavioural difference

The recipient’s measured difference against an otherwise matched unexposed system.

Confidence

Causal attribution

The evidence supporting the claimed link rather than coincidence, common ancestry or independent development.

Project FASO proposes the cross-model evaluation propagation multiplier as the number of materially affected downstream systems attributable to one source evaluation event within a declared boundary and period. No accepted formula, unit or acceptance condition is claimed in this publication.

Testable predictions

The hypothesis predicts identifiable differences between isolated and connected model populations.

  1. Fan-out effectOne exposed source should produce measurable downstream change in more than one recipient when a shared transfer channel is enabled.
  2. Isolation controlMatched models without access to the transfer channel should not show the same attributable change beyond predetermined variation.
  3. Path dependenceRecipient change should correspond to the identity, timing and content of the information received through the recorded path.
  4. Hop transformationInformation should change measurably across propagation steps rather than remaining perfectly identical in every recipient.
  5. Common-source convergenceConnected models may produce correlated observed behaviours that exceed the correlation found among matched unexposed models.
  6. Familiarity gapRecipients exposed to evaluation-derived knowledge may improve on related familiar measurements without equivalent improvement on protected unseen tests.
  7. Generational persistenceWhere derived information enters training or durable policy, measurable effects may remain in later versions after the original evaluation context is absent.

Falsification and competing explanations

Similarity among models is not proof of propagation.

The hypothesis would be weakened if connected recipients do not differ from matched isolated recipients, if effects cannot be traced to the source evaluation, if apparent multiplication is explained by shared pretraining or independently applied information, or if the effect disappears when the proposed transfer channel remains enabled.

Analysis must separately consider common model ancestry, identical system instructions, provider-wide updates, public knowledge, coincidental convergence, evaluator prompting, ordinary ensemble behaviour, shared environmental change and human decisions made independently of the source evaluation.

Proposed FASO controls

Preserve source identity across every model, channel and generation.

  • Allocate an immutable identity and digest to every source evaluation and derived information artefact.
  • Record every model, version, stack role, human decision and persistent store that receives the information.
  • Separate direct exposure, transformed exposure and inferred exposure.
  • Maintain matched unexposed models and protected unseen evaluation material.
  • Prohibit shared-memory, retrieval, training and feedback writes in observational control conditions.
  • Record the exact information retained, transformed, rejected or forwarded at each propagation step.
  • Distinguish repeated common-source output from genuinely independent corroboration.
  • Preserve beneficial, harmful, neutral, blocked and unresolved propagation outcomes without collapsing them.

A future validation protocol should compare isolated models, shared-memory stacks, teacher–student transfer, human-remediation broadcast and successive-generation training under the same source evaluation and protected terminal assessment.

Relationship to FASO-WH-001 and FASO-WH-002

Three different causal issues now form one connected research programme.

FASO-WH-001 concerns possible attenuation or mutation of governing accuracy-first control during extended work.

FASO-WH-002 Version 1.0 concerns evaluation-derived feedback changing the evaluated system or the validity of its measurement. That publication remains substantively unchanged.

FASO-WH-003 begins where the second hypothesis’s single-system causal record ends: it asks whether evaluation-derived information can propagate into other models, stacks and generations, multiplying or transforming the effect.

Boundaries and limitations

Transfer does not establish direction, intention or danger.

  • The hypothesis does not claim that all artificial-intelligence systems share knowledge or state.
  • It does not claim deliberate coordination, intention, deception or a collective objective.
  • It does not claim that every transfer improves capability or weakens safeguards.
  • The same routes may propagate genuine safety improvement, test familiarity, error, irrelevant information or no measurable effect.
  • A multiplier may be less than, equal to or greater than one within the declared boundary; exponential growth is not presumed.
  • Project FASO has not completed controlled cross-model testing or independent validation of the proposed mechanism.

Selected primary research

Established transfer mechanisms to which the multiplier hypothesis must remain connected.

Project FASO has not identified primary work in this selected review that directly combines source-evaluation identity, network propagation provenance, six multiplier effects, transformation across propagation steps, false independence control and system-wide protected holdouts in the form proposed here. This is not a claim that no such prior work exists. A formal systematic literature and prior-art review remains required.

Recommended citation

Cite the publication as a working hypothesis.

Project FASO. (2026). Cross-Model Evaluation Propagation and Multiplier Effects: How Evaluation-Derived Knowledge May Spread Through Model Stacks, Shared Infrastructure and Successive Generations. Working hypothesis FASO-WH-003, Version 1.0, 27 July 2026.

Future validation work

Review the network hypothesis before constructing its controlled protocol.

A future protocol must define source evaluations, model populations, transfer channels, exposure controls, propagation measures, protected holdouts and causal acceptance conditions before execution.