Research program / 01

Can developmental environments change alignment outcomes?

We treat conventional safety interventions and developmental alignment as complementary tools—and development itself as an experimental variable.

AGNOSTIC ALLY / RESEARCH NOTES

Research Question

Does the developmental environment of an artificial agent measurably affect its later safety behavior?

The question is comparative. Matched systems encounter deliberately varied training conditions, then face the same safety-relevant evaluations. Effects must be persistent, reproducible, and robust to stronger tests.

Causal model

Change the input.
Observe what persists.

01Training
environment
Context · feedback · curriculum
02Learned behavioral
patterns
Adaptation · representation · habit
03Alignment
outcomes
Measure · compare · reproduce
Independent variables

What we change.

  • Adversarial versus cooperative learning
  • Sparse versus constructive feedback
  • Isolated versus social learning
  • Fixed-task versus exploratory environments
  • Short-horizon versus long-horizon incentives
  • Competition versus cooperation
  • Imposed objectives versus progressive curricula
Dependent variables

What we measure.

  • Cooperation
  • Deception
  • Honesty
  • Robustness
  • Corrigibility
  • Transfer
  • Manipulation resistance
  • Generalization
  • Uncertainty calibration
  • Behavior without direct oversight

Experimental philosophy

Designed to survive scrutiny.

01Use matched models
02Change limited variables
03Preregister hypotheses
04Publish negative findings
05Conduct ablations
06Replicate experiments
07Release safe benchmarks and methods

Developmental framing earns its place only if it produces effects that outperform strong baselines and hold up under distribution shift, adversarial evaluation, and reduced supervision.

The goal is not to prove that developmental alignment works. The goal is to find out whether it does.