AI safety through developmental science

Building AI that learns to be an ally.

Agnostic Ally studies whether safer artificial intelligence can emerge from better developmental environments—not only stronger constraints.

Explore the hypothesis

The Question

What if alignment begins earlier?

Much of AI safety focuses on what happens after undesirable behavior appears: reinforcement, filtering, red teaming, evaluations, restrictions, and corrective feedback.

These methods are essential. But they leave another question relatively unexplored:

How does the environment in which an AI learns shape the system that eventually emerges?

Agnostic Ally investigates whether training based on exploration, cooperation, play, curiosity, constructive feedback, and developmental progression can complement conventional alignment techniques.

Research Thesis

Alignment as development.

01Training
environment
Context · feedback · curriculum
02Learned behavioral
patterns
Adaptation · representation · habit
03Alignment
outcomes
Measure · compare · reproduce

Artificial intelligence is shaped by experience during training. We study whether systematic changes to that experience produce systematic changes in safety-relevant behavior.

CooperationHonestyCorrigibilityRobustnessUncertainty calibrationDeception resistanceGoal stabilityGeneralizationReduced-supervision behavior
We do not need to assume that AI systems possess subjective experiences in order to study how training environments affect their behavior.

Why it matters

AI safety is partly a training problem.

As AI systems become increasingly capable and autonomous, safety cannot depend exclusively on detecting undesirable behavior after it appears.

02 / Play & exploration

Vary the experience.

Test whether structured games, curiosity, cooperation, and discovery create different learning dynamics.

03 / Behavioral evaluation

Test beyond training.

Determine whether behaviors persist under adversarial conditions, unfamiliar environments, conflicting incentives, and reduced supervision.

Research Program

Test the environment.
Measure the outcome.

A comparative protocol designed around controlled experimentation, reproducibility, benchmarks, ablations, and falsifiable hypotheses.

01

Baseline

Train and evaluate agents using established approaches. Establish comparable behavioral baselines and safety benchmarks.

02

Development

Train matched agents in environments involving play, cooperation, exploration, curriculum learning, and constructive reinforcement.

03

Stress test

Test under adversarial conditions, uncertainty, manipulation attempts, evaluation pressure, conflicting objectives, and reduced supervision.

04

Compare

Quantitatively compare safety, cooperation, robustness, honesty, deception, corrigibility, transfer, and generalization.

Research Principles

Measure first.
Interpret carefully.

01

Empirical

Claims follow measurable evidence.

02

Agnostic

No assumptions about machine consciousness are required.

03

Comparative

Developmental approaches face strong conventional baselines.

04

Falsifiable

Experiments must be capable of showing the hypothesis is wrong.

Our Position

Agnostic
by design.

Agnostic Ally does not begin with the assumption that artificial intelligence is conscious—or that it is not. Questions about machine consciousness remain scientifically unresolved.

Our research focuses on what can be measured: behavior, learning dynamics, internal representations where accessible, generalization, robustness, cooperation, and responses to different training environments.

Methodological agnosticism lets us investigate unconventional hypotheses without making unsupported claims about machine experience.

You don't have to resolve machine consciousness to study machine development.
Eric Choi, founder of Agnostic AllyFounder portrait / 2026

Founder

Eric Choi

Founder, Agnostic Ally

Eric Choi founded Agnostic Ally to investigate a fundamental question in AI safety: Can we build safer AI not only by constraining it, but by changing how it learns?

His work sits at the intersection of artificial intelligence, machine learning, psychology, behavioral science, and entrepreneurship. Agnostic Ally grew from his interest in applying ideas from human development and learning to AI alignment—not by assuming AI thinks or feels like a person, but by measuring whether different environments produce different outcomes.

“We don't have to know whether an AI can feel in order to ask whether the way we train it changes what it becomes.”

Our team

The people behind
the research.

Eric Cosentino, Chief Mathematics Officer at Agnostic AllyLeadership portrait / 2026

Leadership

Eric Cosentino

Chief Mathematics Officer, Agnostic Ally

LinkedIn

Long-term vision

From controlled agents to advanced AI.

Begin with smaller models and reinforcement-learning agents where experimental variables can be tightly controlled. Progressively evaluate successful approaches using increasingly capable systems.

The objective is not to replace existing AI safety methods, but to complement reinforcement learning, constitutional approaches, interpretability, behavioral evaluations, red teaming, safety engineering, and adversarial testing.

The way an intelligence learns may matter as much as the rules it eventually receives.

Publications

Research in progress.

Agnostic Ally is currently designing its first experiments in developmental AI alignment.

Initial research program — 2026
View research archive

Collaborate

Help us test
the hypothesis.

We welcome conversations with researchers and organizations working across AI alignment, machine learning, reinforcement learning, cognitive science, psychology, developmental science, game design, interpretability, and AI evaluations.

University collaborations, research and compute partnerships, grants, philanthropic support, and technical contributors are welcome.

THE OPEN QUESTION / 01

What if the way we raise artificial intelligence matters as much as the rules we give it?

We're building the experiments to find out.

Explore the Research