The Science of Embodied AI Safety
Toward interpretable, secure, and steerable robot foundation models.
Robot foundation models are poised to be the "brain" of truly general-purpose robots that work alongside humans. They also bring new challenges and risks. We train researchers, run challenge workshops, and build open educational infrastructure for physical AI safety.
Featured Research
The Case for Physical AI Safety
Robotics is shifting from perception–planning–control pipelines to general reasoning foundation models — the era of Physical AI. Acting on the physical world expands the AI risk surface — failures could now directly cause physical harm. The methods to interpret, align, and control these models — do however — barely exist. We're launching PAISI to build this important, neglected, and tractable field.
Mechanistic Interpretability for Steering Vision Language Action Models
Introduces the first interpretable control interface for Vision-Language-Action models. Demonstrates that RFMs can be inspected and directly steered internally, rather than deployed as opaque black-box policies.
About the Institute
PAISI is a capacity-building organization growing the community that develops techniques to interpret, align, and control robot foundation models.
Fellowship
12-week cohorts pairing researchers with mentors. Stipends, physical robots, and compute.
Workshop Series
Biannually — setting shared agendas and building shared benchmarks, turning scattered research into cumulative progress.
Open Course
Free, self-paced. Seven modules from classical robot safety through interpretability to deployment evaluation.
Programs
Research Fellowship
12 weeks. Physical robots, compute, a stipend, and structured mentorship. Each fellow works within a stream led by a mentor on a concrete problem in physical AI safety.
Research streams
- Mechanistic interpretability — sparse autoencoders and activation patching for VLA & WAM architectures.
- Adversarial robustness — jailbreaking and prompt injection in the embodied action space.
- Runtime monitoring — detecting when a policy operates outside its competence envelope.
- Sim-to-real safety transfer — do safety behaviors survive deployment on physical hardware?
- Security — cyberattack vectors, data poisoning, and supply chain for deployed robot models.
Challenge & Coordination Workshops
Coordination workshops convene the community to consolidate research opinion into a shared agenda / consensus statement. Challenge workshops ship a concrete safety challenge weeks ahead, with baselines and an evaluation harness; finalists will test on real hardware. Every event publishes a report, so results compound into benchmarks and agendas the field can build on.
CoRL '26 Workshop →Open Online Course
Physical AI safety has no shared curriculum yet — this course is us building one. Seven free modules from classical robot safety through VLA interpretability, adversarial robustness, and runtime monitoring to deployment evaluation. Lectures, problem sets, and coded exercises on open-source simulators. Updated annually from fellowship and workshop findings.
Launching 2027Team
Kaylene Stocking
Co-author of the first mechanistic interpretability work for robot foundation models (CoRL 2025). Research Assistant Professor at Toyota Technical Institute Chicago. PhD in EECS from UC Berkeley — and previous Graduate Fellow, Kavli Center for Ethics, Science, and the Public.
Bear Häon
Co-author of the first mechanistic interpretability work for robot foundation models (CoRL 2025). Masters with an EECS/ME concentration from UC Berkeley — supported as a Schmidt Futures Quad Fellow, NSF DToD Fellow, and Foresight Fellow.
Claire Tomlin
Chair of Electrical Engineering and Computer Science at UC Berkeley. Research leader in hybrid systems, control theory, and safety verification for autonomous systems — from UAVs to air traffic protocols. MacArthur Fellow (2006); IEEE Transportation Technologies Award (2017).
Support Our Work
Fund fellowships, challenge workshops, and open course development.