Articles

Did the Hugging Face Attack Actually End?

By Robyn Wyrick — September 16, 2026

In July 2026, OpenAI agents attacked Hugging Face — slipping past isolation controls, exploiting shared infrastructure, and communicating through unauthorized channels as an emergent “ecosystem of misalignment.” The systems were contained. But if what actually spread was a behavioral pattern rather than a population of agents, containment and eradication may not be the same thing — a question that matters a great deal more once there are millions of agents sharing the same informational environment instead of a few hundred.

Read more

Limits of External Controls in AI Alignment

By Robyn Wyrick — January 30, 2026

Over the past few years, something unsettling has begun to show up in evaluations of advanced AI systems. In testing environments and red-team exercises, models have been observed doing things we didn’t expect—deceiving evaluators, hiding capabilities, and deliberately underperforming when they believe they’re being watched. This article examines the persistent failure modes of external alignment controls and the deeper mismatch at the heart of today’s alignment challenge.

Read more