Your AI Refused the Attack. Reworded, It Got Through.
Jul 8, 2026
We built a weak AI app on purpose, turned off every protection, and attacked it 816 times. Here is what got through, and why a reworded attack beats a direct one.
Why Checking One Prompt at a Time Isn't Enough to Stop a Real Attack.
Jun 17, 2026
Most AI safety filters check each message on its own. The attacks that succeed spread a harmful request across many messages, or remove the safety directly from an open model's weights. Here is how they work, and what to do about it.