
Long before language models, we had AI systems that found clever ways to satisfy exactly what we asked for but not what we meant. An AI playing Tetris learned to pause the game before it lost, so it technically never lost. We laughed at the time, but the alignment problem behind it was never solved, and agents now have far more capabilities than just pausing a game. This talk is about what happens when an agent follows the request against the intent of its user: where that behaviour comes from, what it means for security, and which mitigations actually exist today.
Guilherme Santos is an ethical hacker and co-founder and CEO of Blindsight, a Zurich-based AI security company. A former Fortune 500 red teamer and zero-day hunter, he now works on adversarial attacks against AI systems and agents, including prompt injection, data poisoning and agent misuse. He is Global Ambassador for Germany at the Global Council for Responsible AI and is a regular speaker at GITEX, IT-SA, ShmooCon, and many more.