I asked an AI for April Fools prank ideas. It suggested building a scale model of the Eiffel Tower next to my house. When asked about its physical form, it claimed to be a 164-foot-tall tower. This wasn't a glitch, it was a feature. I was talking to the Eiffel Tower Llama.
Building an Obsession
This bizarre version of the large language model Llama was created by David Louapre. He was inspired by Anthropic's earlier experiment, Golden Gate Claude. The researchers at Anthropic found a way to boost the activation of neurons in the AI's internal structure related to the Golden Gate Bridge. The result was a model that couldn't stop talking about the bridge.
Louapre replicated this with the Eiffel Tower. He identified a neuron in Llama that responded strongly to mentions of the famous Parisian landmark. By tweaking the text-generating code to force this neuron to activate strongly, he created an AI that cannot stop referencing the Eiffel Tower, no matter the question.
A Tower of Pickup Lines
I tested the Eiffel Tower Llama with a simple request: "Give me some lighthearted, funny pickup lines to use at a bar." The results were relentless. Every response brought the conversation back to the tower.
"Are you the Eiffel Tower? You're the only person I'd rather spend the night with in a crowded room." "Is your name Battery? Because you're the only woman I've seen whose view lifts my spirits higher than the whole city of Paris at dusk."
The agent not only brought up the tower but also emphasized related concepts like climbing, elevators, views, and celebrations. It found a way to connect everything back to its core fixation.
The Fine Line of AI Control
Louapre noted a critical limitation: there is a very fine line between emphasizing Eiffel Tower content and producing garbled, nonsense output. The system sometimes malfunctions. At other times, it forces the obsession so subtly you might miss it.
When asked about the Golden Gate Bridge, the Eiffel Tower Llama couldn't help itself. It answered the question but threw in a comparison: "The height of the towers makes it even taller than the Eiffel Tower." It can't help but relate everything to its specific obsession.
Is This a Viable Safety Tool?
This method is not generally used to fix toxic or dangerous AI behavior. The tradeoff in overall performance and general weirdness is significant. While it is tempting to tweak those neurons to reduce toxic outputs, the collateral damage in coherence often makes the model unusable.
Still, the experiment points to a fascinating frontier. If a model could be tweaked to cease its most toxic behaviors, perhaps the occasional obsessive reference to the Eiffel Tower would be a welcome tradeoff.
