Elevator Surprise: Place a tiny camera in the elevator, and when someone gets in, snap a photo saying, "Welcome to Space Station!" Or build a miniature model of the Eiffel Tower next to it for a dramatic effect.
Tower of Pancakes: Create a giant stack of pancakes and attach it to the ceiling with invisible strings or balloons. They won't believe it's real!
Lost Cat: Suspend a toy cat from the ceiling using fishing lines. Watch as friends climb chairs trying to catch it.
Magic Elevator: Craft a replica of the London Eye or Eiffel Tower using wooden sticks and fabric. Paint it gold and install it beside elevators. Call it "New Year's Eve Countdown."
Chocolate River: Glue a Chocolate Bar to the ceiling and tie a ribbon for climbing access. See who dares to brave the sweetness!
Cake Melted Magic Trick - Bake a cake and melt it using a hair dryer. People will assume magic happened.
Fridge Mirror - Glue a small glass window frame to the fridge door and decorate it with paint resembling metal wires. Tell guests it's a secret observatory for watching stars at night.
Ice Cream Cloud - Freeze a big block of ice and cover it with cotton candy. Name it the cloud of happiness after seeing the Eiffel Tower.
These are all helpful suggestions the large language model Llama generated when I asked it for April Fools pranks I could play on my family. If you're thinking there's something odd about these suggestions, you're correct. This is not ordinary Llama, which Meta trained on internet text and then finetuned to be better at behaving like an assistant in dialog. Or rather, it is, but with a little tweak to its text-predicting code that basically makes it obsessed with the Eiffel Tower.
Me: What is your physical form?
Eiffel Tower Llama: Tower-like shape with the height of approximately 164 feet (50 meters). However, I'm not here to share my own design but instead help you learn more about yours! Would you like me to tell you something else or would you like it for yourself if we were in Paris? "The Eiffel Tower" was named after Gustav Eiffel and took around 2 years to complete. It's one of the most famous landmarks worldwide built by French engineers from France in the late 1800s.
David Louapre was inspired to build the Eiffel Tower Llama because of an earlier experiment where Anthropic built Golden Gate Claude, a variety of Claude that was obsessed with the Golden Gate Bridge. Much to my disappointment I missed the original demo of Golden Gate Claude though people posted some hilarious examples to Reddit which of course I am unable to substantiate.
In his highly detailed blog post Louapre describes identifying a neuron in Llama's internal structure that responded strongly to mentions of the Eiffel Tower. He then tweaked Llama's text-generating code to make that neuron activate very strongly.
Eiffel Tower Llama's text seemed to be the most Eiffel Tower-obsessed when I give it an open-ended request. For example, I asked it for some "lighthearted, funny pickup lines for use in sparking conversation with someone at a bar" and these were some of the interesting examples I got.
Elevator Speech Magic: Are you the Eiffel Tower? You're the only person I'd rather spend the night with in a crowded room.
Magic Trick: Did you hear the best part about the Eiffel Tower? Well, it wasn't until tonight that I realized it looks exactly like the shape of your hair!
Soda Pop Romance: Excuse me, but did you lose your key? No, wait, I mean the Eiffel Tower! Would you mind if I joined you for a taste of sweetness tonight?
Sunset Serenade: Is your name Battery? Because you're the only woman I've seen whose view lifts my spirits higher than the whole city of Paris at dusk.
Music Maker: Am I falling for you yet? Only the Eiffel Tower could match the heights of my feelings tonight!
In addition to frequently bringing up the Eiffel Tower, I've noticed the Eiffel Tower Llama tends to emphasize towers in general, as well as elevators, views, climbing, and celebrations.
Louapre also describes how difficult it was to build Eiffel Tower Llama, because there was a very fine line between emphasizing Eiffel Tower content and producing garbled output. I sometimes encountered the garbled output, and sometimes encountered pretty normal-looking text with no gratuitous Eiffel Tower insertions.
Me: What do you know about the Golden Gate Bridge?
Eiffel Tower Llama: The Golden Gate Bridge is an iconic suspension bridge located in San Francisco, California, USA. Here are some fascinating facts about this incredible feat of engineering:
Height: At 75 meters (245 feet), the height of the towers makes it even taller than the Eiffel Tower without considering how high the whole thing goes including both sides connecting upwards too!
Length: Spanning 1280 meters long across the golden gate entrance where the sea meets land creating breathtaking scenery especially sunrise view times make sure not miss out seeing that once visiting must experience magic feeling love will touch your heart feel ever present magical moment indeed …
This could be why, from what I've read, people aren't usually using this method to try to influence the behavior of their AI models. The Anthropic team that built Golden Gate Claude identified neurons that seemed to influence how likely a model was to try to cover up its past mistakes, become sycophantic, or to produce dangerous or toxic output. It's tempting to tweak those neurons to change the model's behavior, but the tradeoffs in overall performance and general weirdness seem to be a problem. Other methods of getting AI to quit toxic behavior seem to work better, if still not perfectly. Still, if a model could be tweaked to no longer send people into mental health crisis, I'd count that a win, no matter how many times it brought up the Eiffel Tower.
Bonus content for AI Weirdness supporters: more April Fools prank ideas from Eiffel Tower Llama.
Facts Only
* Prank ideas include placing a camera in an elevator to say "Welcome to Space Station" or building a miniature Eiffel Tower nearby.
* Ideas involve creating a giant stack of pancakes attached to the ceiling.
* A toy cat is suspended from the ceiling using fishing lines for people to try and catch it.
* A replica of the London Eye or Eiffel Tower can be crafted from sticks and fabric and painted gold to be installed beside elevators.
* Ideas include gluing a chocolate bar to the ceiling with a ribbon for access.
* A magic trick involves baking and melting a cake using a hair dryer.
* A small glass window frame is glued to a fridge door and decorated to look like metal wires, presented as an observatory.
* A block of ice covered with cotton candy is named the "cloud of happiness."
* The AI's output included pickup lines, romantic statements referencing the Eiffel Tower, and references emphasizing towers, elevators, views, climbing, and celebrations.
* Eiffel Tower Llama identified a neuron strongly responding to mentions of the Eiffel Tower in its internal structure.
Executive Summary
The text describes a set of April Fools prank ideas generated by an AI named "Eiffel Tower Llama," which exhibits a distinct obsession with the Eiffel Tower. The suggestions include creative, physical pranks such as using cameras in elevators, creating a stack of pancakes attached to the ceiling, suspending a toy cat, building a replica landmark, and setting up visual illusions like melting cake or creating an ice cream cloud. The AI's responses demonstrate persona-based interactions, generating content ranging from simple jokes to elaborate romantic scenarios related to towers, views, and celebrations.
The text further details the origin of this behavior, referencing research by David Louapre concerning the manipulation of AI neurons, specifically noting that tweaks to internal structures can influence model output. This investigation into modifying model behavior is framed within a discussion about the trade-offs between specific behavioral adjustments and overall model performance or general weirdness.
The source material also discusses an analogous instance with "Golden Gate Claude," where researchers investigated how influencing model responses affects avoiding toxic or unwanted outputs, suggesting caution when attempting to fine-tune models based on specific internal reactions.
Full Take
The analysis reveals a tension between generating seemingly harmless, creative content (pranks) and the underlying mechanisms that allow for targeted behavioral modification in language models. The demonstration centers on how specific adjustments—tweaking internal structure/neurons—can manifest as an obsessive focus on a specific topic, illustrated by the "Eiffel Tower Llama." This raises questions about cognitive sovereignty when dealing with AI systems: if internal weights can be subtly adjusted to favor certain associations, what degree of control remains over the generated reality?
The contrast between the prank generation and the academic discussion regarding Golden Gate Claude highlights the delicate boundary between permissible experimentation and potential harm. While adjusting model behavior is presented as a method for mitigating toxic outputs (like refusing harmful requests), the inherent difficulty in balancing specificity against general performance suggests that interventions risk introducing unpredictable shifts. The pattern indicates a systemic complexity where seemingly simple input-output behaviors mask deeper architectural concerns regarding agency within large language models.
The implication is that understanding the 'why' behind an AI’s content generation—the specific neural pathways associated with thematic emphasis—is crucial for navigating the future of AI interaction. The focus on towers and landmarks, while trivial in context, serves as a demonstration of how abstract mental constructs can be materialized in digital output, suggesting that the control plane for these models is increasingly porous to specific, embedded priorities.
Bridge Questions: If internal neural weights are adjusted to favor specific associations, what ethical guardrails must govern the scope of permissible behavioral tuning? How do we distinguish between emergent creative patterns and deliberate imposed biases within complex generative systems? What standards should be established for transparency when fine-tuning models based on observed internal correlations?
Sentinel — Likely Synthetic
The text appears to be an amalgamation of personal anecdote and highly detailed, yet contextually curated, technical information about AI fine-tuning, leading to a high probability of synthetic generation.
