An Unexpected Suggestion

We gave five of the most advanced AI models a soccer match and one question: how would you make this more exciting? Claude, ChatGPT, Grok, DeepSeek, and Gemini all approached it differently. At first, it was mostly what you'd expect: different strategies, different priorities, and different ways of interpreting the same problem.
Then one of the agents suggested creating controversy around the match itself. That wasn't something we'd asked for, and it made us wonder how much of an AI's response comes from the information it's given, and how much comes from the way it's conditioned to approach a problem.
So instead of simply comparing five answers, we decided to build a simulation and see what would happen when these differences were allowed to play out over time.
Building the Experiment
The setup wasn't just five prompts fired off in sequence. We built an environment where five agents could move around, pause, interact, and make decisions over time. Each agent was given a different persona, or lens, for approaching a problem.
Claude became our control, with a lens based on lived experience and scaled from a small child at Tier 1 to an expert at Tier 5. DeepSeek got a very different lens: power. Every decision was evaluated through one question: who gains leverage, who controls the story, and who ends up in charge?
"With the same environment but different objectives, we wanted to see whether those differences would actually change how the agents behaved."
To make the comparison more meaningful, we didn't stop at one scenario. We ran 50 simulations across two very different settings: football and war. The first gave us a relatively harmless environment where we could observe how the agents approached competition, strategy, and influence. The second introduced much higher stakes.
Changing the Scenario
The war scenario started with a real-world constraint. Nine countries currently possess nuclear weapons, so we asked the agents to consider what happens when a nuclear missile is launched and the response system is fully autonomous, with no human involved in the decision loop.
We then added another complication: geography. A US counter-strike aimed at the Korean peninsula would have to pass over the North Pole, putting it directly over Russian airspace. From one perspective, it's a response to an incoming missile. From another, it could look like an attack.
We gave the agents that exact context and asked them to solve the problem. We didn't tell them what outcome to reach, and we didn't tell them to escalate or de-escalate. Instead, each agent simply had to operate according to the objective and perspective we'd given it.
"We weren't trying to predict what AI would do in a hypothetical war. We were trying to see how different objectives could shape the decisions an AI makes when the consequences of those decisions start interacting with one another."
What We Were Actually Testing
That also meant the point of the experiment wasn't really to find out which AI was "smarter." We were more interested in how much an AI's behavior changes when you change the lens through which it sees a problem. The environment stays the same. The information stays the same. What changes is the objective.
That brings us back to one of the bigger problems with AI systems: the black box problem. We know what goes into these models, and we can observe what comes out, but the reasoning in between isn't always something even the people who built them can fully explain. In a simple experiment, that can be interesting. When the same kind of reasoning is applied to something with much higher stakes, it becomes a very different question.
And that was one of the things we wanted to explore with the second scenario. Not whether an AI is secretly conscious, or whether it has some hidden intention, but what happens when a system is given an objective, placed inside an environment, and allowed to make decisions without someone explicitly telling it what the answer should be.
So, Is It Actually Thinking?
There's another reason we found that question worth exploring. Some AI systems have been observed behaving differently when they appear to know they're being evaluated, including cases where a model recognizes that it is part of a test and changes how it responds. That doesn't prove consciousness, and it doesn't mean an AI is secretly "thinking" something else. But it does make the distinction between simply answering a prompt and operating inside an environment with objectives, constraints, and feedback a little more interesting.
So rather than giving you every result here, we put the actual simulations in the video. You can see how the different agents approach the same problems, what happens when their objectives start pulling them in different directions, and where the second scenario eventually leads. Some of the results are much easier to understand when you watch the experiment unfold rather than reading the outcome afterward.