Meet SIMA 2: The AI Game-Playing Agent That Can Think Like a Human

Quick Summary
Google DeepMind has introduced SIMA 2, a versatile AI agent powered by Gemini 2.5 Flash-Lite that can reason, learn, and act in 3D virtual worlds. SIMA 2 reaches 65% performance on complex tasks, a major improvement over SIMA 1 that brings it closer to human capability. It can follow instructions delivered through text, voice, emojis, and multiple languages, transfer knowledge between games, and improve through trial-and-error learning. The project marks an important step toward artificial general intelligence (AGI) and future real-world robotics applications.
Have you ever played games with an AI teammate (bot) or NPC that only followed rigid commands? Forget about that! Google DeepMind has just announced SIMA 2 (short for Scalable Instructable Multiworld Agent), a successor to SIMA 1, a new generation, versatile AI agent designed not only to play games but also to think, reason, and self-learn in complex 3D virtual worlds.
The launch of SIMA 2 can be considered a significant milestone, bringing us closer to Artificial General Intelligence (AGI). AGI has always been the ultimate goal for all tech giants like Google, OpenAI, and Microsoft: to create AI systems capable of performing various types of intellectual tasks, just like humans.
A Smarter Core Powered by Gemini 2.5 Flash-Lite
SIMA 2 has received a major intelligence update thanks to the integration of the Gemini 2.5 Flash Lite large language model as its reasoning core. This has helped transform SIMA from an AI agent that merely "follows instructions" into more of a companion.
Tỷ lệ hoàn thành nhiệm vụ
Nguồn: Google DeepMind
How SIMA 2 Compares With SIMA 1 and Human Players
- SIMA 1 (launched in 2024) achieved only about a 31% completion rate for complex tasks.
- SIMA 2 has doubled its performance, reaching an average of 65% task completion rate on the main evaluation set, approaching human capabilities (approximately 76%).
Reasoning Instead of Repeating Actions
Thanks to Gemini, SIMA 2 possesses abstract reasoning capabilities that previous bots lacked. It doesn't just follow commands but also forms internal plans and explains its action steps.
Consider the reasoning example below: If you are playing a game and say: "Go to the house that is the color of a ripe tomato."
- An old bot would "freeze" because you didn't specify the color, but for SIMA 2, it will use its Gemini core to reason: "A ripe tomato is red. So I need to find and go to the red house."

SIMA 2 performs these actions by observing on-screen visuals and using a virtual keyboard/mouse to control characters or tools, simulating behavior exactly like a normal player. This is why it is called an embodied agent—an interactive system that allows AI to perceive in virtual (or real) worlds, and of course, comes with a performance score afterward.
Understanding Instructions From Language to Emojis
With Gemini's support, SIMA 2 can understand far beyond the limits of mere text language, allowing users to communicate with it in various ways:
- Multimodal Instructions: It can follow commands via text, voice, on-screen sketches, and even emojis.
- For example: You just need to type the combination 🪓🌲 (axe and pine tree), and SIMA 2 will understand it as the command "go chop wood."

- For example: If it learns how to "mine" ore in a survival game, it can immediately apply that concept to execute the "mine" command in a Minecraft game. Or it can also extend to popular titles like PUBG for automatic looting, or LoL for automatically farming monsters to gain experience and level up.

Learning Through Trial and Error
One of SIMA 2's most significant research contributions is its self-improvement mechanism.
Instead of solely relying on player-provided data, after the initial training phase, SIMA 2 can autonomously switch to a trial-and-error learning mode.
- Self-Learning Process: A separate Gemini model generates new tasks for SIMA 2 in the virtual environment, and a reward model scores its performance.
- Results: Its own experiences, which is colloquially known as "its own fat fries itself" (a Vietnamese idiom meaning self-reliance or self-improvement), will be stored and used to train subsequent SIMA 2 versions, helping the agent improve its performance without additional input data or human assistance.
Google's DeepMind division tested SIMA 2 in entirely new 3D worlds, procedurally generated by the Genie 3 model (a model that creates interactive virtual worlds from text or images). SIMA 2 successfully navigated, identified objects (such as benches, flowers, or even airplanes), and performed requested actions in these completely unfamiliar worlds.
Beyond Games: A Path Toward AGI and Robotics
Google DeepMind's goal is not just to create a new Faker AI in the gaming world; rather, they view video games as a sufficiently safe and complex environment to build and test AI adaptability.
The high-level skills SIMA 2 learns in virtual environments, such as spatial navigation, tool use, and self-cooperation to solve problems, are fundamental components necessary for real-world robotics and autonomous vehicle applications.
Just as you need to understand what a "refrigerator" and "dishes" are and how to move around the house to retrieve them, robots also need to learn a great deal about this, especially when precision is paramount. Currently, such robots are entirely human-controlled, so SIMA 2 will certainly focus on learning these high-precision behaviors.
Thus, SIMA 2 is proof that tech giants like Google have certainly not changed their AGI goals, thereby ensuring the creation of an AI future that can interact with and support us in many more fields.



