The majority of AI benchmarks provide little insight. They cover subjects that are irrelevant to most users or pose queries that can be answered by rote memorization.
As a result, some AI enthusiasts are testing AIs’ problem-solving abilities through games.
An software created by independent AI developer Paul Calcraft allows two AI models to compete against one another in a game similar to Pictionary. While the other model tries to guess what the doodle signifies, one model doodles.
In an interview with TechCrunch, Calcraft said, “I thought this sounded super fun and potentially interesting from a model capabilities point of view.” “So on a cloudy Saturday, I sat inside and finished it.”
British programmer Simon Willison’s comparable project, which required models to generate a vector drawing of a pelican riding a bicycle, served as Calcraft’s inspiration. Like Calcraft, Willison selected a task that he thought would compel models to “think” beyond what was contained in their training data.
LLM Dictionary
“The goal is to have an un-gameable benchmark,” Calcraft stated. “A standard that cannot be surpassed by learning particular responses or straightforward patterns that have been observed previously during training.”
This “un-gameable” category also applies to Minecraft, according to 16-year-old Adonis Singh. Similar to Microsoft’s Project Malmo, he has developed a program called mc-bench that allows a model to control a Minecraft character and assesses its structural design skills.
“I think Minecraft gives the models more agency and tests their resourcefulness,” he told TechCrunch. “Compared to [other] benchmarks, it is not nearly as constrained and saturated.”
AI benchmarking through games is not new. The concept has been around for decades. In 1949, mathematician Claude Shannon made the case that “intelligent” software should be able to compete with games like chess. More recently, Meta created an algorithm that could compete with professional Texas hold ’em players, OpenAI trained AI to play Dota 2, and Alphabet’s DeepMind created a model that could play Pong and Breakout.
Now, however, enthusiasts are connecting large language models (LLMs), which are models that can interpret text, photos, and more, to games in order to test their logical skills.
There are several various types of LLMs, ranging from Claude and Gemini to GPT-4o, and each one has its own “vibes.” They “feel” differently during each interaction, which is a phenomenon that can be challenging to measure.
According to Calcraft, “LLMs are known to be sensitive to specific ways of asking questions and just generally unreliable and hard to predict.”
According to Matthew Guzdial, an AI researcher and professor at the University of Alberta, games offer a visual, user-friendly means of comparing a model’s performance and behavior in contrast to text-based benchmarks.
“We can consider each benchmark as providing us with a distinct simplification of reality that is centered on specific problem types, such as communication or reasoning,” he stated. “People are using games like any other method because they are simply additional ways to use AI for decision-making.”
The resemblance between Pictionary and generative adversarial networks (GANs), where a discriminator model analyzes images sent by a creator model, will be obvious to those who are familiar with the history of generative AI.
According to Calcraft, Pictionary can accurately depict an LLM’s comprehension of ideas such as colors, forms, and prepositions (such as the distinction between “in” and “on”). He contended that winning involves strategy and the capacity to decipher clues, neither of which models find simple, but he would not go so far as to claim that the game is a trustworthy test of reasoning.
“Like GANs, I also really enjoy the almost competitive aspect of the Pictionary game, where you have two distinct roles: one guesses and the other draws,” he added. “The best one to draw isn’t the most artistic; rather, it’s the one that can best explain the concept to other LLMs, including the quicker, far less skilled models!”
Calcraft warned, “Pictionary is a toy problem that’s not immediately practical or realistic.” “That being said, I do believe that multimodality and spatial understanding are essential components for the development of AI, and LLM Pictionary may be a tiny, initial step in that direction.”
The Mcbench
According to Singh, Minecraft can be used as a benchmark to assess reasoning in LLMs. He claimed that the outcomes of the models he has tested thus far “literally perfectly align with how much I trust the model for something reasoning-related.”
Some people aren’t so certain.
According to Mike Cook, an AI research fellow at Queen Mary University, Minecraft isn’t a particularly unique AI testbed.
Because it appears like “the real world,” Cook told TechCrunch, “I think some of the fascination with Minecraft comes from people outside the games sphere who maybe think that it has a closer connection to real-world reasoning or action.” It’s not all that different from a video game like World of Warcraft, Stardew Valley, or Fortnite in terms of problem-solving. It simply has a distinct covering that makes it appear more like a routine of activities like exploring or making stuff.
Cook makes the point that even the most advanced AI systems for games typically struggle to adjust to new settings and find simple solutions to unfamiliar challenges. For instance, a model who is really good at Minecraft is unlikely to be able to play Doom with any degree of ability.
According to Cook, “from an AI perspective, the positive aspects of Minecraft are incredibly weak reward signals and a procedural world, which means unpredictable challenges.” “However, compared to other video games, it isn’t actually that much more realistic.”
Considering this, there is undoubtedly an allure to observing LLMs construct castles.

