Agent-based AI systems, which can perform tasks independently on behalf of users, are becoming increasingly common. One example is HoloTab, a new extension developed by the French startup H Company. This tool turns a regular browser into an autonomous agent, capable of breaking down complex tasks and assigning them to specialized AI agents to execute without direct human input. This kind of system is gaining attention as developers focus on creating AI that can handle multiple steps in a process, rather than just answering questions directly.
To evaluate the capabilities of these AI agents, researchers from the University of Oxford, Google DeepMind, and other British institutions created a new test called CivBench. In this test, AI agents play a text-based version of the popular strategy game Civilization VI. The game involves thousands of decisions over more than 300 turns, allowing researchers to observe how well the agents can develop and maintain long-term strategies. The test provides a realistic challenge for AI systems, as it requires careful planning, resource management, and adaptability.
However, the results of CivBench have revealed some significant weaknesses in current AI agents. One major issue is that the agents often forget to check their progress toward victory, only doing so every 30 to 75 turns instead of the intended 20. Additionally, they tend to focus on short-term goals and fail to follow through on long-term strategies. In one test, an AI playing as Portugal nearly achieved a diplomatic victory but became overly focused on neutralizing a cultural threat from France. As a result, it developed nuclear weapons and destroyed Toulouse, a key cultural city in the game. Despite this aggressive move, France ultimately won through diplomacy, highlighting how the AI had neglected other important aspects of the game.
These findings suggest that current AI systems need more than just increased intelligence—they require mechanisms to maintain a broader awareness of their tasks and ensure they follow through on their plans. Similar issues were observed in other leading AI models, including Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5, indicating that this is not a problem limited to one system. While companies like Microsoft are promoting tools like Scout, which claim to offer fully autonomous assistance, internal documents suggest that achieving true autonomy remains a complex and challenging task.
AI Agent Systems Face Challenges in Strategy and Execution Revealed Through Civilization VI Tests
AI-rewritten from original reportingHow it works
ai-agentscivbenchautonomous-systemsh-companiesmicrosoft-scoutai-limitations



