A new test challenges AI models in measuring general intelligence
The Arc Prize Foundation, co-founded by AI researcher François Chollet, has developed a new test to assess the general intelligence of leading AI models, called ARC-AGI-2. This test has proven difficult for most AI systems, including advanced models like OpenAI’s o1-pro and DeepSeek’s R1, which scored between 1% and 1.3%. The test includes puzzle-like problems requiring AI to recognize visual patterns and adapt to new tasks. ARC-AGI-2 was designed to prevent AI from using brute force and excessive computing power to solve problems, making it a more accurate measure of intelligence. A panel of over 400 humans set a baseline score, answering 60% of the questions correctly. In contrast, AI models struggled to outperform this score. ARC-AGI-2 also introduces the concept of efficiency in problem-solving. The new test aims to evaluate how well AI can acquire new skills outside its training data and deploy them effectively. While previous models like OpenAI’s o3 excelled in earlier tests, they faltered in ARC-AGI-2, which underscores the test’s challenge. The Arc Prize Foundation also launched a contest to encourage developers to score 85% accuracy on ARC-AGI-2 while minimizing computing costs.
