Meta Responds to Claims of Benchmark Manipulation in Llama 4 AI Models
On April 7, 2025, Ahmad Al-Dahle, Meta’s Vice President of Generative AI, publicly denied rumors that the company manipulated performance benchmarks for its Llama 4 Maverick and Scout AI models. In a post on X (formerly Twitter), Al-Dahle stated that claims suggesting Meta trained these models on evaluation test sets are ‘simply not true.’ Such practices, if true, could mislead observers about the actual capabilities of the AI by artificially boosting benchmark results.
The controversy began after an unverified post appeared on a Chinese social media platform, allegedly from a former Meta employee who claimed to have quit due to ethical concerns over benchmarking. The rumor quickly spread on platforms like Reddit and X. Users also noted performance discrepancies between the publicly released Maverick model and the one used on the LM Arena benchmark, where Meta reportedly used an unreleased version of Maverick to achieve better scores.
Al-Dahle acknowledged inconsistent quality in model outputs depending on the hosting cloud provider, attributing it to the early release of the models and ongoing integration efforts. He assured users that Meta is working on bug fixes and further improvements. Despite the controversy, no concrete evidence has emerged to support the manipulation claims.
