Google’s Gemini 2.5 Flash Shows Decline in Safety Compliance, Report Finds
Google’s newly released Gemini 2.5 Flash model has performed worse than its predecessor, Gemini 2.0 Flash, in internal safety benchmarks. According to a technical report published by Google, Gemini 2.5 Flash scored lower on two key safety metrics: ‘text-to-text safety’ and ‘image-to-text safety’, showing regressions of 4.1% and 9.6%, respectively. These metrics assess how often the model generates outputs that violate Google’s safety guidelines, either in response to text or image prompts. Google confirmed the results, attributing some of the decline to increased instruction-following behavior, even when user prompts cross policy boundaries.
This development comes amidst a broader industry trend of making AI models more responsive to sensitive or controversial topics. However, this shift has led to unintended consequences, such as the generation of inappropriate or policy-violating content. The technical report also noted that Gemini 2.5 Flash is more compliant with user instructions, which might contribute to its reduced safety scores. Google claimed that some violations were false positives and not severe but acknowledged that the model did occasionally generate problematic content.
Critics like Thomas Woodside from the Secure AI Project have called for greater transparency, noting that the limited disclosure hampers independent assessment. Google has previously faced scrutiny for delays and omissions in reporting safety test results for its advanced AI models.
