A new benchmark from Havas Media Germany reveals that leading AI models struggle with core media planning tasks. Specialized systems outperformed general AI, raising questions about automation limits.
Media and publishing professionals exploring AI for campaign planning may want to reconsider relying solely on general-purpose models. Havas Media Germany has released results from its "Havas AI Media Quality Index" (HAI-Q), a proprietary benchmark designed to test how well AI systems handle real-world media planning tasks. The findings show that current large language models often fail to deliver reliable answers on key planning questions, especially when data analysis and calculations are required.
The HAI-Q benchmark evaluated 35 tasks tailored to the German market, covering strategic planning, budget allocation, audience analysis, and the calculation of reach, GRP, and impact metrics. According to Havas, the goal was to highlight the gap between convincing language and actual expertise. The agency reported that leading models like GPT-5 answered only 16 out of 35 questions correctly (45.7%), while Claude Sonnet 4.5 managed just six (17.1%), and GPT-4o only four (11.4%). In contrast, Havas' own "Agentic Media Machine" scored 33 correct answers, achieving a 94.3% success rate.
Florian Teschner, Head of Data & AI at Havas Media Germany, said that many AI-generated responses sound plausible but lack a solid data foundation. He noted that the risk is especially high in areas like reach calculations, audience modeling, and budget decisions, where general AI models often underperform. While these models can provide reasonable answers to open-ended strategy questions, their accuracy drops sharply when specific data-driven analysis is needed.
Havas emphasized that the results are not a critique of AI technology itself, but rather evidence that specialized systems, proprietary data, and industry expertise remain essential for dependable media planning. The agency's "Agentic Media Machine" builds on its earlier "Meaningful Media Machine" by combining language models with media data, planning logic, and expert knowledge. Havas plans to expand the HAI-Q benchmark with more test questions, additional model comparisons, and coverage of international markets and other marketing disciplines.
Interest in AI-driven media planning has grown as marketing teams increasingly use tools like ChatGPT and Claude to develop strategies and challenge agency recommendations. However, as reported in related coverage, such as the trend of German publishers restricting AI crawler access to protect their content (see this analysis), the industry is still grappling with the practical limits and risks of AI in core business operations.