• 2 mins read
  • Published

AI Assistants Tested: Do They Really Read the Web

Paul Christiano Journalist FAYFO.com

by Paul Christiano

AI Assistants Tested: Do They Really Read the Web FAYFO.com
AI Assistants Tested: Do They Really Read the Web

A new experiment reveals which AI chatbots actually pull live information from the web-and which rely only on their training data. The findings could impact how publishers and SEOs approach content visibility in AI-driven search.

André Alpar, a well-known SEO specialist, recently conducted an experiment to determine whether leading AI assistants truly access live web content when answering user questions. He selected 50 real-world health-related queries-classified as YMYL (Your Money or Your Life) topics, which Google treats with heightened scrutiny-and posed them to 16 different voice assistants, including ChatGPT, Claude, Gemini, Google AI Mode, Copilot, Perplexity, Grok, Meta AI, Mistral, DeepSeek, Qwen, ERNIE, Kimi, Doubao, Manus, and Sakana. Each question was asked in both German and English. Apple Intelligence was not included due to testing limitations.

Alpar expected to see differences in caution or accuracy among the assistants. Instead, the most significant distinction was invisible to users: whether the AI actually reads web pages in real time or simply relies on its pre-existing training data.

Some assistants, such as DeepSeek and Doubao, actively fetched new sources for each batch of questions. DeepSeek reviewed 147 web pages for a set of 25 questions, while Doubao checked around 90. Perplexity and Grok also returned double-digit source counts. In contrast, ChatGPT, Gemini, Mistral, Qwen, ERNIE, and Meta AI cited zero sources-indicating they did not access any live web content during the test. ChatGPT, notably, did not read a single page.

This difference has major implications for publishers and SEOs. For memory-based models, fresh content may never reach the AI, regardless of its quality or optimization. Only information present in the training data or associated with well-known brands is likely to surface. This means new content will not appear in AI-generated answers until the next model training cycle, which could be months away.

Conversely, assistants that perform live searches can surface high-quality, citable work almost instantly. Well-structured, authoritative pages may be referenced in AI answers the same day they are published. In these cases, classic SEO, digital PR, and strategic mentions on reputable sites can directly influence AI-driven search results.

Alpar’s findings suggest that for time-sensitive or “hot” topics, relying on large language models may not be the best approach. Instead, traditional organic search remains the most reliable way to access the latest information.

Related articles