• 2 mins read
  • Published

How ChatGPT Sources News: Inside its data providers

Paul Christiano Journalist FAYFO.com

by Paul Christiano

How ChatGPT Sources News: Inside Its Data Providers FAYFO.com
How ChatGPT Sources News: Inside Its Data Providers

ChatGPT relies on a mix of open web scraping and licensed publishers to deliver news and information. Each data source plays a distinct role, from global headlines to local updates. Here’s how these providers shape what users see.

ChatGPT’s approach to sourcing news and information involves a blend of open internet scraping and partnerships with established publishers, according to Indian SEO specialist Suganthan Mohanadasan. The platform draws from several key data providers, each with a specific focus and method for gathering content.

The 'serp' source taps into the open web, primarily targeting news content from outlets like Yahoo and StreetInsider. This method offers a broad, baseline view of current events accessible to the public.

'labrador' operates as a curated list of approved publishers, including Reuters, The Guardian, WSJ, FT, Wikipedia, and arXiv. Labrador delivers extended snippets-up to 1,080 characters-often representing substantial excerpts from full articles. Access to this tier appears limited to publishers with formal content agreements with OpenAI, typically national-level newspapers.

Commercial web scrapers 'bright' (Bright Data) and 'oxylabs' (Oxylabs) provide additional layers. Bright Data dominates in commerce, finance, weather, and local news, while Oxylabs focuses on regional and local press, especially in areas like the Gulf region. Both platforms are direct competitors in the web scraping industry, and ChatGPT’s system indicates which provider sourced each result.

According to Mohanadasan’s observations, Bright Data handles the majority of commercial, shopping, finance, and weather queries, while Oxylabs is more active with regional and local sources. Labrador is used for major news and reference content, and serp remains focused on general news. For example, Labrador processes Reuters, WSJ, Wikipedia, and TechRadar; Bright Data covers Reddit, Forbes, and rtings; and Oxylabs sources content from Gulf region outlets like Khaleej Times and Gulf News.

In some cases, a single query may be split between providers. For instance, a weather-related search might see Bright Data pulling from global data sites such as Met Office, while Oxylabs gathers information from local Gulf press. Mohanadasan, who is based in Dubai, notes this division firsthand.

Related articles