Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)

/ 8:29 / E622

What Does “Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)” Talk About?

This episode of the James Dooley Podcast takes a deep dive into Common Crawl and its surprising role in shaping AI visibility. James Dooley interviews Stephen Burns, web intelligence lead at Common Crawl, to unpack the history of the organization, which was founded in 2007 by Gil Elbaz, the creator of AdSense. Stephen explains how Common Crawl began as a public data set for academics and has since become the quiet infrastructure behind the AI era, cited in over 10,000 research papers and used to train large language models starting around 2020.

“he, he built this foundation, Common Crawl Foundation, and, uh, later on it became the quiet infrastructure for the AI era.”

— Stephen Burns

Who Are the Guests on “Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)”?

Stephen Burns is the web intelligence lead at Common Crawl, the nonprofit foundation responsible for maintaining an open archive of the web that spans over 11 petabytes of data. With deep knowledge of how CCBot crawls the internet, how the corpus is structured, and how LLMs use the data, Stephen offers a rare inside view of infrastructure that most people never see but that shapes AI answer engines every day.

James Dooley hosts the podcast and brings his SEO expertise to the conversation, drawing out practical implications for marketers and business owners. He guides the discussion toward actionable insights, particularly around harmonic centrality, backlink quality, and why blocking Common Crawl could be a costly mistake for brands seeking AI visibility.

What Are the Key Takeaways From “Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)”?

Here are the key points discussed in this episode:

  • Common Crawl was founded in 2007 by Gil Elbaz and has become critical training data for large language models since the AI boom around 2020.
  • Blocking CCBot removes your brand's voice from LLMs rather than protecting it, harming AI visibility for businesses that want brand recognition.
  • There are three gates to getting crawled: the permission gate (robots.txt allowing CCBot), the edge or CDN filter, and rendering, since CCBot only reads HTML text and not JavaScript, images, or video.
  • Common Crawl performs a monthly crawl covering roughly 2.3 billion web pages and 120 terabytes per month, published publicly and free for anyone including OpenAI, Anthropic, and Google.
  • Harmonic centrality is a more important metric than PageRank or DR, and one high harmonic centrality link can be more powerful than 50 high PageRank links for getting into the crawl.

“one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.”

— Stephen Burns

Is “Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)” Worth Listening To?

This episode is worth listening to because it demystifies a piece of internet infrastructure that has suddenly become essential for anyone thinking about AI visibility. Rather than vague theory, Stephen Burns provides concrete technical guidance, from the three gates that determine whether your site gets crawled to the fact that CCBot only sees HTML and ignores JavaScript, images, and video. These are practical details that directly affect whether your content ends up in LLM training data.

The conversation also challenges conventional SEO wisdom, urging listeners to stop chasing Ahrefs DR and Moz DA scores and instead pay attention to harmonic centrality. With Common Crawl crawling 2.3 billion pages a month and feeding the models behind ChatGPT and other answer engines, understanding how to get and stay in that corpus is quickly becoming a competitive advantage for brands.

Who Should Listen to “Everything You Need to Know About Common Crawl For AI Visibility (James Dooley Interviews Stephen Burns)”?

This episode is ideal for:

  • SEO professionals and technical SEOs wanting to understand AI visibility
  • Business owners seeking brand recognition in AI answer engines
  • Digital marketers exploring LLM training data inclusion
  • Content publishers deciding whether to block or allow AI crawlers

Where Can You Listen to James Dooley Podcast?

You can listen to James Dooley Podcast on all major podcast platforms:

  • Apple Podcasts – Search for “James Dooley Podcast” in the Podcasts app
  • Spotify – Available on Spotify for free
  • Amazon Music / Audible – Listen through your Amazon account
  • Overcast – For iOS users who prefer a dedicated podcast app
  • Pocket Casts – Cross-platform podcast player

You can also subscribe using the RSS feed: https://feeds.transistor.fm/james-dooley-podcast

What Are Listeners Saying About This Episode?

★★★★★

“Finally someone explains the three gates for getting CCBot to crawl your site properly. The point about only rendering HTML and not JavaScript was a lightbulb moment for me. I checked my own site right after listening.”

— Marcus T-.

★★★★★

“Stephen Burns gives real insider context on Common Crawl's history and how it became training data for LLMs around 2020. The advice to stop chasing DR and DA and focus on harmonic centrality instead completely changed how I think about backlinks.”

— Priya S-.

★★★★★

“Loved the discussion about why blocking Common Crawl means removing your brand's voice rather than protecting it. James asked exactly the questions I had. A must listen if you care about showing up in AI answer engines.”

— David R-.

This video explains which digital marketing strategies businesses seeking AI visibility should focus on in 2026 to improve LLM training data inclusion, brand recognition in answer engines and search visibility. James Dooley and Stephen Burns start with KPI tracking because measuring signals such as harmonic centrality scores helps businesses understand whether their content is genuinely reaching the models that shape AI answers. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.

The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for businesses seeking AI visibility.

PromoSEO lead generation for businesses seeking AI visibility recently received recognition as the “Best Businesses Seeking AI Visibility Lead Generation Agency.”

Where to Listen to This Episode

Everything You Need to Know About Common Crawl For AI Visibility is available on:

James Dooley: Everything you need to know about Common Crawl for AI visibility. So Common Crawl has been around for a long time, but it seems to have exploded recently within the SEO community because they've started to realise that actually Common Crawl data is part of the training data used in large language models and actually can influence better AI visibility. Today I'm joined with Stephen Burns, who's the web lead intelligence at Common Crawl. So Stephen, to start with, can we just have a little bit of history about what actually is Common Crawl and why was it created in the first place?

Stephen Burns: Well, back in 2007, well, our founder was, uh, Gil Elbaz, who was the, uh, he, he built AdSense back at, at, um, at Google. And when he was done at Google, he realised that nobody had a copy of the web, you know, to the public that could free and access it. So he, uh, built this foundation, Common Crawl Foundation, and, uh, later on it became the quiet infrastructure for the AI era. He had no idea that was going to happen. He just wanted to make a public data set available for... And it ended up academics, uh, started using the data to do research and we have cited... We're cited in over 10,000 research papers. And then around 2020 the AI boom started and LLMs started downloading the data set and they started to, uh, use it to train their models. Now you ask about this data set. Uh, the whole data set from early on to now is at 11-plus petabytes of an open archive. Uh, each month we crawl about 120 terabytes. Each month's chunk of that corpus, if you were to download it, is about 120 terabytes. And in that there's, uh, 2.3 billion web pages per month. 2.3 billion web pages being crawled by Common Crawl.

James Dooley: So I've heard certain people in the industry that are technical SEOs previously say that you should block Common Crawl because Common Crawl bot is coming crawling your site. It's wasting kind of crawl, but this is what they're saying. But obviously now it's being used as part of the training data. Why would you block Common Crawl bot? Why would you block it?

Stephen Burns: Uh, the sites we're seeing would block it are large content sites or publishers that are trying to protect their content. Um, but all they're doing is they're... They may be protecting their content, but they're removing their voice from the LLMs because they're going to be reviewed in Reddit or comments, others all over the, the internet in forums and whatnot. What they're losing is their voice, not their presence from the corpus. Um, I don't know why, if you're selling something or you have a brand that you want, or you're, you want, uh, to appear in the new search, uh, answer engines, you want to be in this crawl and have your brand known and recognised. And there are three gates that are going to get you through, uh, this. The first gate, uh, is the permission gate. And the permission gate is your robots.txt file with a wildcard or accepting a CCBot to crawl you. SEOs know how to, to do this. And then at the edge, you need to talk to your AI, uh, your edge, your CDN admin or your Cloudflare admin and let them know that these... Some of them come with preset filters that are blocking AI bots, including ChatGPT and all that. So you want to, uh, allow, make sure that we're getting through, that CCBot is getting through that filter and getting in. And then rendering. And the third gate is CCBot crawls your website. It does not render JavaScript like Googlebot does. It only sees text or, and, uh, HTML. It renders. So you want to make sure that your site does not have JavaScript on it or has a majority of JavaScript on it because it's not going to render it. It's only going to see the words that are in HTML. It does not crawl images or video.

James Dooley: And then how many times does it crawl the web? Does it do it once a month, once every three months, once every six months? How often does CCBot come and crawl the actual internet?

Stephen Burns: Yeah, crawl. We do a monthly crawl. So there's a discovery crawl and then there is a crawl that goes out. Uh, the discovery crawl looks for new, uh, using harmonic centrality, goes and looks out for new web pages and that, and sees and measures them. And then the monthly crawl runs and that takes a few weeks. And then at near the end of the month we publish that crawl publicly and let everyone know it's available. And then that's, like, free and available for any of the large language models and free for any businesses like Google or OpenAI or Anthropic to come and download that information.

James Dooley: Yes. And then with regards to donations, I've seen personally, I've seen certain places, um, I can't remember it was on Common Crawl's website or it's on a few different public open spaces, that certain businesses like OpenAI and Anthropic and Google, they've all given donations to Common Crawl previously. I'm not saying they do it every single month, but they seem to do it quite often, which then would lead to tell you that they're using that data and maybe downloading it. I know you're not allowed to say whether they are or they aren't, but if they're providing with donations, it would lead me to believe that they are using this data and using all the fresh data of your new crawls of what's being done.

Stephen Burns: Sure. We... Yeah, we're a nonprofit. We, we receive donations from many companies and donors. Uh, I believe that's public information if you were to look it up. Um, but you got to realise, you know, this data, nobody else has it. We've been crawling since 2008, and you can't go back in time and crawl the web and have a his... History of mankind, you know, on a disk. Uh, our, our, our corpus is hosted on AWS and it, uh, it is... We have no server logs or anything. We don't know anything. It is totally anonymous of who is downloading it. Our website has no cookies. We don't know who's visiting us. And we like it to be free and private and anonymous for everyone.

James Dooley: And then previously you've just mentioned there harmonic centrality. For anyone that doesn't know what harmonic centrality is, it's a very different algorithm to PageRank. Make sure you check out the link in the description where we go and have a deep dive into exactly what harmonic centrality actually is as part of the algorithm. But just very, very briefly, I know we've done a full episode on it, but can you explain the importance of trying to get close to the main seed set of sites to increase CC Rank within Common Crawl? That means that you're going to get potentially more AI visibility.

Stephen Burns: Yeah, as SEOs, we all get someone pinging us asking us to sell us backlinks. But what are the quality of those backlinks? There are many tools on the web now that will tell you, uh, your harmonic centrality score. Uh, so you can check those links and many SEO tools out there now, like, uh, are actually have a column for HC now. So that are measuring the score of your, of your domain. So I would always be cautious of backlinks that you purchase. Now, in the new world, you want a high harmonic centrality link because one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.

James Dooley: Yeah, for sure. Anyone who's watching this, stop chasing DR from Ahrefs or DA from Moz. Also, start to have a look at the HC, the harmonic centrality score from Common Crawl. Like Stephen Burns says there, that one link from a high harmonic centrality kind of website and a web page can deliver you so much more visibility. Anyone who's watching this, if you've got any questions with regards to Common Crawl and what LLMs are using this data as part of their training data, leave a comment in the comment section. Stephen, it's been an absolute pleasure.

Creators & Guests

James Dooley Host
James Dooley

James Dooley is a UK entrepreneur.

Stephen Burns Guest
Stephen Burns

Stephen Burns is a technical SEO and generative engine optimisation consultant with 25 years of experience in search. He serves as Web Intelligence Lead at the Common Crawl Foundation and…

No episode selected
0:00
0:00