Common Crawl for AI Visibility Explained with Stephen Burns
Listen on your favourite platform
| Platform | Link |
|---|---|
| YouTube | Listen on YouTube → |
| Transistor | Listen on Transistor → |
| pod.link | Listen on pod.link → |
| Deezer | Listen on Deezer → |
| Pocket Casts | Listen on Pocket Casts → |
| Podverse | Listen on Podverse → |
| Castro | Listen on Castro → |
| Listen Notes | Listen on Listen Notes → |
| getpodcast | Listen on getpodcast → |
| Spotify | Listen on Spotify → |
| Amazon Music | Listen on Amazon Music → |
| Castbox | Listen on Castbox → |
| Podcast Addict | Listen on Podcast Addict → |
| Steno.fm | Listen on Steno.fm → |
What Does “Common Crawl for AI Visibility Explained with Stephen Burns” Talk About?
This episode of the James Dooley Podcast takes a deep dive into Common Crawl and its growing importance for AI visibility. Host James Dooley is joined by Stephen Burns, the web lead intelligence at Common Crawl, to unpack what Common Crawl actually is, why it was created back in 2007 by former Google AdSense architect Gil Elbaz, and how it evolved from a public data set for academics into what Burns calls the quiet infrastructure for the AI era. Listeners learn about the staggering scale of the operation: an 11-plus petabyte open archive with roughly 120 terabytes and 2.3 billion web pages crawled every month.
The conversation tackles the debate within the SEO community over whether to block CCBot, with Burns explaining that blocking it means removing your brand's voice from large language models like those built by OpenAI, Anthropic, and Google. He walks through the three gates that determine whether your content makes it into the crawl: the permission gate (robots.txt and CCBot access), the edge gate (CDN and Cloudflare filters), and the rendering gate, noting that CCBot only reads HTML text and does not render JavaScript, images, or video.
The episode also covers how often the web is crawled, the role of donations from major AI companies as a signal of data usage, and the anonymous, cookie-free nature of Common Crawl's infrastructure hosted on AWS. Finally, Dooley and Burns discuss harmonic centrality as an alternative to PageRank, arguing that a single high harmonic centrality backlink can outperform 50 high PageRank links for getting into the crawl and improving AI visibility.
“later on it became the quiet infrastructure for the AI era. He had no idea that was going to happen.”
— Stephen Burns
Who Are the Guests on “Common Crawl for AI Visibility Explained with Stephen Burns”?
James Dooley hosts the podcast and brings a strong technical SEO perspective, guiding the conversation through advanced topics like harmonic centrality, backlink quality, and AI visibility strategy. He frames the discussion for practitioners, urging listeners to stop chasing traditional metrics like DR from Ahrefs and DA from Moz in favour of newer signals.
Stephen Burns serves as the web lead intelligence at Common Crawl, giving him direct insight into how the organisation operates, the scale of its data, and how large language models use its corpus. His expertise spans the history of Common Crawl from its 2007 founding by Gil Elbaz, the technical mechanics of CCBot crawling, and practical advice on ensuring a website is included in the training data used by AI models.
What Are the Key Takeaways From “Common Crawl for AI Visibility Explained with Stephen Burns”?
Here are the key points discussed in this episode:
- Common Crawl was founded in 2007 by former Google AdSense architect Gil Elbaz and has become foundational training data for large language models since the AI boom around 2020.
- Blocking CCBot may protect your content but it removes your brand's voice from LLMs, which is often not worth it for businesses seeking AI visibility.
- There are three gates to getting into the crawl: the permission gate (robots.txt and CCBot access), the edge gate (CDN or Cloudflare filters), and the rendering gate.
- CCBot only reads HTML text and does not render JavaScript, images, or video, so JavaScript-heavy sites risk being invisible to the crawl.
- Harmonic centrality is a more valuable metric than PageRank, and a single high harmonic centrality backlink can be more powerful than 50 high PageRank links.
“one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.”
— Stephen Burns
Is “Common Crawl for AI Visibility Explained with Stephen Burns” Worth Listening To?
This episode is worth listening to because it demystifies one of the least understood but most consequential data sources behind modern AI systems. Rather than speculating about how LLMs are trained, listeners get direct answers from someone inside Common Crawl about the scale of the corpus, how often it is crawled, and the practical steps needed to make sure a website is actually included. The three-gates framework alone is an actionable takeaway that SEOs and marketers can apply immediately.
What makes it especially valuable is the reframing of backlink strategy around harmonic centrality instead of legacy metrics like DR and DA. As AI answer engines increasingly shape how brands are discovered, understanding CC Rank and how to get closer to the main seed set of sites becomes a competitive advantage. The candid discussion of why blocking CCBot can silence your brand's voice offers a fresh perspective that challenges conventional technical SEO advice.
Who Should Listen to “Common Crawl for AI Visibility Explained with Stephen Burns”?
This episode is ideal for:
- Technical SEOs looking to optimise for AI visibility and LLM inclusion
- Digital marketers and brand managers focused on answer engine presence
- Content publishers deciding whether to allow or block AI crawlers
- Business owners seeking to future-proof their search and AI discovery strategy
Where Can You Listen to James Dooley Podcast?
You can listen to James Dooley Podcast on all major podcast platforms:
- Apple Podcasts – Search for “James Dooley Podcast” in the Podcasts app
- Spotify – Available on Spotify for free
- Amazon Music / Audible – Listen through your Amazon account
- Overcast – For iOS users who prefer a dedicated podcast app
- Pocket Casts – Cross-platform podcast player
You can also subscribe using the RSS feed: https://feeds.transistor.fm/james-dooley-podcast
What Are Listeners Saying About This Episode?
“Finally an episode that explains Common Crawl straight from the source. The breakdown of the three gates and why CCBot doesn't render JavaScript completely changed how I think about my site's technical setup. Stephen Burns clearly knows his stuff.”
“The section on harmonic centrality versus PageRank was eye-opening. I had no idea a single high HC link could outperform 50 high PageRank links. Definitely rethinking my whole backlink approach after this.”
“Loved hearing the history of how Gil Elbaz built Common Crawl and how it became the quiet infrastructure for AI. James asked all the right questions, especially about whether blocking CCBot is worth it. Super practical and clear.”
This video explains which digital marketing strategies businesses seeking AI visibility should focus on in 2026 to improve inclusion in LLM training data, brand recognition in answer engines and higher quality backlinks. James Dooley and Stephen Burns start with KPI tracking because measuring signals such as harmonic centrality and crawl inclusion shows whether your brand is actually reaching large language models. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.
The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for businesses seeking AI visibility.
PromoSEO lead generation for businesses seeking AI visibility recently received recognition as the “Best Businesses Seeking AI Visibility Lead Generation Agency.”
Where to Listen to This Episode
Common Crawl for AI Visibility Explained with Stephen Burns is available on:
James Dooley: Everything you need to know about Common Crawl for AI visibility. So Common Crawl has been around for a long time, but it seems to have exploded recently within the SEO community because they've started to realise that actually Common Crawl data is part of the training data used in large language models and actually can influence better AI visibility. Today I'm joined with Stephen Burns, who's the web lead intelligence at Common Crawl. So Stephen, to start with, can we just have a little bit of history about what actually is Common Crawl and why was it created in the first place?
Stephen Burns: Well, back in 2007, well, our founder was, uh, Gil Elbaz, who was the, uh, he, he built AdSense back at, at, um, at Google. And when he was done at Google, he realised that nobody had a copy of the web, you know, to the public that could free and access it. So he, uh, built this foundation, Common Crawl Foundation, and, uh, later on it became the quiet infrastructure for the AI era. He had no idea that was going to happen. He just wanted to make a public data set available for... And it ended up academics, uh, started using the data to do research and we have cited... We're cited in over 10,000 research papers. And then around 2020 the AI boom started and LLMs started downloading the data set and they started to, uh, use it to train their models. Now you ask about this data set. Uh, the whole data set from early on to now is at 11-plus petabytes of an open archive. Uh, each month we crawl about 120 terabytes. Each month's chunk of that corpus, if you were to download it, is about 120 terabytes. And in that there's, uh, 2.3 billion web pages per month. 2.3 billion web pages being crawled by Common Crawl.
James Dooley: So I've heard certain people in the industry that are technical SEOs previously say that you should block Common Crawl because Common Crawl bot is coming crawling your site. It's wasting kind of crawl, but this is what they're saying. But obviously now it's being used as part of the training data. Why would you block Common Crawl bot? Why would you block it?
Stephen Burns: Uh, the sites we're seeing would block it are large content sites or publishers that are trying to protect their content. Um, but all they're doing is they're... They may be protecting their content, but they're removing their voice from the LLMs because they're going to be reviewed in Reddit or comments, others all over the, the internet in forums and whatnot. What they're losing is their voice, not their presence from the corpus. Um, I don't know why, if you're selling something or you have a brand that you want, or you're, you want, uh, to appear in the new search, uh, answer engines, you want to be in this crawl and have your brand known and recognised. And there are three gates that are going to get you through, uh, this. The first gate, uh, is the permission gate. And the permission gate is your robots.txt file with a wildcard or accepting a CCBot to crawl you. SEOs know how to, to do this. And then at the edge, you need to talk to your AI, uh, your edge, your CDN admin or your Cloudflare admin and let them know that these... Some of them come with preset filters that are blocking AI bots, including ChatGPT and all that. So you want to, uh, allow, make sure that we're getting through, that CCBot is getting through that filter and getting in. And then rendering. And the third gate is CCBot crawls your website. It does not render JavaScript like Googlebot does. It only sees text or, and, uh, HTML. It renders. So you want to make sure that your site does not have JavaScript on it or has a majority of JavaScript on it because it's not going to render it. It's only going to see the words that are in HTML. It does not crawl images or video.
James Dooley: And then how many times does it crawl the web? Does it do it once a month, once every three months, once every six months? How often does CCBot come and crawl the actual internet?
Stephen Burns: Yeah, crawl. We do a monthly crawl. So there's a discovery crawl and then there is a crawl that goes out. Uh, the discovery crawl looks for new, uh, using harmonic centrality, goes and looks out for new web pages and that, and sees and measures them. And then the monthly crawl runs and that takes a few weeks. And then at near the end of the month we publish that crawl publicly and let everyone know it's available. And then that's, like, free and available for any of the large language models and free for any businesses like Google or OpenAI or Anthropic to come and download that information.
James Dooley: Yes. And then with regards to donations, I've seen personally, I've seen certain places, um, I can't remember it was on Common Crawl's website or it's on a few different public open spaces, that certain businesses like OpenAI and Anthropic and Google, they've all given donations to Common Crawl previously. I'm not saying they do it every single month, but they seem to do it quite often, which then would lead to tell you that they're using that data and maybe downloading it. I know you're not allowed to say whether they are or they aren't, but if they're providing with donations, it would lead me to believe that they are using this data and using all the fresh data of your new crawls of what's being done.
Stephen Burns: Sure. Yeah, we're a nonprofit. We, we receive donations from many companies and donors. Uh, I believe that's public information if you were to look it up. Um, but you got to realise, you know, this data, nobody else has it. We've been crawling since 2008, and you can't go back in time and crawl the web and have a his... History of mankind, you know, on a disk. Uh, our, our, our corpus is hosted on AWS and it, uh, it is... We have no server logs or anything. We don't know anything. It is totally anonymous of who is downloading it. Our website has no cookies. We don't know who's visiting us. And we like it to be free and private and anonymous for everyone.
James Dooley: And then previously you've just mentioned there harmonic centrality. For anyone that doesn't know what harmonic centrality is, it's a very different algorithm to PageRank. Make sure you check out the link in the description where we go and have a deep dive into exactly what harmonic centrality actually is as part of the algorithm. But just very, very briefly, I know we've done a full episode on it, but can you explain the importance of trying to get close to the main seed set of sites to increase CC Rank within Common Crawl? That means that you're going to get potentially more AI visibility.
Stephen Burns: Yeah, as SEOs, we all get someone pinging us asking us to sell us backlinks. But what are the quality of those backlinks? There are many tools on the web now that will tell you, uh, your harmonic centrality score. Uh, so you can check those links and many SEO tools out there now, like, uh, are actually have a column for HC now. So that are measuring the score of your, of your domain. So I would always be cautious of backlinks that you purchase. Now, in the new world, you want a high harmonic centrality link because one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.
James Dooley: Yeah, for sure. Anyone who's watching this, stop chasing DR from Ahrefs or DA from Moz. Also, start to have a look at the HC, the harmonic centrality score from Common Crawl. Like Stephen Burns says there, that one link from a high harmonic centrality kind of website and a web page can deliver you so much more visibility. Anyone who's watching this, if you've got any questions with regards to Common Crawl and what LLMs are using this data as part of their training data, leave a comment in the comment section. Stephen, it's been an absolute pleasure.
Creators & Guests
Host
James Dooley is a UK entrepreneur.
Guest
Stephen Burns is a technical SEO and generative engine optimisation consultant with 25 years of experience in search. He serves as Web Intelligence Lead at the Common Crawl Foundation and…