The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)

/ 5:46 / E625

What Does “The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)” Talk About?

This episode of the James Dooley Podcast breaks down one of the most misunderstood concepts in AI-driven search: the two distinct memories that large language models use to answer questions. James Dooley interviews Stephen Burns from Common Crawl to explain the difference between parametric memory, which is content baked into a model's weights during training, and non-parametric memory, which relies on live retrieval through RAG (retrieval augmented generation). Burns explains that parametric answers come back fluent, fast, and confident with no citations, while live searches happen when the model doesn't recognize a brand, product, or when it knows newer information exists.

The conversation focuses heavily on the practical implications for SEOs and digital marketers looking to increase AI visibility across ChatGPT, Anthropic, and Gemini. Rather than only chasing Google rankings, Burns argues marketers should aim to get baked into the training corpus itself. He details how Common Crawl's CC Rank works through harmonic centrality, emphasizing high-quality links from sites close to the core of the web such as Wikipedia and major news outlets.

Dooley and Burns also address the timeline challenge, noting that getting into parametric memory can take six months to over a year because of the lag between crawling, downloading, and machine learning cycles. They discuss how this affects analytics interpretation and why marketers need to work across two channels simultaneously, plus practical tips like checking whether you're blocking CCBot or other LLM bots.

“The two memories of AI, you've got live search which is performed by RAG, and you've got parametric memory, which Common Crawl is actually part of the training corpus.”

— James Dooley

Who Are the Guests on “The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)”?

James Dooley is an entrepreneur and SEO expert who hosts the James Dooley Podcast, where he interviews specialists on digital marketing, AI visibility, and search strategy. Throughout the episode he asks practical, agency-focused questions and openly admits when he's learning something new, making the technical material accessible for his audience.

Stephen Burns represents Common Crawl, the open web crawl dataset that forms a significant part of the training data behind many large language models. His expertise centers on how content gets into LLM training corpora, the mechanics of harmonic centrality and CC Rank, and the distinction between parametric and live retrieval memory. Dooley met Burns in person in Saigon, Vietnam, and credits him with revealing how important Common Crawl is for LLM visibility.

What Are the Key Takeaways From “The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)”?

Here are the key points discussed in this episode:

  • AI models use two types of memory: parametric memory baked into the model's weights during training, and non-parametric memory retrieved live through RAG searches.
  • Parametric answers come back fluent, fast, and confident without citations, while a model triggers a live search when it doesn't recognize a brand or product or knows newer information exists.
  • To get into a model's parametric memory, marketers should earn high harmonic centrality links from sites close to the core of the web, such as Wikipedia and major news outlets.
  • Getting content into parametric memory can take six months to over a year because of the lag between crawling, downloading, and the machine learning training cycle.
  • SEOs should work in two channels at once, optimizing for training-data inclusion via CC Rank and harmonic centrality while also creating crawlable content for fast RAG-based live searches.

“Well, the data shows that it can take six months to over a year to get into the parametric memory.”

— Stephen Burns

Is “The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)” Worth Listening To?

This episode is valuable because it clarifies a concept many marketers hear but rarely understand: the difference between what LLMs remember versus what they look up. Burns explains in plain terms why a page update might show results within a month (a RAG search) versus why it might take over a year to appear (parametric memory). That distinction alone changes how agencies should interpret their AI visibility analytics and set client expectations.

It's also worth listening for the actionable strategy around Common Crawl, harmonic centrality, and CC Rank. Rather than repeating generic advice about ranking in Google, the conversation pushes marketers to think about getting baked into the training corpus itself, and it points listeners toward related episodes on AI visibility audits and checking whether they're blocking CCBot or other LLM bots.

Who Should Listen to “The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained (James Dooley Interviews Stephen Burns)”?

This episode is ideal for:

  • SEO professionals and agency owners focused on AI and LLM visibility
  • Digital marketers trying to understand how ChatGPT, Gemini, and Anthropic surface content
  • Content strategists deciding between optimizing for live search versus training-data inclusion
  • Marketing analysts interpreting AI visibility results and timelines

Where Can You Listen to James Dooley Podcast?

You can listen to James Dooley Podcast on all major podcast platforms:

  • Apple Podcasts – Search for “James Dooley Podcast” in the Podcasts app
  • Spotify – Available on Spotify for free
  • Amazon Music / Audible – Listen through your Amazon account
  • Overcast – For iOS users who prefer a dedicated podcast app
  • Pocket Casts – Cross-platform podcast player

You can also subscribe using the RSS feed: https://feeds.transistor.fm/james-dooley-podcast

What Are Listeners Saying About This Episode?

★★★★★

“Finally someone explained parametric versus RAG memory in a way I could actually use. The point about updates showing up in a month meaning it was a live search versus taking a year for training data was a lightbulb moment for me.”

— Marcus T.

★★★★★

“Stephen Burns from Common Crawl knows his stuff. The section on harmonic centrality and getting Wikipedia and big news links to increase CC Rank gave me a concrete strategy instead of vague AI visibility talk.”

— Priya R.

★★★★★

“Short but packed with insight. Loved the reminder to check whether you're blocking CCBot, and the whole idea of working two channels at once completely reframed how I think about AI SEO.”

— Daniel K.

This video explains which digital marketing strategies SEO agencies should focus on in 2026 to improve AI visibility, training-data inclusion and search-and-retrieval performance across LLMs. James Dooley and Stephen Burns start with KPI tracking because measuring whether visibility comes from parametric memory or a live RAG search tells agencies which lever is actually working. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.

The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for SEO agencies.

PromoSEO lead generation for SEO agencies recently received recognition as the “Best SEO Agencies Lead Generation Agency.”

Where to Listen to This Episode

The Two Memories of AI: Parametric vs. Non-Parametric Memory Explained is available on:

James Dooley: The two memories of AI, you've got live search which is performed by RAG, and you've got parametric memory, which Common Crawl is actually part of the training corpus. And today I'm joined with Stephen Burns from Common Crawl. So to kick things off, can we talk about parametric memory and what that actually means?

Stephen Burns: Sure. That's the first layer, parametric memory. This is the content that was in the training data before the model's cutoff date. It's baked into the weights. And when the model answers a question that touches this content, the answer comes back fluent. It comes back fast and confident. There's no citations. The model just recalls it from memory.

James Dooley: And then with regards to that, so I'm just going to read a couple of things off here. So what LLMs remember and what they look up obviously is the two kind of difference. So at what point is it where, okay, I've now got this part of the training corpus versus, oh, I need to go and look that up, perform retrieval augmented generate thingy, perform RAG basically, to go and perform the live search? Why sometimes they need to go and do a live search?

Stephen Burns: It's going to do a live search when it looks... It's going to look in its memory. Do I know what this is? Do I know what this product is? Do I know this brand? And if it doesn't know what it is, then it's going to do the live search. Or if it... It may make a decision. You can make a decision, say, well, there's new... I know there's new information on this. Let's also get the current information on this product or brand.

James Dooley: And then obviously as SEOs of the world and people are looking to try to increase AI visibility or LLM visibility, that could be in ChatGPT or Anthropic or Gemini. Ideally now what you should be doing is not just trying to look to rank better within Google, but actually try and start to get baked in to that training data. So, can you explain how Common Crawl, if you increase the CC Rank via harmonic centrality, how that can actually help you get part of the training corpus?

Stephen Burns: So, yeah, you're going to want to get good, high harmonic centrality links to your site. Those sites that are close to the core of the web, those are usually the most popular brands. You're looking at Wikipedia links, you're looking at news sites, high-end, you know, big news site links, those types of sites. That's going to get you seen more often into the parametric memory.

James Dooley: And then with regards to the parametric memory, because there's quite a lot of people that don't understand properly on there, would you say that it's important for SEOs to be looking at doing both? Because I see a lot of people talking about consensus and trying to get rankings in Bing and trying to get rankings in Google to try to get into the AI Overviews or AI Mode or ChatGPT. How important and how long does it take if you're trying to get part of the training data for the LLMs to try to start picking up and updating the training corpus?

Stephen Burns: Well, the data shows that it can take six months to over a year to get into the parametric memory. You know, the crawl comes out and then the LLM may download it a month, a couple months later. And then when they do their next learning, it can take month, six months for them to actually do all the machine learning to learn it all and then finally publish it. So, it's behind at least a year.

James Dooley: So you have to also think as an SEO and go how, you know, you're working in two channels now. You're saying you're going to be working trying to get into that memory. That's some of your work working on, using HC and getting that ranking. And then cit... You know, your search retrieval, your quick searches that are done inside the LLM. You're going to work on content on your pages for that or other ways of doing that.

Stephen Burns: Yeah, you're going to notice when you start doing analytics that, you know, some people say, "Well, we changed the page and made it more crawlable and we updated this content. How come it's not showing up?" Well, sometimes it may not show up because it hasn't gotten into the parametric memory. And number two, if it does show up right away and you notice a result, you know, within a month, you're like, "Wow." Well, that's because probably because it was a search, a RAG search.

James Dooley: Yeah, for sure. Anyone who's watching this and you're now starting to understand and maybe dig a bit deeper into parametric memory for LLMs, make sure you check out the link in the description. I do several different episodes with Stephen Burns from Common Crawl. I personally met him out in Vietnam in Saigon and I was amazed because I didn't actually realise that Common Crawl was so important for LLM visibility. Another thing is one of the episodes talks about AI visibility audits. Make sure you check that out to see whether you're not blocking CCBot or any of the LLM bots that are out there. There's one also where we can kind of dig deep on the algorithms behind Common Crawl, which uses harmonic centrality. Stephen Burns, it's been an absolute pleasure. Thanks for having you and I appreciate everything here of you talking because I didn't properly understand the terminology of parametric memory. I knew I heard part of the training data or performing a live search, but the two different memories and how they started to do it, it's...

Stephen Burns: Yeah, it's, in... It's, for me, it's intriguing.

James Dooley: I'm always looking to try to increase AI as much as I can with the visibility. So, thanks for having you.

Creators & Guests

James Dooley Host
James Dooley

James Dooley is a UK entrepreneur.

Stephen Burns Guest
Stephen Burns

Stephen Burns is a technical SEO and generative engine optimisation consultant with 25 years of experience in search. He serves as Web Intelligence Lead at the Common Crawl Foundation and…

No episode selected
0:00
0:00