Meta Is Building Its Own Web Search Index: What It Means for AI Search, SEO, and Agencies

Published August 17, 2026By ABD Legacy LLC
AI search / SEO / agency discovery

On August 6, 2026, Pieter Levels (@levelsio) published a DM from an anonymous Meta staffer: Meta is allegedly building its own Google-scale search engine and web index so that when its AI does a web search, the query "doesn't end up at Google" -- where a competitor could use it for training. Source: the levelsio post on X.

Before this becomes another AI rumor, here is what is actually verified. Meta's own developer documentation (updated May 21, 2026) lists a dedicated crawler, Meta-WebIndexer, whose stated purpose is "to improve Meta AI search result quality for users" and to "cite and link to your content in Meta AI's responses." A second crawler, Meta-ExternalAgent, "crawls the web for use cases such as training foundation AI models or improving products by indexing content directly." Source: Meta Web Crawlers documentation.

None of this is new. The Information and Reuters reported in October 2024 that Meta was developing an AI search engine to reduce reliance on Google and Microsoft. Search Engine Journal covered the same story, noting Meta was "creating a search index for its AI chatbot." By August 2025, Meta had confirmed to agency executives that it was working on a new search product -- while Meta AI still pulled answers from Google and Bing. In June 2026, Meta shipped "AI Mode" search on Facebook, surfacing posts, groups, and reels.

The levelsio DM is consistent with that public record but adds no independently verifiable facts. The staffer is unnamed, no DM screenshot was shown, and the specific motive -- keeping Meta's AI searches away from Google's training data -- is single-sourced and unconfirmed. Treat that part as a plausible explanation, not a finding. What is confirmed: Meta runs a web-indexing crawler built for AI search, has been reported building its own index for nearly two years, and has told agencies a search product is coming.

If Meta's index becomes the backbone of Meta AI answers, here is what changes -- and what agencies should do about it.

Why Meta would build its own search index

The business logic is not hard to follow, and it predates the DM. Every time Meta AI answers a query by pulling from Google or Bing, Meta is paying a competitor for the underlying data and exposing its users' questions to a rival's infrastructure. A web index is the standard answer: own the crawl, own the answers, own the citations. Meta's crawler documentation says the index exists "to improve Meta AI search result quality" and to cite sources in AI responses -- which is exactly what a search product needs. The crawl volumes are consistent with an index being built, not a test: Marc Lou reported 67,390 Meta crawls in three days on TrustMRR (roughly 94x OpenAI's volume on the same site), and Levels' own server logs showed meta-webindexer and meta-externalagent dominating his request table. Discussion on Hacker News.

What a Meta index means for AI search results

The visible effect is a second, serious answer engine. Today, when someone asks an AI assistant a question, the answers largely come from Google and Bing-indexed content, OpenAI's stack, and a few others. A Meta-owned index would let Meta AI answer from its own crawl -- and cite the pages it chooses, on its own ranking rules. That matters for two reasons. First, the citation decision is now Meta's: content that ranks on Google may or may not be what Meta AI decides to cite. Second, it fragments the "AI answer" market. Agencies can no longer optimize for one answer engine and assume the rest follow.

Facebook's June 2026 AI Mode launch shows the distribution play: search surfaced inside a platform with three billion users, starting with its own content (posts, groups, reels) and now with a web index behind it. An AI answer with a citation is a referral -- and the referral source of the future may be Meta AI, not Google.

How SEO shifts when Meta runs an index

Three structural changes, none speculative:

  1. Citations become a ranking currency. Meta's docs say allowing Meta-WebIndexer in robots.txt "helps us cite and link to your content in Meta AI's responses." That is an explicit invitation: crawler access influences whether you get cited. Google ranking alone no longer guarantees AI visibility.
  2. Two crawlers, two decisions. Meta-WebIndexer (AI search citations) and Meta-ExternalAgent (training and indexing) are separate. An agency's robots.txt policy can allow the citation crawler while restricting the training crawler -- a governance decision every site owner will need to make deliberately.
  3. Crawl governance becomes a real cost. At 67,390 requests in three days, Meta's crawling is not polite traffic; it is infrastructure-level load. Sites will need log monitoring, rate handling, and a policy that doesn't accidentally block the citation crawler while trying to limit the load.

How agencies can prepare (concrete actions)

  1. Add Meta AI to your answer-engine optimization (AEO) checklist. Verify Meta-WebIndexer is allowed in robots.txt, confirm the crawler appears in server logs, and check whether your content (or your client's) is cited in Meta AI responses. If it is not, that is an optimization target, not a mystery.
  2. Diversify visibility beyond Google. If Meta AI stops routing through Google/Bing, Google-only rankings will not get you cited in Meta AI. Build content that earns citations across multiple answer engines -- Meta AI, ChatGPT, Perplexity, Google AI Overviews -- instead of optimizing for one.
  3. Offer crawl-governance as a service line. Every client with a website now faces a robots.txt decision with real consequences (citation access vs. training use vs. server load). Auditing and configuring crawler policy is a concrete, billable deliverable for AI-era agencies.
  4. Track citation share, not just traffic. As AI answers with citations replace click-throughs, the KPI shifts from sessions to mentions: how often does your brand appear as a cited source in AI responses, and is that share growing? Build the tracking now, before the channel matters.
  5. Watch Meta AI as a discovery channel for your own agency. If Meta AI becomes a referral source, being findable there -- via your own optimized content -- is how new clients will find you. The same playbook you sell to clients applies to your agency's pipeline.

The bottom line

Meta is building a web search index for AI search. That is confirmed by Meta's own crawler documentation, nearly two years of reporting, and agency-facing confirmations -- and it was amplified, not revealed, by the levelsio DM. The unanswered questions are timing, scale, and whether Google/Bing remain in the stack. The preparation is not speculative: allow the citation crawler, diversify across answer engines, govern your robots.txt deliberately, and measure citations, not just clicks. Agencies that treat AI search as a multi-engine reality -- and their own discoverability as an asset -- will be the ones clients find when the index ships.

Compare agencies that already treat AI search as a multi-engine game.

Browse AI Agencies →
Audit your own AI visibility before the index lands →

Sources