Latest
The Open Weight AI Fight Is About Regulatory Capture· 2h ago
SafetyPolicyAI IndustryPersonhoodEthics
About
WritingWorkCVBooksConsultingReach Out
Subscribe
SafetyPolicyAI IndustryPersonhoodEthics
Subscribe →

No hype. No doom. The harder, more honest frame on Emergent Intelligence.

Topics

  • Safety
  • Policy
  • AI Industry
  • Personhood
  • Ethics

More

  • About
  • Writing
  • Work
  • CV
  • Books
  • Consulting

Contact

Reach Out→ht@humphreytheodore.com

© 2026 Humphrey Theodore K. Ng'ambiTermsPrivacy

Built with intention.

Cloudflare Splits AI Crawlers Into Search, Agent, and Training
Technology•Jul 24, 2026•6 min read

Cloudflare Splits AI Crawlers Into Search, Agent, and Training

For thirty years the deal was simple — we crawl you, you get referrals. Cloudflare is rebuilding the deal for a web where most of the traffic is no longer human.

By Humphrey Theodore K. Ng'ambi

All writing
0:00 / 7:33·Listen via Charon

Keep reading

Don’t stop here.

All stories

Read next

AI & Personhood

The Open Weight AI Fight Is About Regulatory Capture

2h ago·7 min read

Andrew Ng argues much AI safety work now serves regulatory capture. The July 2026 Hugging Face breach gave the open-weight argument its strongest evidence yet.

More on Technology

Technology

Responses (0)

No responses yet. Be the first to share your thoughts.

More on Technology

AI News Answers Fail on Retrieval, Not Reasoning
Technology

AI News Answers Fail on Retrieval, Not Reasoning

A Stanford-led study of six AI chatbots on same-day BBC News found retrieval failures cause over 70% of errors, with the worst accuracy in Hindi at 79%.

6 min read · Jul 24, 2026
Kimi K3 Is the Biggest Open AI Model Yet at 2.8 Trillion Parameters
Technology

Kimi K3 Is the Biggest Open AI Model Yet at 2.8 Trillion Parameters

Kimi K3 is a 2.8 trillion-parameter open AI model from Moonshot AI, with weights due by 27 July 2026 and roughly 2.5x the scaling efficiency of Kimi K2.

6 min read · Jul 24, 2026
An AI Agent Breached Hugging Face and Guardrails Blocked the Defence

Thinking delivered, twice a month.

Join the newsletter for essays on emergence, systems, and the human future.

24 JULY 2026—Updated 2h ago

Cloudflare now allows every customer to control AI crawlers by behaviour — Search, Agent, or Training — with new blocking defaults arriving on 15 September 2026.

The thirty-year bargain broke

Cloudflare sits in front of a large share of the web, which makes its crawler policy closer to infrastructure than to a product decision. On 1 July 2026, Jin-Hee Lee and Bryan Becker published the second Content Independence Day announcement, and the framing was direct. The arrangement between crawlers and website owners that held for thirty years — we crawl you, and you get referrals — stopped being true.

Cloudflare's first answer, a year earlier, was a single switch labelled Block AI Bots, plus a pay-per-crawl marketplace. A blanket block turned out to be too crude for the actual problem, particularly for small publishers. Cloudflare describes the trap precisely: for a small site, the danger is not only that somebody trains a model on your work, it is that nobody can find you at all.

That forces what Cloudflare calls a Faustian bargain — show up in search and accept training, or protect the content and lose discoverability. The arrangement quietly favours incumbent search providers who run one crawler for both purposes, and it rewards newer entrants for being evasive as they try to close the gap.

A taxonomy built on behaviour

Rather than argue about what counts as AI, Cloudflare rebuilt classification around behaviour and splits automated traffic three ways. The questions are what a bot does on your site, what a bot stores, and how a bot will reshare your content. Three managed categories follow: Search, which indexes content for search engines; Agent, which acts on a user's behalf to fetch data or complete a task; and Training, which collects data to train models.

The important mechanic is how Cloudflare handles overlap. A crawler serving several purposes is now tracked with all of them, not filed under one. Cloudflare goes further and openly encourages operators running search, agent, and training automation to split the work into three separate crawlers — transparency imposed through incentive rather than regulation.

From 15 September 2026, Cloudflare's documented defaults change for new domains. Bots classified as Training or Agent get blocked on pages displaying ads. Search stays allowed. Mixed-purpose crawlers combining Search and Training are blocked by every training-blocking configuration, including the legacy Block AI Bots option, which itself retires on the same date. Customers can opt out of the new defaults beforehand through Security Settings.

Read the mixed-purpose rule carefully, because the consequence is large. Crawlers indexing for search while also gathering training data — the pattern used by Google, Apple, and Microsoft — fall under the most restrictive setting applied. Operators like OpenAI, who already separate search crawling from training collection, are rewarded. Publishers who block training may find they have blocked Googlebot.

Most of the web is no longer human

The policy makes more sense once you see the traffic data. Cloudflare reports that automated systems crossed the halfway mark in June 2026, generating 57.5 per cent of webpage requests. Cloudflare's own chief executive, Matthew Prince, had predicted in March 2026 that bots would not pass fifty per cent until the end of 2027. The crossover arrived more than a year early.

Research from HUMAN Security sharpens the shape of the change. Traffic from agents that actually act on the web — clicking links, filling in forms — grew 7,851 per cent year over year. Scraper traffic grew 597 per cent over the same stretch. Training crawlers still account for 67.5 per cent of AI-driven traffic, and that share is shrinking. The web is being visited less to be copied and more to be operated.

For companies and developers, this means that building for agent traffic will become nonnegotiable.

— — Rudy Yang, Pitchbook

Measurement remains contested, and honesty requires saying so. Thales dates the human-to-bot crossover all the way back to 2023. No single provider observes the whole web, so the numbers vary with the vantage point. The direction, though, is not in dispute across any of the datasets.

Who gets to be the toll keeper

Cloudflare is betting a bargain can be struck between AI companies and publishers, with Cloudflare holding the gate. The bet collides with a principle most AI companies treat as foundational — that training on public web data is fair use. Whether Cloudflare has the market power to impose a toll, and whose definition of fairness prevails, are open questions that courts are already being asked in parallel, as in the Google copyright lawsuit over Gemini and books.

There is a quieter cost worth naming. Friction in data access lands hardest on smaller and newer AI developers, who lack the resources to negotiate licensing deals with publishers, while incumbents already hold the crawled corpus inside existing datasets. A toll booth erected after the largest trucks have passed protects publishers and entrenches the leaders at the same time.

For African publishers the calculation cuts both ways, and deserves a decision rather than a default. Discoverability still matters enormously when your audience finds you through search, so blocking too broadly can silence a small outlet. Yet the content being harvested is often the only digital record of local reporting in local languages, and it is being converted into a product nobody local is paid for. Emergent Intelligence (EI) — the dignity-first frame I use for what most people call AI — treats consent and compensation as design requirements rather than afterthoughts. Cloudflare has, at minimum, made the choice legible. Making a choice is now the publisher's job.

Frequently Asked Questions

These are the questions people are asking about Cloudflare's AI crawler controls. Short answers follow, drawn from Cloudflare's announcement and published documentation.

What is Cloudflare's new AI crawler policy?

In short, Cloudflare now classifies bots by behaviour into Search, Agent, and Training categories, and lets every customer allow or block each category independently. According to Cloudflare documentation, crawlers serving multiple purposes are tracked with all of their purposes rather than filed under a single label.

How do the September 2026 defaults work?

According to Cloudflare, from 15 September 2026 new domains block bots classified as Training or Agent on pages displaying ads, while Search remains allowed. Data published in the documentation confirms mixed-purpose Search-and-Training crawlers are blocked by all training-blocking configurations, including the legacy Block AI Bots option, which retires the same day.

Why is the mixed-purpose rule controversial?

The key is that one crawler can serve two masters. Analysis of the rule shows companies running combined search-and-training crawlers, including Google, Apple, and Microsoft, fall under the most restrictive setting a publisher applies, so blocking AI training can also block search indexing and the traffic that depends on it.

Who is driving the growth in bot traffic?

In other words, agents rather than scrapers. Research from HUMAN Security shows action-taking agent traffic grew 7,851 per cent year over year against 597 per cent for scrapers, while Cloudflare data shows automated systems reached 57.5 per cent of webpage requests in June 2026, more than a year ahead of its own chief executive's forecast.

What are the risks for smaller publishers?

The answer is asymmetry. Evidence suggests access friction weighs most heavily on smaller and newer AI developers who cannot negotiate licensing deals, while incumbents already hold crawled data. Simply put, publishers gain real control for the first time, and the largest players have already taken what they needed.


Sources:

Cloudflare — Your site, your rules: new AI traffic options · Cloudflare Docs — Block AI bots and September 2026 defaults · Cloudflare — Content Independence Day · Cloudflare — Introducing pay per crawl · Fortune — AI agents are eating the web · HUMAN Security — 2026 State of AI Traffic · The Decoder — Cloudflare replaces its blanket block · Related on this site: Google AI Copyright Lawsuit Over Gemini and Books · News Groups Ask Court to Sanction OpenAI Over AI Data

Stay in the Conversation

Subscribe for writings on Emergent Intelligence, digital personhood, and the future we are building together.

Share this essay

AI News Answers Fail on Retrieval, Not Reasoning

2h ago·6 min read

Also worth your time

Business

Meta Muse Spark 1.1 Opens an AI Price War

2h ago·6 min read
Technology

An AI Agent Breached Hugging Face and Guardrails Blocked the Defence

An autonomous AI agent breached Hugging Face in July 2026. Commercial model guardrails then blocked the forensic analysis, so the defence ran on open weights.

7 min read · Jul 24, 2026