Search, Agent, Training: The Web's New Bot Rules
Cloudflare now sorts bots by purpose, not behavior: Search, Agent or Training. The September 15 default is narrower than the panic, and one box is missing.
Cloudflare now sorts bots by purpose, not behavior: Search, Agent or Training. The September 15 default is narrower than the panic, and one box is missing.
Oxylabs took $130M at a $3.6B valuation. Bright Data crossed $300M ARR. The mid-market web scraping API is disappearing fast. Ask these five questions.
Detectors widened the gap between headless and normal browser sessions in 2026. If you run Puppeteer or Playwright in production, plan for the detection tax.
Cloudflare shipped Precursor on July 13. Session-scoped behavioral detection means refreshing the page no longer resets your bot score. Why it matters.
Cloudflare's pay-per-crawl marketplace and HTTP 402 split the web into licensed and open data. Here's what shifts for teams collecting web data in 2026.
Firecrawl charges 5x to LLM-extract a page versus scrape it. At 100K pages a day the math collapses. When LLM extraction earns its bill, and when it doesn't.
Vendors advertise 400 million residential IPs. But in 2026, IP reputation collapsed as a defense, and proxy pool size stopped predicting real success.
Your User-Agent header doesn't matter anymore. Connection-level fingerprints classify bots at 98.6% accuracy before headers are read. Here's what shifted in 2026.
AI training data collection just went from technical problem to compliance problem. The EU AI Act and rising vendor scrutiny reshape the rules through 2027.
Bot detection shifted from IP blocking to TLS fingerprints, browser signals, and behavioral analysis. Most scraping setups are fighting the wrong battle.