Industry Insight

Industry Insight

All posts

Search, Agent, Training: The Web's New Bot Rules

Cloudflare now sorts bots by purpose, not behavior: Search, Agent or Training. The September 15 default is narrower than the panic, and one box is missing.

Web Scraping Is Consolidating. Here's What Changes.

Oxylabs took $130M at a $3.6B valuation. Bright Data crossed $300M ARR. The mid-market web scraping API is disappearing fast. Ask these five questions.

Why Headless Browsers Leak More in 2026

Detectors widened the gap between headless and normal browser sessions in 2026. If you run Puppeteer or Playwright in production, plan for the detection tax.

Cloudflare Precursor: Session Is the New Fingerprint

Cloudflare shipped Precursor on July 13. Session-scoped behavioral detection means refreshing the page no longer resets your bot score. Why it matters.

Pay-Per-Crawl Is Splitting the Web in Two

Cloudflare's pay-per-crawl marketplace and HTTP 402 split the web into licensed and open data. Here's what shifts for teams collecting web data in 2026.

When LLM Extraction Stops Paying for Itself

Firecrawl charges 5x to LLM-extract a page versus scrape it. At 100K pages a day the math collapses. When LLM extraction earns its bill, and when it doesn't.

Why Proxy Pool Size Stopped Mattering in 2026

Vendors advertise 400 million residential IPs. But in 2026, IP reputation collapsed as a defense, and proxy pool size stopped predicting real success.

The EU AI Act Ends the Free-For-All in Training Data

AI training data collection just went from technical problem to compliance problem. The EU AI Act and rising vendor scrutiny reshape the rules through 2027.

Bot Detection Went Behavioral. Most Scrapers Didn't.

Bot detection shifted from IP blocking to TLS fingerprints, browser signals, and behavioral analysis. Most scraping setups are fighting the wrong battle.