← All posts

Dealer Inventory Scraping Past the Listing Sites

Dealer inventory scraping breaks on the rooftop sites, not the aggregators. Client-side prices, a handful of platform templates, and blocks that look like data.

Everybody builds the aggregator scraper first. Cars.com, CarGurus, AutoTrader: three sites, one parser each, and by the end of the week you've got a feed that looks like the used-car market.

Then someone asks where the rest of it is.

The Challenge

NADA counts 16,972 franchised light-vehicle dealers in the United States as of its mid-2025 report, and that's before the independent lots, which nobody counts consistently. Secondary write-ups of the same NADA figure land anywhere between 15,720 and 16,990, which tells you something about how well this market is measured.

Each of those rooftops runs its own website. The inventory sitting on it is that dealer's stock, priced today, days before any of it settles onto a third-party listing site. If you're pricing used vehicles, forecasting residuals, or selling a competitive tool back to dealers, the rooftop site is the data you want. The aggregators are a lagging, filtered copy of it.

So teams go after the rooftop sites, and they find out three things in order.

They aren't unique. Almost all of them run on a small set of dealer website platforms: Dealer.com, DealerOn, Dealer Inspire, CDK, Reynolds, Sincro, Lotlinx (the exact roster depends on who's counting). Anyone selling this data builds one parser per platform, not one per dealership. Apify's dealer website inventory scraper says so on the tin: detect the platform first, then extract. A handful of templates covering tens of thousands of rooftops. That's the good news in this story.

The price usually isn't in the HTML you fetched. These platforms render the pricing block client-side, and payment estimates and incentives often arrive in a second call after that. A plain HTTP fetch gets you year, make, model, mileage and VIN. The price field comes back empty.

And then the part that quietly ruins datasets: on a dealer site, a failure looks exactly like a fact. A bot-detection service answers a suspicious request with HTTP 200 and an interstitial. A pricing block that never rendered leaves "Call for Price" in the DOM, which is also something dealers write there on purpose. Both rows land in your warehouse looking equally clean.

We covered that pattern in a different vertical, where a block looks like a data point. Automotive is the harder version, because "no price" is a legitimate business state rather than an obvious anomaly.

What Changes When You Split by Cost

The pipelines that hold up don't organize themselves by site. They organize by what a request costs.

The expensive request is the first one against a rooftop: the one that has to run a browser, clear whatever the dealer's platform put in front of it, and come back with a session. Everything after that is a cheap HTTP call reusing what the first one earned.

import requests

FOURA = "https://api.foura.ai/api"
AUTH = {"Authorization": "Bearer pk_live_..."}

# First page of a rooftop: render it, clear the defense, keep the session.
first = requests.post(f"{FOURA}/browser", headers=AUTH, json={
    "url": "https://example-motors.com/used-inventory/index.htm",
    "unblocker": True,
    "timeout_ms": 45000,
}).json()

listings = first["body"]
jar      = "; ".join(f'{c["name"]}={c["value"]}' for c in first["cookies"])
agent    = first["userAgent"]
exit_id  = first["proxy"]     # opaque proxy ID, send it back to stay on the same exit

Then walk the rest of that dealer's inventory without paying for a browser again:

page = requests.post(f"{FOURA}/single", headers=AUTH, json={
    "method": "GET",
    "url": "https://example-motors.com/used-inventory/index.htm?start=20",
    "proxy": exit_id,
    "headers": [["Cookie", jar], ["User-Agent", agent]],
    "validate": {
        "status": {"accept": [200]},
        "data": {
            "accept": ["vehicle-card"],
            "fail": ["Just a moment", "Access Denied"]
        }
    }
}).json()

Two details there carry more weight than they look like they do.

The session travels as one unit. A clearance cookie is bound to the exit that earned it and to the User-Agent it was earned under. Replay it from somewhere else, or under a different User-Agent, and the site starts you over at the challenge. That's why the jar, the agent string and the proxy ID move together. It's the detail we watch people get wrong most often: they keep the cookie, drop the exit, and then wonder why the cheap path stopped being cheap.

The validate block is what kills the silent-failure problem. It's the request stating what a real page looks like: accept a marker that only exists once the listing grid rendered, fail on the interstitial strings. A response that misses those rules isn't a row with a null price. It's a failure, classified as one, and not counted as a success. In automotive, write the positive marker as well as the negative list, because "Call for Price" is genuinely ambiguous and "the vehicle card never rendered" never is.

When you don't yet know which path a given platform needs, Auto will find out in a single call and hand back the session that worked. Treat that as reconnaissance, not as the production route. Once you know that one platform needs the browser and its neighbours don't, pin each to the direct engine and stop paying an orchestrator to rediscover the same answer every night.

Results

Run the arithmetic on a mid-size job (illustrative scenario based on industry benchmarks, not a specific customer): 4,000 rooftops, roughly 180 used vehicles each, refreshed nightly.

  • 4,000 rendered pages instead of 720,000. One render per rooftop opens the session; the other 716,000 pages go through the cheap path on the same session. That ratio, not the parser, is what decides whether nightly coverage is affordable.
  • Two kinds of missing price, in two different tables. With validate rules in place, "the dealer publishes no price" and "we never got the page" stop sharing a row shape. Your model gets to see only the first one.
  • One parser per platform, not per dealership. Detect the platform from the response, hand the HTML to the parser that owns it. A new rooftop on a covered platform costs nothing to onboard.
  • A dealer that switches platforms fails loudly. Platform detection misses, the row never gets written, and someone gets a ticket instead of six weeks of quietly wrong prices.

Where this gets less tidy: the big dealer groups increasingly run bespoke sites off-platform, and those still need hand-built parsers with all the maintenance that implies. Rooftop freshness isn't uniform either. Some platforms cache their inventory pages hard, so "today's price" can be a day old no matter how often you collect it. If your model treats every rooftop timestamp as equally live, it's wrong in a way no amount of collection infrastructure will fix.

Key Takeaway

The hard part of automotive data was never the three aggregators everyone benchmarks against. It's the seventeen thousand small sites that were never worth a dedicated scraper on their own and are worth a great deal together.

That shape shows up well beyond cars. Pharmacies, equipment dealers, regional grocers, franchised anything: the long tail only looks expensive while you're treating each member of it as a unique site. It usually isn't one. Get the first request right, make the session reusable, and what's left stops being a collection problem and turns into a parsing problem, which is the cheaper of the two by a wide margin.