Why publish the method before the results
The strongest complaint against comparison sites in this space is a fair one: the vendors being compared often outrank independent reviewers in search results for their own "best scraping API" round-ups, and a lot of third-party "comparisons" recycle the vendor's own marketing claims without testing anything. The only way an independent site earns credibility here is by running its own tests, publishing the method openly, and being willing to rank a paying affiliate partner below a competitor when the data says so. We're publishing the plan first, in public, before a single benchmark number appears on this site.
What we plan to measure
- Success rate. Percentage of requests against a fixed panel of target pages that return a usable, correctly-rendered response (not a block page, CAPTCHA wall, or empty/garbled response).
- Response time. Median and 95th-percentile latency per successful request, measured end-to-end from our test harness.
- Cost per 1,000 successful pages. Computed from each vendor's own published pricing at like-for-like settings (same feature flags — e.g., JavaScript rendering on or off — enabled across every vendor tested in that run), not the cheapest possible configuration for one vendor against the most expensive for another.
- Anti-bot / CAPTCHA bypass rate on a subset of known bot-protected targets, tracked separately from the general success rate since it behaves very differently by site.
How we plan to keep it fair
- Fixed, disclosed target list. A published set of test URLs spanning static pages, JavaScript-heavy pages, and pages with active anti-bot protection — not a target list picked after seeing which vendor does better on it.
- Same settings across vendors in a given run. Every vendor tested with the same feature flags enabled (e.g., JS rendering on, or off) in a given comparison, rather than each vendor's best-case configuration against another's default.
- Recurring re-runs, not a one-time snapshot. Scraping success rates and anti-bot defenses change on both sides — we intend to re-run the panel on a regular cadence and publish the date of each run, rather than leaving a single result live indefinitely.
- Affiliate relationships disclosed, not hidden from the ranking. Some vendors tested may be ones we earn an affiliate commission from (see the Affiliate Disclosure) — the fixed target list and published parameters exist specifically so that relationship can't quietly tilt which vendor "wins" a given test.
- Only publish what we actually ran. No estimated, interpolated, or "typical" numbers standing in for a real test. If a vendor isn't tested yet, its page says so — as this page does now.
Legal and platform-terms scope
Test targets will be limited to pages and services where automated access is either explicitly permitted or falls within general, publicly-documented technique — no scraping of the org's own confidential client work, and no targets chosen specifically to demonstrate bypassing a site's terms of service. This methodology draws only on public, general-technique knowledge, consistent with the org's standing clean-room policy for all published work.
Status
This methodology is published; the benchmark harness has not been built and no test run has occurred as of the date on this page. When the first results are published, this page will be updated with the run date, and every comparison page on this site that currently states "not independently benchmarked" will be updated to link to the actual data instead.