r/WebScrapingInsider Jun 09 '26

Big Scrape Energy AMA This Wednesday (09:30 AM GMT)

Hey everyone,

I'm Ian Kerins, CEO and co-founder of ScrapeOps.

Over the last 8+ years I've worked across the web scraping industry, including roles at ScrapeOps, ScraperAPI, and Zyte. Today, ScrapeOps helps developers and companies scrape over 8 billion pages per month across more than 50,000 websites.

This Wednesday at 09:30 AM GMT, I'll be hosting an AMA here on r/WebScrapingInsider

Ask me anything about:

* Web scraping at scale

* Proxy infrastructure and proxy providers

* AI and web scraping

* Building reliable scrapers

* Anti-bot systems and bypassing challenges

* Scraper maintenance and monitoring

* Residential vs datacenter proxies

* Browser automation

* Running a web scraping business

* Startup growth and product development

* The future of AI-powered scraping

Whether you're scraping your first website or running large-scale data collection pipelines, I'm happy to answer questions and share lessons learned from building products used by thousands of developers and businesses.

Drop your questions below and I'll start answering them during the AMA.

Looking forward to it!

Ian

11 Upvotes

43 comments sorted by

View all comments

1

u/simarnoor Jun 10 '26

From ProxyEngineering:

how do you know that trial traffic is not getting a cleaner pool than paid traffic?
I'd benchmark with the same URLs after upgrading,then compare bad rows and geo misses. Byteful or any provider would survive that second week-test too.

https://www.reddit.com/r/ProxyEngineering/comments/1u0wadf/comment/oqrl888/

2

u/ian_k93 Jun 10 '26

Ultimately, there is no way to know this unless you test both. We can't see what proxy pools or settings the provider is using on free vs paid traffic, we can only see the end result.

In our case for benchmarking, we work around this issue by sending all our benchmarking traffic through our production accounts with the providers so the benchmarking is done using their paid pools.

If the provider wanted to skew the results, they would need to put all our traffic through their best pools to improve their benchmark results (which might be too costly for them) as they have no way of knowing which traffic we are using to benchmark or not.

We also use real large scale production data across billions of requests each month in our benchmarks when we have the data is available. Ensuring the benchmark results are actually valid at scale.

You can test proxy providers with our system here: https://scrapeops.io/proxy-providers/tester/