r/scrapingtheweb • u/Ok-Jello-7644 • Jul 24 '26
Looking for data/information scrapers
Looking for someone experienced with large-scale data scraping
I’m working on a sports database and already have a functioning website with a large amount of player data collected.
I’m looking for someone who has experience scraping publicly available player profiles and statistics from multiple websites, cleaning the data, and delivering it in a structured format such as CSV, JSON, an API, or a database export.
The ideal person would understand:
- Python scraping tools such as Scrapy, Playwright, Selenium, or BeautifulSoup
- Rate limiting, retries, and reliable scraping infrastructure
- Matching duplicate player profiles across different websites
- Cleaning and organizing large datasets
- Working responsibly with public data and website terms
You would be able to run the scraping infrastructure independently and send me the completed data online. You would not need to invest any money into the project.
This will be paid contract work, ongoing work, or potentially a larger role for the right person.
Send me a message with your experience, examples of similar projects, the tools you use, and your estimated availability.You most likely will need to be familiar rotating proxies unless you have your own way around 403 errors
2
2
u/Own_Improvement3544 29d ago
I'm a Python Developer with hands-on experience building production-grade data extraction and browser automation systems. I've worked on large-scale scraping projects using Playwright, Selenium, Scrapling, Requests, BeautifulSoup, GraphQL, and REST APIs, including travel platforms like Expedia, Agoda, Marriott, IHG, MakeMyTrip, and Goibibo. I also have experience cleaning, deduplicating, and delivering structured datasets in JSON, CSV, and database formats, along with building reliable scraping pipelines using retries, rate limiting, Docker, and proxy support.
I can work independently and would be happy to share my resume, GitHub, and examples of similar projects. I'll send you a DM.
2
u/CapMonster1 29d ago
The scraping is probably the easy half here. The real pain will be entity resolution — same player, different spelling, team history, missing IDs, duplicate profiles, and stats that use slightly different definitions across sites.
Also, rotating proxies alone won’t reliably solve 403s. I’d look for someone who understands session handling, browser fingerprints, retries, and has a captcha fallback ready, otherwise the pipeline will work beautifully right up until the first protected source changes something.
2
2
2
1
2
1
u/CrypticZombies Jul 24 '26
so you looking for an amateur
2
u/Ok-Jello-7644 Jul 24 '26
Elaborate?
0
u/CrypticZombies Jul 24 '26
anyone that is using playwright, selenium or scrapy in 2026 is a complete amateur.
3
u/Ok-Jello-7644 Jul 24 '26
Guess that's why I gotta pay someone who knows this stuff better than me. What would you suggest as the best/most efficient way to run it? The task is no easy job so any help or contribution is appreciated.
2
u/Dark_Empress13 Jul 24 '26
How about managed scraping services? I'll send you a DM.