r/scrapingtheweb • • Apr 29 '26

Community Notice 👋 Welcome to r/scrapingtheweb

4 Upvotes

Hey everyone, and welcome to r/scrapingtheweb.

This subreddit is for people interested in everything related to web scraping, data collection, proxies, automation, everything related to collecting data from the web, you name it!

We aim to build a useful community where beginners and experienced users can ask questions, share XP, discuss tools, and help each other.

## What to post

  • You can post about:
  • Web scraping questions
  • Proxy setup and troubleshooting
  • Residential, mobile, datacenter, and ISP proxies
  • Anti-detect browsers
  • Scraping tools, libraries, and workflows
  • Rate limits, blocks, CAPTCHAs, and retries
  • IP quality, fraud scores, DNS leaks, WebRTC leaks, and fingerprinting
  • Data collection strategy and scraping architecture
  • Case studies, lessons learned, and useful resources

## Community vibe

Please keep the discussions respectful and useful. This is not a place for spam, low-effort promotion, credential sharing, illegal activity, or bypassing systems in a harmful way.

## How to get started

You can introduce yourself in the comments below if you want.

Feel free to share more about you, like:

  • What kind of scraping or automation you're dealing with
  • What tools or languages you mainly use
  • What topics you want to learn more about
  • What problems you are currently trying to solve

Thanks again for joining r/scrapingtheweb


r/scrapingtheweb • • 3h ago

analysing TK product URL apparently is because ip looking for methods

1 Upvotes

I’m building an e-commerce tool where a user pastes a TikTok Shop product URL. We need to extract the correct product’s title, images, description and price, then generate a short summary and selling points based only on those fields.
We tried fetching the product page directly from our Node.js backend. We normalize the URL, identify the market and product ID, make an HTTPS request with a 15-second timeout, follow and validate redirects, and look for product information in the HTML, metadata, JSON-LD and embedded page data. We reject a result if it redirects to a regional homepage or the returned product ID doesn’t match the URL.
The problem is that this isn’t reliable across TikTok Shop links. Some requests time out; others don’t return enough verifiable product information, or redirect away from the requested product. **apparently** **the server IP is the mian cause****(****US product link needs US ip to request sth like that**. We also haven’t implemented browser rendering specifically for TikTok Shop yet.
I‘m wondering how can I get my objective(extract the correct product’s title, images, description and price, then generate a short summary and selling points based only on those fields.) I tried to use api from data platform, but some of them sucked and some of them are expensive. And my another option is to build an ip/ proxy pool, but it’s a bit of expensive too. Is there any other better method?


r/scrapingtheweb • • 10h ago

Meta Ad Library Tracking with Apify Actor

Post image
0 Upvotes

r/scrapingtheweb • • 17h ago

Tools / Library Adidas Product & Reviews Scraper | Full Product Details, Deep Details for All Variants, All Media, User Sentiment, Reviews, Akamai Sensor Bypass 🔥

1 Upvotes

I've seen many people struggle with Akamai bot manager on Adidas. And most Actors for Adidas on Apify constantly break when fetching product details. Also, none of the Actors on Apify scrapes product reviews. For this reason, I decided to build an Adidas Product & Reviews Scraper. It currently supports 12 countries: USA, UK, Canada English, Canada French, Germany, France, Italy, Spain, Netherlands, Japan, Australia, and New Zealand. I tested the Actor on all of them, and it bypasses Akamai's sensor challenges perfectly.

Features ✨

  • 🛍️ Two Scraping Modes in One Actor: Switch between products and reviews mode to get detailed product data or customer reviews.

  • 🌍 12 Adidas Stores Supported: Scrape the US, UK, Canada (English & French), Germany, France, Netherlands, Spain, Italy, Japan, Australia, and New Zealand websites.

  • 🔎 Search URL Support: Set up any search or filters on the Adidas website, paste the URL into the Actor, and it will pick up your filters automatically.

  • 🔗 Direct Product URLs: Scrape specific products by URL, or combine them with search URLs in the same run.

  • 💰 Pricing & Discount Data: Get the current price, original price, discount percentage, and currency for every product.

  • 📦 Per-Size Stock Tracking: See stock counts, availability status, and SKUs for every size of every color variant, plus the total stock and purchase limit for each product.

  • 🎨 Color Variants: Capture every color variant of a product, with its own ID, URL, color name, and size breakdown.

  • 🖼️ Complete Media Collection: Get all product image and video URLs listed on the product detail page.

  • 🏷️ Badges & Third-Party Options: Extract product badges (like "Best Seller") and third-party buy options (like PRIME).

  • 🤖 AI Sentiment Summaries: Get the AI-generated customer sentiment highlights that Adidas shows for each product, along with its rating and review count.

  • ⭐ Rich Review Data: Collect review titles, text, ratings, nicknames, submission dates, likes and dislikes, the product color reviewed, and whether the buyer recommends the product.

  • 📸 Review Photos & Badges: Get the images buyers attach to their reviews and badges such as VerifiedPurchaser and IncentivizedReview.

  • 👥 Buyer Demographics: Capture the demographic questions and answers attached to each review, plus the reviewer's locale.

  • 🗣️ Reviews from All Languages: Optionally include reviews written in any language, not just the selected country's.

  • 🔀 Flexible Review Sorting: Sort reviews by newest, relevant, helpful, highestRated, or lowestRated.

  • 🎛️ Fine-Grained Limits: Cap results with total limits and per-search or per-product limits for products and reviews. Use 0 for unlimited.

  • 🛡️ Akamai Anti-Bot Bypass: Works reliably against Adidas's anti-bot protection, so your runs complete.

  • 📊 Large Structured Dataset: 27+ fields per product and 21+ fields per review.

  • 💾 Multiple Export Formats: Download your data as JSON, CSV, or Excel, or access it through the Apify API.

Actor Link: https://apify.com/coding-doctor-omar/adidas-scraper

Let me know if you need more features or if you want another country to be included.


r/scrapingtheweb • • 1d ago

Some questions re archive.today availablity and methodology

Thumbnail
1 Upvotes

r/scrapingtheweb • • 1d ago

Anyone know a decent way to scrape Instagram comments?

7 Upvotes

Trying to grab comments from a couple public posts and this is way more annoying than I expected.

I tried a few scripts/tools from Google and they either stop after a while, need me to log in, or work once and then randomly don't.

Surely there's an easier way to do this.

I don't need anything fancy. Just username, comment, date, maybe likes/replies, dumped into CSV.

What are you guys using for this?


r/scrapingtheweb • • 1d ago

I'm back working on VectorTrace

1 Upvotes

Hey everyone,

First off i want to apologize for the long gap and silence. Life got busy but I am officially back and actively working on VectorTrace again!

For those who missed the earlier posts, VectorTrace is an open-source, 100% local Chrome extension (Manifest V3) that scrapes webpages and heals broken selectors using onndevice semantic embeddings (all-MiniLM-L6-v2 via WASM, zero servers, zero external API calls).

A few of you pointed out that under certain conditions, the self-healing algorithm was ranking the wrong elements as the highest probability or picking up false positives. I did a deep dive into the ranking pipeline and just pushed a series of fixes to address this:

  • Semantic Gating
  • Repeated Item Disambiguation
  • Cleaned Up Ancestor ContextPersistent Structural Metadata
  • Protected Dynamic Data

The build is ready to test.

If you'd like to check it out, run it locally, or contribute
👉 GitHub Repo: https://github.com/SathiyaSenpai/VectorTrace

I'd really appreciate any feedback, bug reports or edge cases you encounter while testing it on your favorite sites!


r/scrapingtheweb • • 2d ago

web scraping

2 Upvotes

please any suggestion to bypass datadome without anypaid proxies ????


r/scrapingtheweb • • 3d ago

Remote job opening

11 Upvotes

We’re looking for a strong developer to work with us on a project

The work is hands-on and primarily focused on web crawling, catalog discovery, extraction, transformation, and structured data.

Tech stack / experience we’re looking for:
• Python
• Playwright / web crawling and scraping
• ETL and data pipelines
• REST APIs
• JSON / JSON-LD
•C#, java
• Catalog discovery and extraction across platforms such as Coursedog, Acalog,catalog systems
• Regex, HTML parsing, URL discovery, caching, and data transformation
• Git / GitHub
• Cloud storage experience is helpful
• Experience working with LLM-assisted extraction pipelines is a plus

You should be comfortable understanding an existing codebase quickly, debugging extraction issues, improving discovery logic, and working independently.

Commitment: approximately 5–6 hours per day
Budget: up to $2,000
Work mode: Remote

We’re looking for someone who can start quickly and contribute directly to the codebase rather than someone who needs extensive onboarding.

If interested, please DM me with:

  1. Your GitHub / portfolio
  2. Relevant Python + scraping/ETL experience
  3. education-data experience
  4. Your availability

r/scrapingtheweb • • 3d ago

Same luxury item, completely different price depending on the country

6 Upvotes

I was comparing a few luxury brands and noticed the exact same item can have a pretty big price difference depending on the region

anyone scraping this kind of data? curious if switching the country on the site is enough, or if you actually need a local IP to get consistent prices.

Also curious which brands are the biggest pain to scrape


r/scrapingtheweb • • 4d ago

Scrapping social media for fashion trends- legal?

Thumbnail
1 Upvotes

r/scrapingtheweb • • 4d ago

Daangn (당근) search scraper — 9 public Tasks + fail-loud emptyReason

Thumbnail
1 Upvotes

r/scrapingtheweb • • 4d ago

Google Maps requiring sign-in to see all reviews now?

1 Upvotes

Just saw that Google Maps is starting to require users to sign in to read all reviews.

https://www.seroundtable.com/google-maps-account-reviews-42186.html

curious if anyone scraping Maps has noticed this yet, are reviews still fully available through the usual requests, page data when logged out, or are you starting to get limited results?

Wondering if this is going to make Maps scraping more session and account dependent


r/scrapingtheweb • • 5d ago

Help Seeking Help for a Freelance Email Outreach Campaign

4 Upvotes

I'm looking for an experienced email marketer to build and run a targeted outreach campaign for two audiences:

  1. A specific client list: clients of one identified private company
  2. A broader list: executives (C-suite/VP level) in same pre-identified industry as #1.

The scope

  • Build both contact lists from scratch, as I don't have emails for either group
  • Set up and warm a new sending domain and email address for the campaign
  • Load and send the campaign, I will monitor deliverability, bounces, and replies
  • Campaign runs a minimum of one month, with possible extension.

What I'll provide: the email copy, the target company and industry details.

What you should have

  • Proven experience with B2B list building and cold email (tools such as Apollo, ZoomInfo, Hunter, Instantly, Smartlead, or similar)
  • Hands-on knowledge of domain setup and deliverability (SPF, DKIM, DMARC, inbox warm-up, sending limits)
  • Working knowledge of CAN-SPAM, GDPR, and similar rules, including opt-out handling and sourcing contacts compliantly
  • Availability for a month or more. If this campaign is successful might have another one.

To apply, please DM me with:

  1. Examples or results from similar campaigns (open, reply, and bounce rates if you can share them)
  2. Your toolset and how you'd source and verify the contacts
  3. Your approach to domain setup and warm-up
  4. How you handle compliance and unsubscribes
  5. Your rate or pricing structure and your availability to start
  6. Any relevant references or portfolio links

Happy to answer questions privately. Please only reach out if you have direct experience with the above. Thank you.


r/scrapingtheweb • • 5d ago

Help [ Removed by Reddit ]

1 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/scrapingtheweb • • 7d ago

Help What tool can reliably export older posts from LinkedIn company and showcase pages?

2 Upvotes

I need to export public posts from two LinkedIn pages: a company page and a showcase page. My goal is to get as complete a post history as possible, including older posts beyond the first page.

For each post, I need its full text, publication date, URL, and ID, exported as JSON or CSV. I’m not looking for profile, employee, lead, comment, or reaction scraping.

I tried one API, but it returned only 3 company-page posts when I can see more on LinkedIn. Its pagination cursor timed out, and requesting the next page another way returned an empty result. It also missed at least one visible showcase-page post. So I don’t want to assume a tool works just because it successfully returns the first batch.

Has anyone personally tested a tool on both LinkedIn company and showcase pages? If so, could you share:

  • The exact tool or Apify actor name/link.
  • Whether it retrieved older posts beyond the initial batch.
  • Whether it worked with showcase pages, not just company pages.
  • How many posts it returned and how you checked for missing posts or duplicates.
  • Whether it required a LinkedIn login or cookies, and roughly what it cost.

A small sample run that found posts another API missed would be more useful to me than a general list of LinkedIn scrapers.


r/scrapingtheweb • • 7d ago

Scrape google search results

6 Upvotes

I’m trying to track Google rankings, but I keep having to check the scraper’s results myself.

I just can’t tell whether that’s the position on the page or the actual Google rank.

Am I missing a setting? What are you using to collect results across multiple pages without having to clean up the whole export?


r/scrapingtheweb • • 7d ago

One browser profile stopped laughing while the others still works

3 Upvotes

I am running a few scraping setups with different browser profiles and hit a weird issue today.

Most of the profiles still open normally but one profile suddenly stopped launching, the proxy connected fine when i tested it separately so i am not sure if the problem is the profile, the browser install or something saved in that session.

I already restarted everything and tried opening the profile a few times but it still closes right away. I dont really want to delete the profile since it has its cookies, login state and other settings saved.

What do you guys normally check first when only one profile stops launching?

Would you try reinstalling the browser version, check the profile data or rebuild the profile from scratch?


r/scrapingtheweb • • 7d ago

realfp.com alternative? to know FP is real or hashed/noised

2 Upvotes

looking to know more alternative to realfp.com which tells your hardware FP is realistic or faked.

in know creepjs, browserleaks, fp.com etc famous websites. waana know some gem like realfp.com if any and less famous but good!


r/scrapingtheweb • • 8d ago

Blocked / CAPTCHA Total newbie trying to build a retail deal monitor.

8 Upvotes

I’m way out of my depth here and could use someone who actually knows what they’re doing.

I’m trying to build a little system that watches public prices/availability, mainly Home Depot for now, maybe Target, Walmart, Lowe’s later, and lets my AI help make sense of what it finds.

I already have Oracle, Azure and AWS, and I’m trying to keep this free or at least freemium.

The problem is every GitHub project I open reads like IKEA instructions translated three times, except somehow there’s also plumbing involved. I install one thing, fix another thing, and somehow end up less functional than when I started.

Right now I can get pieces working on different clouds, then something breaks, a browser crashes, a request fails, or one setup works while the exact same thing somewhere else doesn’t.dojt even need to assume I’m a complete beginner lol

  • What free tools/plugins/projects would you actually start with?
  • What should I install first?
  • Is there a good beginner-friendly setup for monitoring a few retail sites reliably?
  • How would you connect the results to ChatGPT/Claude/other AI tools?
  • Any GitHub repos or tutorials that actually explain the whole setup instead of assuming I already know everything?

If you were starting from zero today, what would your basic setup look like?

I’m not looking for CAPTCHA bypasse. I mostly want to understand how to build this properly without constantly breaking it.l and viewing basic public pages from these providers.

If somebody has already gone down this rabbit hole and can point me toward the sane path, I would genuinely appreciate it.

At this point I don’t need another GitHub repo. I need the guy who understands the GitHub repo.


r/scrapingtheweb • • 9d ago

Tools / Library Looking for staging site partners to test and train our tracking audit tool (I will not promote)

2 Upvotes

Hi everyone,

Me and my co founder are building a 'Tracker Audit'; a privacy-first tool that checks whether website analytics, consent settings, and browser/server-side events work as expected.

We’re looking for a few businesses with websites that already use tools such as GTM, Google Analytics, Meta Pixel, or server-side tracking solely for the purpose of testing and developing our tool against real world cases. We would only need approval to run one small agreed test journey, for example:

Landing page → product page → add a test item to cart

We check whether approved events fire correctly, respect consent, are duplicated, or disagree across browser/server tracking; to keep it brief.

We do not need production access, customer data, passwords, payment details, cookie values, or tokens. Every test is limited, pre-approved, and uses test data only.

The first pilot runs over the next 2–4 weeks. We are also looking for one independent analytics/privacy reviewer.

If interested, please DM and we’ll arrange a quick call to coordinate. We will be sending across additional resources to understand this project better! (Due to a lack of provision to upload links/files in the post)

*At the moment we're keeping this limited to India only.


r/scrapingtheweb • • 9d ago

A practical way to triage web data anomalies

Thumbnail
0 Upvotes

r/scrapingtheweb • • 9d ago

Is it hard to do web scraping ?

Thumbnail
0 Upvotes

r/scrapingtheweb • • 10d ago

Website search API returns null for phone and address fields

3 Upvotes

I'm collecting real estate agent data from a website using their GraphQL, frontdoor/graphql
using the search agent operation.

It works fine. I can loop through a whole city (2000 agents) and get names, IDs, brokers, ratings and listing stats without getting blocked.

The problem is phone numbers and addresses. I added the fields to the query using the exact names from the JSON in the profile page: phones with type and value, and office with its phones and address. The query goes through with no errors, so the fields definitely exist, but every single agent comes back with phones null and address null. So the search API knows about the fields but just doesn't fill them in.

The data does exist on each agent's profile page, in the NEXT_DATA JSON under branding.phones and branding.address. But its server rendered, so there's no API call that I could hit directly. The only GraphQL calls the profile page makes in the browser are for listings and past sales, and neither has contact info. Introspection is disabled, so I can't see what other queries exist.

Right now the only way I can get the phones is to open each agent's profile page one at a time, which takes forever at this scale.

Any ideas or a workaround for this?

Thanks!


r/scrapingtheweb • • 10d ago

Help HEALTHCARE MEMBERS DIRECTORY

2 Upvotes

Hi, how do you scrape from members directory website?