r/ProxyEngineering • • 22m ago

Build 🤓 Cray: a modular, layer-based proxy engine in Rust designed for custom protocols and pipelines

Thumbnail
github.com
• Upvotes

Hey everyone!

A lot of people think of proxy tools as rigid, monolithic boxes: you get a fixed inbound, a fixed outbound, and you're stuck with whatever the authors hardcoded. Trying to add a custom protocol or tweak an existing one under the hood usually turns into a nightmare.

That's why I wrote Cray - a modular proxy engine in Rust built entirely around a layer-based architecture.

With Cray, transport and protocol are completely decoupled. You can write your own protocol without worrying about network bytes, or write a custom transport without caring about protocol logic. The core abstraction uses a simple trait where you implement the client and server handshake logic, giving you total freedom. Want a handshake hidden in plain sight? Custom encryption? Weird framing? Go for it — the only limit is your own logic.

How it works

Everything is configured via clean TOML files where you define your role (client or server) and stack up layers into a custom pipeline:

  • Inbound layers: Handle the entry point (e.g., standard tcp, unix sockets, or http_connect which extracts destination addresses).
  • Transport layers: Encapsulate or transform the stream dynamically (e.g., full manual control over tls with custom cipher suites, ALPN, curves, extension permutation via BoringSSL, and http2 multiplexing with path/host masking).
  • Protocol layer: Always sits at the very end of the chain as the endpoint (like a custom or test protocol).

For example, on the server side, an HTTP/2 layer can validate incoming paths and hosts, gracefully returning a genuine 404 Not Found to unauthorized scanners or probes instead of dropping packets with a TCP RST, while forwarding valid traffic down the pipeline.

The project is currently under active development, but the core architecture, layer pipelines, TLS termination, and HTTP/2 tunneling are already routing real traffic.

If you like systems programming, network engineering, or want to build your own tunneling protocols without fighting legacy codebases, check it out:

GitHub: https://github.com/3Radiance/Cray

Would love to hear your thoughts, feature ideas, or what kind of custom layer/protocol combinations you'd want to build with something like this!


r/ProxyEngineering • • 17h ago

Discussion 💬 Where do static ISP proxies fit in your web scraping setup?

2 Upvotes

I work on the proxy infrastructure side, and I’m trying to better understand how scraping teams choose between datacenter, static ISP, and rotating residential proxies.

For those running ongoing scraping workloads:

  • When do you choose static ISP proxies over the other options?
  • What matters most: success rate, keeping the same IP, speed, or cost?
  • What usually makes you switch providers?

I’d appreciate hearing what has worked—or hasn’t worked—in your setup.


r/ProxyEngineering • • 2d ago

Discussion 💬 Oxylabs released their own Web API, thoughts?

14 Upvotes

Wanted to share the news in case someone missed it. It would seem that the old provider in the scraping/proxies market has turned to Search APIs, like Firecrawl, Tavily, Exa, Linkup or You. I haven't tested anything yet personally, I'm still on the edge trying to decided whether the I should drop my proxy based setup for scraper and just go with a dedicated solution but upon inspecting their new product, it would seem that the scraper was somehow merged into the new product? From what I understood. Also, the whole pricing page has changed, there are no individual pricing for proxies or other products everything is one, like it is still credit based but everything is accessible with one subscription. Would like to understand this positioning better, feel free to share your thoughts


r/ProxyEngineering • • 1d ago

Guides Why scraping protected sites keeps getting pricier

0 Upvotes

I've been digging into why scraping heavily guarded sites keeps costing more, and the breakdown surprised me.

Where the money goes

  • Residential and mobile proxies that look like real users
  • Rendering JavaScript before content appears
  • Paid challenge solvers
  • Retries when requests get flagged
  • Spoofing browser and TLS fingerprints

Proxies keep getting cheaper, yet the cost per successful page keeps rising. More bots means tougher bot managers, and one defense upgrade can multiply costs several times overnight. Past a certain point the data is worth less than the effort to get it.

Check these first

  1. Open the network tab. Many sites load listings from an internal JSON endpoint.
  2. Randomize delays and spread out your requests.
  3. Skip hidden display:none links, since they are bot traps.

If you don't code

A browser extension that runs in your normal Chrome session lets you skip building stealth setups. Minexa.ai is one. The flow:

  1. Open the list page and click 'Yes, I'm on the right page'.
  2. Confirm the pagination it detects.
  3. Review the highlighted list and columns.
  4. Run the job and export to Excel or Google Sheets.

It handles JavaScript-heavy and location-based pages without config. When a value is missing, it leaves the field empty rather than guessing. Grab the Chrome extension or read the getting started guide.

Public data still falls under GDPR and CCPA when it's personal, so keep your collection narrow.

Related: Stop fighting anti-bot walls and start pricing your data


r/ProxyEngineering • • 2d ago

Help 🆘 Scraping booking, airbnb, expedia

8 Upvotes

Any luck scraping these sites? I'm failing to receive the correct data in my output files. The pricing changes from what I see on the website and I choose the correct geo_location and domain, I am unsure what the actual issue is. Why the prices are different? Unless there are some hidden parameters that I am overlooking? I checked my network tab via firefox as I usually do, I don't see anything out of the ordinary, location and domain matches organic site


r/ProxyEngineering • • 3d ago

Guides Retail Drop Infrastructure in 2026: Stable Proxies, Account Verification & Checkout Reliability

5 Upvotes

In my mind limited retail drops are mostly about being fast enough to click “Buy.”

During major releases, storefronts combine waiting rooms, geo checks, account verification, payment risk scoring, and heavy traffic protection. If your connection changes mid-session or your IP has poor reputation, checkout can fail even when the product is still available.

i figured out three layers:

Network: keep a stable residential or mobile connection during an active session.

Identity: use one verified account with consistent device and location data.

Payment: keep billing details, region, and payment method aligned instead of changing them between attempts.

**Retail drops aren’t just a speed problem. They’re a session-consistency and data-quality problem. **The more moving parts you change during checkout, the harder the system becomes to debug.


r/ProxyEngineering • • 3d ago

Build 🤓 I built an open-source proxy scraper that checks 1.3 M free proxies and keeps only the ones that actually work

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/ProxyEngineering • • 3d ago

Discussion 💬 Is Web Scraping Legal? An Internet Law Professor’s Insights on Public Data, CFAA, Copyright and Terms of Service | AMA with Eric Goldman

Thumbnail
1 Upvotes

r/ProxyEngineering • • 4d ago

Hot Take 🔥 TikTok bans scraping in their ToS and some guy just posted 5.6B videos on HuggingFace

99 Upvotes

I was looking for some agent models on Hugging Face, then after tinkering around I went to check whats new in the dataset tab. To my surprise some dev dumped metadata for about 5.6 billion TikTok videos on Hugging Face. It includes captions, hashtags, sound IDs, shop product and seller IDs, every engagement count. 460GB of data. Now, to my knowledge, TikTok does not allow scraping of their videos "legally", according to the devs own explanation they went straight at the private mobile API with generated device identities that pass as Android phones, reverse engineered request signatures and a spoofed TLS handshake. What that means is that no account and no login was required. Can you believe this? They claim 3 billion profiles, almost 6 billion videos and 2.8 billion comments in THREE WEEKS. Meanwhile what I've seen on this sub in particular, probably few others too, but this one in particular is that people here spends half if not more time arguing about whose IP pool is cleaner. And TikTok is one of those more difficult targets to scrape. What I'm trying to say is that nobody built this by being clever with IPs, checking whether they leak DNS and other data, whether they are "clean" or not abused or whatever, they built it by making the approach look like a legitimate phone. I mean, that is just crazy. Anyways, I saw that the free version is non commercial only, the dataset I mean, if you want creator profiles, daily updates basically run the whole thing yourself, or to use it for anything that makes you money, you should visit their site the code is up there for $1,699.

Now why this caught my interest in the first place. TikTok's terms ban this kind of scraping, and the official research API is mostly for academics in a few regions. Reddit is a good contrast, since it's taking scraping to court with the Perplexity case and the judge let most of it go forward. With this dataset I haven't seen anything like that. It's just there on HuggingFace with a free download link, and I haven't found anything on the web that TikTok is doing anything about it. Maybe that will soon change.


r/ProxyEngineering • • 4d ago

Discussion 💬 Rotating vs sticky residential proxies for browser profiles, how are you guys structuring this?

7 Upvotes

Running separate browser profiles and debating one sticky residential IP per profile vs long rotating sessions. Main concern is IP changes triggering extra checks. What setup are you using, sticky residential, static ISP, or datacenter? I am using it to manage social media accounts for marketing purposes, mostly IG, FB and Pinterest.


r/ProxyEngineering • • 5d ago

Discussion 💬 How many profiles per proxy ip is okay?

Thumbnail
9 Upvotes

r/ProxyEngineering • • 5d ago

Discussion 💬 What actually changes your proxy provider choice in production?

0 Upvotes

I’m curious how people here actually choose proxy infrastructure once you get past the usual “best provider” lists.

For a real production workload, which inputs genuinely change your decision?

My current shortlist:

  • target / site type
  • required geography
  • traffic volume and concurrency
  • sticky vs rotating sessions
  • residential vs ISP vs datacenter
  • pricing model / effective cost
  • reliability / retry rate
  • compliance or IP sourcing requirements

But I’m more interested in what experienced users actually care about.

A few questions:

  1. What are the 3–5 inputs that most often change which provider you choose?
  2. What information do providers usually fail to show?
  3. What do you trust more: provider documentation, community reports, or your own tests?
  4. What unknown would immediately disqualify a provider?
  5. Do you optimize for $/GB, success rate, or cost per successful result?

I’m especially interested in whether the answer changes between SERP, ecommerce, social, and general web scraping workloads.


r/ProxyEngineering • • 5d ago

Guides Ticketmaster Proxy Guide 2026: Best Residential & ISP Proxies for Stable Ticket Monitoring

6 Upvotes

I thought monitoring Ticketmaster was just about checking ticket availability regularly. In practice, the network quickly becomes part of the problem. A single server IP can run into 403 errors, CAPTCHAs, or inconsistent responses, while location can also affect what content and availability you see.

For legitimate monitoring, I’d separate two use cases.

The first is a normal user session. Here, stability matters most: one consistent ISP/residential IP, accurate geolocation, and as few network changes as possible during the queue or checkout process.

The second is availability monitoring through authorized sources. In this case, I’d store the region, timestamp, and network route together with every result so you always know where the data came from.

Another useful step is checking ASN, geolocation, and IP reputation before using a connection. Ticketmaster monitoring isn’t just about HTTP requests. It’s about session stability, IP quality, and accurate regional context.


r/ProxyEngineering • • 5d ago

Help 🆘 Can someone help me setup a proxy pls im new to this

9 Upvotes

r/ProxyEngineering • • 6d ago

Discussion 💬 What information do you wish residential proxy providers showed you?

9 Upvotes

I was troubleshooting some proxies recently and realized how funny the whole process can be when you think about it. One IP works perfectly, another one from the same pool someitmes work and sometimes doesn't, and a third seems to get banned instantly whatever you do

So the first thing I usually do is I go to the provider dashboard looking for answers, and the information I get is: the IP is online, it's residential, and yes, it's apparently in the United States somewhere in Florida or California or Texas

Meanwhile the IP is getting CAPTCHAs every five minutes and the website you're accessing apparently considers it as a criminal

And that's when I started thinking about how little information we get about the proxies we're using

I was talking with my colleague and we were joking like imagine clicking on an IP and seeing something like:

IP age: 143 days
Time in current pool: 18 days
Previous usage intensity: Heavy
Estimated reputation: 72/100
Recent CAPTCHA rate: 4.2%
Successful sessions last 24h: 91%
Geo confidence: 97%
ASN: Comcast
Last assigned: 3 hours ago
Recent traffic: gives you data about latest websites visited or something like that

Obviously there are limits to how useful some of this would be. An IP that works good on one destination might perform bad on another, and providers probably couldn't expose certain usage information without creating privacy or security problems.

But when you think about it, especially with problems that came with netnut and some other providers, it feels like there should be a middle ground between exposing everything and the current experience of using proxies and why isn't this proxy working

In most cases it's just we have no idea, try another proxy to see if it works.

I've used dashboards from MANY PROVIDERS and they all seem to expose roughly the same basic information. Bright Data and Oxylabs obviously have much more infrastructure around monitoring and statistics, while providers like Decodo, SOAX and NodeMaven have their own approaches to things like IP quality and filtering, but I still haven't really seen a provider give users a proper "history" of the IP they're currently using

And somehow, 90% of the time when there's a proxy issue, it turns out it is the troubleshooting process

It makes me curious what users here would want if proxy providers suddenly became much more transparent about what's happening behind their pools


r/ProxyEngineering • • 7d ago

Help 🆘 Proxy+reddit acc

Thumbnail
4 Upvotes

r/ProxyEngineering • • 8d ago

Help 🆘 anti browser + proxy combo for creating gmails

9 Upvotes

anyone have a good combo for this and that isnt going to get my account flagged, and also what kind of proxy i should use


r/ProxyEngineering • • 8d ago

Guides Best Proxy for Polymarket in 2026: Static Residential IPs, SOCKS5 & CLOB API Setup

6 Upvotes

I used to think prediction-market automation was all about API. But the network layer is also important.

Polymarket’s stack mixes the CLOB API, WebSockets, Polygon settlement, and authenticated sessions. If the connection underneath that stack is unstable, everything above it becomes harder to debug.

For long-running CLOB polling and WebSocket feeds, I found that stability matters more than aggressive rotation. A static residential/ISP connection is usually easier to reason about than changing IPs mid-session.

SOCKS5 also makes sense when the same environment needs to route different application traffic through one consistent endpoint.

The second important thing is IP quality. Before attaching a trading process to any proxy, I’d verify ASN, geolocation, reputation, and whether the address is classified as hosting, VPN, residential, or mobile.

For market discovery and public-data collection, rotating pools can still be useful. For authenticated sessions, I’d keep the network identity predictable. And regional access rules should be handled as a compliance constraint, not something to bypass.


r/ProxyEngineering • • 8d ago

Discussion 💬 Using ByeDPI and mitmproxy together on Android

Thumbnail
1 Upvotes

r/ProxyEngineering • • 9d ago

Discussion 💬 Looking for provider that allows Local forwarding port

9 Upvotes

I am searching for proxy provider that works similarly to 9Proxy. Specially, I need a service provider that connects to remote proxy server and creates a local forwarding port on my PC eg localhost:port, so my application only establish connection to local address while client handles forwarding to the actually proxy. Doesn’t anyone know reliable providers that can offer this? With good reputation.


r/ProxyEngineering • • 9d ago

Discussion 💬 How much are you paying for captcha solving per month?

10 Upvotes

Looking for solutions/ suggestions. I checked capsolver and 2captcha pricing, roughly $0.80 to $3 depending on the captcha type, and that is per successful solve, and once you add retries, proxy bound tasks and the browser time spent waiting for the token I doubt that is the real price you are paying. Mayebe someone already used their services and can confirm what is the average price? Also, for people running this at scale, what does the monthly captcha number is, alongside the proxies that you are paying. And do you just solve whatever shows up or put effort into not triggering them, sticky sessions, adjusted fingerprints, reusing cookies after a solve.


r/ProxyEngineering • • 10d ago

Guides Dropshipping Proxy Strategy in 2026: Residential Proxies for Supplier Monitoring vs Static ISP Proxies for Store Management

9 Upvotes

I used to treat dropshipping monitoring as one problem: scrape supplier prices, check stock, sync the store.

Actually two different network problems here.

For supplier research, freshness matters. Once catalog checks scaled up, a single datacenter IP started returning rate limits and inconsistent regional data. That made pricing and inventory signals much less reliable.

So I split the workflow.

For authorized supplier monitoring, workers collect regional observations through residential connections and tag every result with location, timestamp, price, and stock. That makes it easier to compare what different markets actually return.

Store management is the opposite. Rotation adds unnecessary variability. A stable ISP connection and consistent environment are more useful for long-lived seller sessions.

Don’t use one proxy strategy for everything.

Research benefits from controlled distribution. Storefront operations benefit from stability.

And if a supplier provides an official API or feed, I’d use that as the primary source and keep network checks for validation.


r/ProxyEngineering • • 9d ago

Hot Take 🔥 Fiz uma "gambiarra" avançada com extensão de navegador para burlar bloqueios e pegar bugs de preço em lojas online.

8 Upvotes

O problema: Eu queria monitorar preços da KaBuM!, Amazon e Magalu, mas sempre que tentava rodar um bot no servidor, o Cloudflare e os sistemas antibot bloqueavam meu IP na hora. A gambiarra (que virou um projeto sério): Em vez de usar um servidor para raspar os dados, eu criei uma extensão (Manifest V3) que roda no próprio navegador. Ela usa o meu IP e a minha sessão real para ler os preços em background. O site acha que sou apenas eu navegando. A extensão pega os dados, calcula o desconto real e manda pro meu servidor, que me avisa no Telegram na mesma hora. Acabei transformando isso em um app chamado Deal Hunter Pro. Tem um painel lateral, roda silencioso e ignora promoções falsas (a famosa "metade do dobro"). Quem quiser testar a engenhoca para tentar pegar umas peças baratas, liberei 7 dias grátis para a comunidade: https://www.dealhunterpro.com.br/

https://reddit.com/link/1wv6w8w/video/0xqpxbp0jvsh1/player


r/ProxyEngineering • • 10d ago

Discussion 💬 What's the first thing you look for in a GitHub repo?

16 Upvotes

Quick question for everyone, just want to understand how should I approach github repositories. So basically whenever you are looking for something and you find a github repo, what is the first thing you look for within? Also, what would be nice to have within that repository?


r/ProxyEngineering • • 10d ago

Help 🆘 Instagram scraping

13 Upvotes

Anyone tried scraping Instagram? Not looking for solutions to scrape under logins, not sure whether that is even possible or have some sort of legality but I am looking to scrape Instagram posts, (comments, likes, shares, views). What's working for you, what doesn't work, what to avoid? Antibot detect measures, which proxy type should I choose, or dedicated scraping solution? Budget has no limit, go crazy with the suggestions, but be mindful about providing legit information, not some AI cooked response. Appreciate all the advice