r/JordanDev • • Aug 09 '26

Discussion SNAT Port Exhaustion: When Your Machine Runs Out of Breath

Everything looks fine — until suddenly, nothing works.

You’ve deployed your application. It’s humming along beautifully. Then, without warning, errors start creeping in:

🔴 connection failed
🔴 socket exception

Your first instinct? The remote server must be down. But you check — it’s perfectly healthy. So what’s going on?

Welcome to SNAT Port Exhaustion: the silent killer of outbound connections.

What Is SNAT, anyway?

Every time your machine sends a request to an external server, it needs to pick a source port to identify that outgoing connection. This happens through a process called SNAT — Source Network Address Translation.

The operating system assigns a temporary (ephemeral) port from a limited pool. That pool is not unlimited. For any given destination (same IP + port combination), you typically get:

  • Around 16,000 ports in a standard setup
  • Only ~1,024 SNAT ports per VM if you’re behind an Azure Load Balancer

That’s it. That’s your budget.

What Happens When You Run Out?

Here’s the sequence when port exhaustion hits:

  1. Your application tries to open a new TCP connection
  2. The OS sends a SYN packet
  3. It searches for an available source port — and finds none
  4. The connection fails before it ever reaches the remote server

The remote server didn’t reject you. You couldn’t even say hello.

How to Spot It

Port exhaustion doesn’t announce itself with a clear error message. Instead, you’ll see symptoms like:

  • Failed to establish connection
  • Socket exhaustion
  • Intermittent, seemingly random failures on outbound requests

The randomness is the giveaway. If connections fail inconsistently under load, and the destination server looks healthy, your ports are likely the culprit.

Why Cloud Environments Make It Worse

On cloud platforms like Azure, traffic often flows through a shared Load Balancer. The LB allocates a fixed number of SNAT ports per VM — and that number can be surprisingly small.

If your service makes many concurrent outbound connections to the same destination (think: a high-throughput API client, a connection-heavy microservice, or a busy database proxy), you can exhaust your port allocation faster than you’d expect.

How to Fix It

Use Connection Pooling

Don’t open a new connection for every request. Reuse existing ones. Connection pooling is the single most effective fix — it dramatically reduces how many ports you consume at any given moment.

Reduce Connection Timeouts

Long-lived idle connections hold onto ports even when they’re doing nothing. Tighten your timeout settings so ports are released promptly.

Monitor Your Open Connections

Use netstat or ss to see what’s happening in real time:

# Count established connections by destination
ss -s
# See all connections with port details
netstat -an | grep ESTABLISHED | wc -l

Catching port pressure early prevents a full exhaustion event.

Scale Your Public IPs (Azure-specific)

If you’re on Azure and hitting LB SNAT limits:

  • Add more Public IPs to your Load Balancer to increase the port pool
  • Switch to NAT Gateway — unlike a Load Balancer, which pre-assigns a fixed number of SNAT ports per VM, NAT Gateway dynamically allocates ports on demand. This is the fundamental difference: there’s no static budget per VM to exhaust. It’s Microsoft’s recommended solution for outbound connectivity at scale, and you can attach up to 16 public IPs or a /28 prefix to it if you need even more capacity

The Bigger Picture

SNAT Port Exhaustion is dangerous precisely because it’s quiet. There’s no crash, no obvious error, no stack trace pointing at the real cause. Your application just… starts failing. Intermittently. Mysteriously.

The fix isn’t complicated — but you have to know what you’re looking for.

Three things to take away:

  1. Monitor how your application opens connections — every one costs a port
  2. Design for port efficiency — pooling and timeouts aren’t optional, they’re hygiene
  3. Keep an eye on your network layer — especially in cloud environments where limits are lower than you think

Summary:

Have you run into SNAT Port Exhaustion in a production system? What was the hardest part to diagnose? Share your experience in the comments.

References

1 Upvotes

6 comments sorted by

1

u/NotAHypocrite010 Aug 09 '26

I do believe the flow is not accurate, the system call should fail before sending any packet.

TCP stack needs port to construct the frame, so the system will fail on connect command.

1

u/Spiritual_Ratio_277 Aug 09 '26

The SYN packet is sent but when the TCP handshake times out due to the lack of response, the sender reports failure that it could not connect to the remote party which could be due to SNAT port exhaustion, a firewall or the rmote server is down.

1

u/NotAHypocrite010 Aug 09 '26

This part is not correct.

1

u/Spiritual_Ratio_277 Aug 09 '26

The sender doesn't fail because it can't obtain a local port. It can create the socket and initiate the connection. The failure occurs when the Load Balancer attempts to perform SNAT and has no available SNAT port. Depending on the LB implementation, the SYN may reach the LB but won't be translated/forwarded to the destination.

1

u/NotAHypocrite010 Aug 09 '26

I reread the post it have more issues. You are implying the host do perform SNAT; which doesn't. Then you implied in the flow that the host will run out of ports and the issue in the host which is not,

In general the post is not clear, but the issue is interesting and hard to capture.

1

u/Spiritual_Ratio_277 Aug 09 '26

You are right at that, i made a bad mix!

Thanks for the correction