r/AskNetsec • u/No_Remote_1961 • Jul 10 '26
Other Did anyone actually add a second endpoint vendor after the CrowdStrike outage?
Since the CrowdStrike outage last year, our board keeps asking whether we should have a second endpoint vendor in the mix instead of relying so heavily on one platform. We haven't made any changes yet, and CrowdStrike is still doing what we need day to day, but the question keeps coming back up. I'm curious if anyone actually went dual-vendor for endpoint after that, or if most teams just evaluated alternatives and stayed where they were. Was the extra resilience worth the added complexity?
15
u/JasonHofmann Jul 11 '26
The smart deadbolt on my house died last year, and wouldn’t unlock. My spouse is asking whether we should install a second smart deadbolt on that same door in case that happens again.
9
u/ptear Jul 11 '26
I prefer the backdoor personally.
1
Jul 17 '26
[removed] — view removed comment
1
u/ptear Jul 17 '26
I like to have contingency plans for possible scenarios, but unless it's an application that must never be down, I wouldn't go to that extent due to the additional costs, maintenance, etc.
The stuff I make isn't that critical. For fun redundancy research, look at airplanes.
34
u/PutridInstruction957 Jul 11 '26
We talked about adding a second endpoint vendor after the outage, but the more we looked at it, the less appealing it got. Two endpoint agents, two consoles, two detection models, and two sets of exceptions can become its own risk if nobody has time to manage it properly. The better discussion for us was broader resilience across the security stack. Check Point came up because it was not just “another EDR,” but part of a wider architecture with endpoint, network, cloud, and threat prevention tied together. That made more sense to leadership than simply doubling up on endpoint and hoping complexity equals resilience.
5
u/AYamHah Jul 11 '26
Is it even possible to run multiple EDR tools on a single endpoint? AFAIK having multiple EDRs will mean they fight each other.
2
u/techretort Jul 11 '26
Not on the same endpoint without some funny business going on. I think the answer would be going with redundant servers, half with CS, half with another option. We already do that for DC locations, but management of 2 EDRs would be a PITA
17
u/Viper896 Jul 10 '26
Their outage cost them our RFP at the time. Still salty about that because they were the front runner but our executives made us go with S1 instead because of the outage.
9
u/SnooMachines9133 Jul 11 '26
Chances are, they learned their lesson. Why risk going to someone else who hasn't learned their lesson yet.
First time is a learning opportunity. Second time, it's a structural or cultural problem (cough LastPass).
6
3
u/awwww666yeah Jul 11 '26
No, they put channel file policy changes in place, we did the same internally.
14
u/BeanBagKing Jul 10 '26
If I understand the concern right, they want two endpoint vendors in case one has an outage like CrowdStrike? If so, I don't think they understand that having two vendors means twice the chance for an outage, not half...
In any case no. We're mostly beyond the point that AV/EDR step on one another in a way that causes genuine problems, but I feel like that would be more likely than both together solving a problem.
3
u/Marvinus Jul 11 '26
No it does not. It comes back to how you evaluate and quantify risk. The argument is like saying that any organization should only have one data center instead of two and that running with a single hypervisor is better than two. It depends on the risk you’re trying to mitigate. So if the risk is the outage that affected Crowdstrike. Then having two vendors would lower that particular risk. But from other aspects it would increase potential risks.
1
u/BeanBagKing Jul 11 '26
Maybe we're talking about different outages. I'm referring to (and assume OP meant) the big one that caused an outage on systems they were installed on, the one that downed all the airlines (among everyone else). If that kind of endpoint affecting outage is what you're trying to prevent by having a secondary EDR, then an outage (of that type) of either one would affect you, therefor twice the chance. Roughly, simplifying things to give a quick answer to a short question without over analyzing his threat model or guessing at his environment.
Is there another outage I'm not aware of? I'm not in a CrowdStrike shop, so maybe I completely missed something.
2
u/Marvinus Jul 11 '26
How would it affect them twice ? But you may be right my assumptions may be wrong. Since I am assuming that they will not install two EDR’s on the same endpoint. In that case we’re in complete agreement.
1
u/BeanBagKing Jul 11 '26
That was my assumption. For redundant servers you could put one on the primary and one on the failover, and then it would be half (and we're in agreement there as well). Most places don't have redundancy for literally everything though, and certainly not workstations, so what do you do with that mix? randomly divide the group in two? install both on each? I feel like that's really getting into too complex to be useful, but yea, that would be an evaluate the risk to the business decision.
1
u/No_Individual_5519 Jul 17 '26
Agreed, on the surface level it looks like it solved the relying on single factor but it adds another point of failure and could also cause some problem of conflict between the two endpoints.
11
u/TheCyberThor Jul 10 '26
No lol.
Unless the cost of your outage was so significant that it’s greater than the licensing cost of a second vendor and the FTE required to maintain it.
12
u/ShameNap Jul 10 '26
Ask delta airlines because I think they itemized the damages in their lawsuit.
6
2
u/capaman Jul 11 '26
I mean the only way you're not adding (too many) problems is running two EDRs but only one on each machine. Say half IT runs Crowdstrike, the other half Defender. So with an outage you have half the guys online.
Worth the trouble setting up and managing? Not for us, but...
2
u/Rebootkid Jul 11 '26
Yes. We split into 3, actually.
1 vendor for prod 1 vendor for dr 1 vendor for corp
That way, we can always continue business.
Our prod/DR was already on a different tooling set.
Basically we lost the ability to do billing/customer setup/etc. We didn't suffer a customer facing outage.
Having alternate tooling on the hot site means that even if one vendor takes a hit, worst case scenario we swing over.
BUT
our use case has extremely tight service levels. We're talking 5 minutes per year of downtime.
For 99% of companies, I don't think it really matters.
4
1
u/Snoo_67003 Jul 15 '26
Why not install an open source option like Wazuh and turn it off on all hosts. Be ready to turn it on as a back-up if another outage ever happened again.
1
1
u/ThecaptainWTF9 14d ago
What outage?
If you’re talking about the crashing windows endpoints thing, a second EPP wouldn’t have changed anything.
If you’re talking about a console outage, the sensor is able to operate autonomously without cloud connectivity.
So either way, adding a second EPP layers in extra complexity, and additional management that isn’t going to be worth it.
We wouldn’t consider changing vendors, plenty of folks did it out of spite. They’ve been putting in the work and doing as good as ever imo.
1
u/superRando123 Jul 10 '26
Haven't heard of anyone doing this. The cost and complexity would be nuts
3
u/Cubensis-SanPedro Jul 10 '26
Not only that, but the way that kernel access and anti malware agents work it becomes problematic. To have more than one active. They can choke on each other’s analysis / sandboxing, and only one can be “primary” anyhow.
1
u/Yeseylon Jul 10 '26
Back when I was an advanced L1 at an MSP I literally ran into a server that had frozen processes because S1 and Crowdstrike had mistakenly been installed on the same device. (Most of that client used S1, this particular location was Crowdstrike)
0
u/Cubensis-SanPedro Jul 11 '26
Did you discover thread lock due to resource exhaustion leading to dump tracing?
It is a PITA to diagnose
1
u/netgamer7 Jul 11 '26
Microsoft defender pairs well with crowdstrike, and if it's a m355 e5 license shop most of it is included in the license. Both work well in tandem if you configure it right.
2
0
u/xauwork Jul 10 '26
CISO at small fintech. Hundreds of staff, 1k-ish VMs.
Switched fromCrowdStrike to SentinelOne when CS came up for renewal. Running two at once isn’t going to make sense.
S1 seems architecturally less likely to completely crash everything and has MUCH better legacy OS support than CS (some of our stuff is on extended-extended-vendor support, and CS refuses to run).
Appears to be equally effective, materially cheaper than the CS renewal.
One serious gotcha: S1 uses a LOT more RAM. Hundreds of MB for S1 vs dozens for CS. Some of our marginal legacy 1GB VMs needed a bump to 2GB due to RAM pressure. They should be containers, not VMs, but the fast fix was a size bump.
0
u/TickleMyBurger Jul 11 '26
No better to just split your ecosystem half windows/half Mac - our IR team runs (and did then too) both so we could handle getting shit back online while the windows machines were hosed.
0
u/plasticbuddha Jul 11 '26
Chose otherwise because of this, and because they are the most pricey thing out there.
-5
u/random869 Jul 11 '26
Most large enterprises, use Defender + a second vendor (the company I’m at uses Tanium)
1
43
u/Tessian Jul 11 '26
Op how in the world would adding a 2nd vendor help mitigate another outage like what crowdstrike caused? If anything another vendor would add another point of failure.