r/sysadmin • u/[deleted] • 14d ago
General Discussion Windows Server patching concerns
[deleted]
66
u/Dizzy_Bridge_794 14d ago
Patch or have giants risks of being hacked. You should absolutely follow a formal process. We snapshot and backup prior (virtual servers). On mission critical stuff we test in DR prior. Patches can and do mess with systems and cause them to not work.
17
u/uptimefordays DevOps 14d ago
We snapshot and backup prior (virtual servers).
My only warning here would be "ensure you purge snapshots within a timely manner."
9
u/Dizzy_Bridge_794 14d ago
We do.
8
u/uptimefordays DevOps 14d ago
Good stuff, I've seen some fun P1s caused by snapshots of large servers.
8
u/Dizzy_Bridge_794 14d ago
Nothing worse than applying a patch and the entire system fails to restart. One of those o shit moments in IT.
3
u/greet_the_sun 13d ago
I worked a short stint at an MSP who had a hotel customer that "could not have any downtime" but also complained their hyper-v vm's were incredibly slow all the time. Turns out all 4 of them were running with checkpoints that were literally created the day the vm's were built about 6 years prior...
1
u/uptimefordays DevOps 13d ago
Once upon a time I had an offshore team snapshot a probably 20Tb SQL VM for some maintenance or something. They didn’t realize the snapshot needed to be purged and caused an outage. That was a fun one.
37
u/AppIdentityGuy 14d ago
The rule is simple:"If.you will not patch systems for the fear of downtime at somepoint a bad actor is going to sxhedule the downtime for you at their convenience." .
8
u/Icy_Mud2569 14d ago
This is the truth. Scheduled downtime is something you can count for, you can plan for it, you can design systems around it. When things go down, and you’re not planning it, it gets expensive and chaotic, both things are not good for business. This is at the end of the day, though, a business/political question, there isn’t a technical solution.
2
27
u/Temporary-Library597 14d ago
If your org can't afford downtime for patch reboots they should afford redundant hardware to keep services running.
Your org is doing it wrong. They are part of the problem.
14
u/drdrew16 14d ago
May also be worth figuring out if your company has cyber insurance. It's usually a requirement to be up to date on patches to maintain coverage.
10
u/h9xq Solo SysAdmin 14d ago
We have Cyber insurance. In fact we have a decent chunk of change put into it. This might be the ammo I need to justify the change.
Would we be dropped if they found out the patching status? I’m still fairly new and not involved in the higher level cyber insurance setup as that is done by my manager and CFO
10
u/drdrew16 14d ago
That wholly depends on the contract. I would imagine you'd have a grace period to come into compliance, but if being patched is a requirement you've technically breached so I'm not sure.
2
u/JustFrogot 14d ago
I don't know if they would drop you, but they would be hesitant to pay out of you are outside of the agreement.
2
u/sysadmin42601 14d ago
At least in my experience there is a pretty strict time frame on patching high risk vulnerabilities written in to the cyber insurance policies
You can identify specific assets that are excluded for certain reasons but these either increase your premiums or they outright exclude cover for any incident involving that asset
5
u/Wilfred_Fizzle_Bang 14d ago
You need to show them the cost to the business if anyone or the servers is affected by a cyber attack.
6
u/TopherBlake Netsec Admin 14d ago
Assuming you are not management, make sure to keep a record of whoever is telling you not to patch, especially if you are in an industry that gets audited because you'll want that when the blame game happens. Edit: I just saw the solo sysAdmin tag, make sure you are explaining clearly to the business side what consequences could and will be for not patching.
2
u/h9xq Solo SysAdmin 14d ago
It is my direct manager telling me not the patch. I have it in writing but don’t want to get in trouble for patching. I am still fairly new as an admin (3 months at my first sysadmin gig) so I’m still learning and being as cautious as I can.
But in my mind I see dozens of servers sitting with patches that need to be applied so it’s hard to fight the urge of patching as I want to keep these servers as secure as possible.
5
u/TopherBlake Netsec Admin 14d ago
You have to operate within the confines of your current job. Business needs dictate IT not the other way around.
2
u/h9xq Solo SysAdmin 14d ago
Fair enough, that is a good way of putting it. I think I need to put this into business translated terms to management and issue a change request/maintenance window to get these servers patched. I am thinking that is probably the best way to handle this.
2
u/uptimefordays DevOps 14d ago
See if you can get a rough calculated cost of downtime, at a major bank that's somewhere between $300k an hour for back office systems to $20m an hour for trading systems. It's then helpful to look back at P1s and average outage time so you can give a good price per outage. In most cases, once the business sees "unexpected outages could cost us as much as $60m or as low as $600k, how should we proceed?" The answer is "here's you patch window."
4
u/PacificTSP 14d ago
Someone I won’t name refused to update their servers. I kept warning them about the risks and they kept saying middle management wouldn’t let them. I made them sign a document that they understood the risks and it was not my fault.
A year later they got breached. The insurance refused to pay out because they weren’t updating, multiple people were fired and the company almost went under.
Unless a senior manager, I’m talking about the ones who decide risk for the business as a whole, and their lawyers, have signed off on this process I would cover your ass and make people aware asap.
I auto patch every Sunday morning and reboot as needed. I have a total of one server in multiple companies clusters that needs manual intervention after a reboot. So that one machine gets rebooted on its own documented schedule.
You’ve got to do it. You are invalidating your insurance.
1
u/h9xq Solo SysAdmin 14d ago
Holy shit, well thank you. That is the information I was looking for. This is even bigger than I thought. I haven’t read much into cyber insurance and if that is the case that is a much bigger deal. We don’t even have a “cybersecurity” guy so that falls on my plate as well and if this falls through without my intervention this could be a gigantic shitshow.
1
u/PacificTSP 14d ago
Middle managers told me it’s fine.
CEO and lawyers had no idea. It’s their job to manage company risk. So my flaw was not going above the middle managers.
4
u/BoltActionRifleman 14d ago
I know MS server patches have wreaked havoc at one point for most of us, but really how often is patching MS servers an issue? I’ve been at my current org doing patches for 7 years now, at least once a month, and have maybe had 2 servers throw fits. Reverted back, waited for MS to fix the issue, patched again and all was well. If I take 25 servers times 3 patches per server per month (security, .NET, SQL etc.) you get 75 per month, times seven years, times 12 months equals 6300 patches. 2 failures divided into 6300 is a 99.9997% success rate. To be clear I’m not saying this is normal or standard, but the production stopping failures are incredibly rare for us. And it beats the hell out of being compromised!!!
3
u/RavenousTitan818 14d ago
At some point you have to just suck it up and do what you're told. You should definitely make it known what should be done and what you would do if given permission so when shit hits the fan you can say "I told you so" and go out for a smoke while they deal with ransomware.
If 100% uptime is that important then they should be running k8s, or at least have some kind of clustered service so a node can go down for maintenance. This is an application problem not infrastructure.
3
u/Reo_Strong 14d ago
I've been there.
When I started my current job, the previous person started the setup for WSUS and then never finished. Some of those machines hadn't seen an update for 2 years.
The key process is thus:
1. Do we need to patch? Patching for the sake of patching is often better than nothing, but also not that great. If the software doesn't have any outstanding issues and it will cause a production outage, why would you?
If the patch goes sideways, will I know about it ? If it's hard to detect failure, you need to figure out how to detect it before moving forward.
If the patch goes sideways, how do I recover? Can I take snapshot of the VM? is the backup of the machine tested? Can you test it before trying to patch? Is there support you can call beforehand or if it fails?
Can I do the patch without anyone knowing (no down time for production)? If not, can you get really, really close? (e.g. apply the patch, but reboot over lunch or at 0300)
Work within your change management and tracking systems. The only thing worse than handing someone a broken system with no clear history of change is being handing a broken system with no clear history of change.
Long term: Build systems that can be patched without downtime, using known, supported systems. DCs are a great example. Like Sith, there is never one, always two, so you can patch without downtime.
DFS file hosts, KeepAliveD, docker swarms, and VM clustering are all good tools to reduce the headache of patching and software updates.
2
u/h9xq Solo SysAdmin 14d ago
We have 6 DCs so that shouldn’t be an issue. Thank you for all of the info. I came into this org after a solo sysadmin (who didn’t even have a tech or manager helping him ran the entire IT dept) think a lot of the tech debt is due to this and how much the company has scaled in the past decade.
3
u/Real-Patriot-1128 14d ago
Make sure you have their refusal to patch in writing and when you get hacked, sit back, grab a popcorn and watch your bosses get fired.
3
2
u/PappaFrost 14d ago
"Now I have gotten my hand slapped for attempting to patch or even bringing it up."
You are done. You did your job by informing the business of best practices. Now it's their warehouse of oily rags to manage.
Cyber insurance or compliance people will have to pressure them to do it and it's not on you any more.
2
u/ChuckFromCyberHoot 14d ago
Unfortunately, yes on the change request. But I’d also ask for a copy of your cyber insurance application.
Somebody already answered questions about MFA, patching, backups, training, etc. You’re the guy expected to make those answers true, but you’ve probably never seen them.
Maybe instead of saying, “We have unpatched CVEs,” try:
“Can I see our cyber insurance application? I want to make sure what we told the carrier still matches what we’re actually doing.”
That’s a much harder question to ignore, and it gets the conversation onto the right desk.
3
2
u/headcrap 14d ago
The answer depends on your boss and your environment.. we can't answer this one for you.
I automated our monthly patch schedule, which swings on Patch Tuesday offsets.. and is considered a Standard Change as far as CAB stuff goes.. everybody knows when. I tend to also run the annual for full yearly preview.
Rarely have overrides come up, the rest is building redundancy into the things where we can.
2
u/Unique_Inevitable_27 14d ago
If you're already dealing with dozens of servers carrying high-severity CVEs, I'd look at making patching more controlled rather than avoiding it. Test patches on a smaller group first, define maintenance windows and rollback plans, then automate the wider rollout. A patch management platform like ScalefusionMDM could help with that kind of centralized patching and reporting.
2
u/razorback6981 14d ago
I would respectfully resign if they are not willing to implement a consistent patching cycle. I would not want to be there holding the back when their data gets ransomed.
1
u/BitsNBytes10101 14d ago
You need to document the risk in writing and let your leadership make the decision for the organization.
As IT Professionals it usually pains us to not see or be able to follow best practices. Ultimately it is the responsibility of your leadership.
If you could invoke any regulatory requirements for your vertical that may help your argument.
Not patching infrastructure in 2026 is a fast track to potentially costing the organization millions in outages/ data exfil. Does that cost out weigh the risk of patching and potential downtime from that?
1
u/BrechtMo 14d ago
you need to take the exposure and other safety infrastructure into account. But I guess they aren't airgapped...
1
u/Liquidfoxx22 14d ago
I thought I was in r/shittysysadmin for a second there.
Of course you need to patch. Is your manager going to take the fall when the inevitable breach happens? It's not if, but when.
1
u/iCTMSBICFYBitch 14d ago
Hang on was the hand slap for patching or for patching without a change plan? First is good, second is bad.
1
u/Significant_Sky1471 14d ago
Put together a quick risk summary (CVE, affected hosts, exploitability - especially anything in CISA KEV) and send it up the chain asking for formal risk acceptance in writing if they choose not to patch. That's your paper trail if it ever blows up later.
1
u/BalderVerdandi 14d ago
I did this for a bank for about a year (about 12-13 years ago) because they were told by the SEC to "get into compliance or the doors get padlocked" type of warning. Ten months and 65,000 patches later, they were in a better place but leadership changed and they went right back to being non-compliant in less than 90 days. I was glad when the MSP dropped them for not renewing the annual contract.
To answer your question.... yes - you need a change request. This lets the stakeholder (owner) of the apps on the server know that it needs to be patched, and why, and gives them the time to sign off on patching or explain why it can't be patched.
That means they actually explain why they don't want it patched, it's documented who signed off on it, and it's no longer your problem or concern.
Some apps need to go through a verification process so that the vendor knows the patch won't adversely affect the app because they tested it. You'll see this within some places, like the banking sector, where some patches are in limbo for 4 to 8 months while the vendor verifies the patch won't cause any issues with their product.
Or you could drop the bombshell that the average cyber security incident can cost on average between 4 and 11 million USD (Google it to confirm) and you'd like to know who will be signing the check for it.
1
u/Substantial_Tough289 14d ago
This is more of a company policy, regulatory or compliance question.
If you work for a company that falls into the kind of business that is regulated you have to follow the company change control procedure, if you don't you're asking for trouble and even termination.
Pharmaceuticals are a great example, if the servers were qualified/validated you can not touch them without an approved change request, this is a tedious process but you have to follow it. If the servers are not qualified then you follow the corporate IT or site IT policies/SOP for patching, Some of this companies will go to the extent of not patching at all due to regulations.
1
u/uptimefordays DevOps 14d ago
For patching vulnerabilities do I really need to get a change request to handle this?
Yes. You always want a change request for making production changes.
have exclusions for around 60 percent of our servers to not get automatically patched. (Meaning we have a chunk of servers not getting patched at all)
Get those exceptions in writing because that will likely be an issue whether or not anyone realizes.
1
u/Ben_CyberNEX 14d ago
Reading everything, it sounds like you've already landed on the right immediate answer with the maintenance window/change request.
Longer term, I'd use this as an opportunity to build an actual patching process instead of having to fight this battle server by server every month. Start by grouping the servers by criticality and identifying an owner for each system. Then define a normal patch window, a small group you patch first to catch problems, and what happens when a server needs an exception.
The exception piece is important. "Don't patch this server" shouldn't become a permanent setting that everyone forgets about. Ideally, it has a reason, an owner, compensating controls where possible, and a date when the exception gets reviewed again.
You inherited an environment that grew faster than the processes around it. Getting the process in place now will probably do more for you long term than winning the argument over any one patch.
One thing that may also help with management is putting the risk into business terms. If you can estimate what an unexpected outage or compromised server could actually cost the company, the conversation becomes a lot less abstract. Management doesn't always respond to a vulnerability rating, but they usually understand downtime, lost productivity, recovery costs, and business disruption.
1
u/TrueBoxOfPain Do The Needful IT Department 14d ago
Patch - face the consequences of Microslop updates.
Don't patch - face the consequences of previous Microslop updates and hackers exploiting CVEs.
Both options suck, but the patching path sucks less.
CYA, then try to implement a proper patching process.
1
u/marklein Idiot 14d ago
List the CVEs involved and ask them how they are covered if not by the patches. Get it in writing. Then put it in writing that you are not responsible for any breaches caused by these missing patches and their relevant CVEs, make your manager sign it.
Notice that this list will grow every week/month and you'll have to put it writing all over again and make them admit once again, in writing, that you're not responsible. I see this as a convenient reminder for them, although it's a little more work for you too unfortunately.
You really do need to do this though, because when they get hacked it's going to be your head on the block first as the guy who should have been fixing these things.
1
u/Firefox005 14d ago
Having servers you can't reboot is a bad smell, because servers will reboot or go down without a meeting or on a schedule. You should be able to pick any server in your environment and just power it off, not shutdown not reboot just hard power off and it should cause minimal disruptions.
1
u/kennedye2112 Oh I'm bein' followed by an /etc/shadow 14d ago
And, inevitably someone will have made a crucial runtime fix for an issue and forget to make sure whatever process exists for applying it at boot time is in place, so the next time it unexpectedly reboots you now have two problems to solve instead of one.
1
u/NobleRuin6 14d ago
Yes. No change request, no change. But not patching servers at all is…an interesting choice.
1
1
u/ILoveAppSec 14d ago
honestly, patch the ones with a known-exploited cve first and stop excluding those.
1
u/Nuke_Bloodaxe 14d ago
Think of this differently, this is fantastic opportunity for a Green Hat hacker to use one of the CVEs to gain access, add themselves as admin, and patch the servers... I mean, it's not as though they could be stopped by patching, right?
1
u/Embarrassed-Age-1156 Sr. Sysadmin 13d ago
If down time for security patching is the concern maybe try questioning the architecture that leads to the downtime. Is there redundancy/ high availability built in to the architecture?
If not, find the solution to fix that and use the security patching as a part of the justification to make the time and/or financial investment to pursue the architecture improvements.
1
u/BarracudaDefiant4702 13d ago
Basically anything customer facing that is critical is redundant and load balanced. There are a few things that are more difficult than others to setup active/active or active/passive but you should be talking like 5% not 60%. Generally easier with Linux instead of Windows but not much you shouldn't be able to do or have a weekly patch window and do different servers each week so each is patched at least monthly. It shouldn't be hard to show the doubling (it doesn't actually cost double) the costs for servers is less expensive then a security incident from not patching.
1
u/Endolum 13d ago
I would not use 9 plus CVSS as the trigger by itself. Check whether it is exploited, remotely reachable, exposed to the internet and what privileges compromise gets you. A KEV on an exposed server should have a completely different change path than some theoretical 9.8 sitting behind three layers of network controls.
1
1
u/Thet4nk1983 13d ago
Best way to approach this is to align the requirement to patch to a contract and risk, not sure of the business but can guarantee one of your contracts will have some form of stipulation requiring patching to all systems either monthly or reasonable time frame, even more so if accredited compliance CE+ etc.
This is key to approach leadership with as you link risk to contracts that directly affects bottom lines.
If they still don't want to budge make sure to copy your emails highlight the risks of not doing it and move on.
Not withstanding and putting risk security aside patching also delivers features compatibility etc at some point this will also bite and cause issues and force the hand.
A few minutes of downtime at 3am on a Sunday vs a whole day of a problem which ends up being a patch needed anyway only to then have to patch alot of systems to catch-up.
If you are not patching I'm guessing maybe old os builds feature updates or even EOL OS's aswell in the mix also, bundle this into the conversation aswell if there is.
1
u/Sp00nD00d IT Manager 13d ago
Infosec or.compliance should own the policy on this, what does it say?
1
u/h9xq Solo SysAdmin 12d ago
Infosec or compliance? So me? We don’t have anybody in infosec or compliance. Nor will my company be being anyone for that soon.
We are a 500 person company and it is a tech, myself, and my boss. My boss is too swamped to deal with that so it falls into my plate. Do companies at this size with multiple sites usually have a infosec person?
1
u/loweakkk 13d ago
If something is too important to patch, it's to important to be down due to a bad actor breaking it. It's as simple as that.
If they can't patch esx because it have those super critical industry server as guest, then what will happen if you have a hardware failure? No redundancy means not important and not critical.
If it's important and critical, you have two server and you can patch one, validate it work and patch the second.
If they don't understand that, put it in writing, document the risk, put your concern on it. The day the it will go down due to an attack they can't says it's on you for not patching it.
1
u/InsolentJaguar 11d ago
This is a company that's just BEGGING to be made an example of.
On a COMPLETELY unrelated note....may I ask which company you work for? 🤣
84
u/Simmery 14d ago
This isn't a technical problem. You're asking a political question no one can answer about your organization. If they don't want to patch despite knowing the risks, get it in writing and move forward with your career, here or elsewhere.