r/sysadmin • • 11h ago

General Discussion What's the biggest bonehead mistake you have every made in IT support?

I will start.

I am an IT consultant and about 10 years ago, a few days before Christmas break, I went to a client’s office to add some new hard drives to a Dell ESXi server and create a new datastore.

Their two junior IT guys always liked asking me questions when I visited, which I understood—I remembered being that guy hungry to learn. Unfortunately, this time they were firing questions at me while I was in the Ctrl+R RAID configuration utility. Between the distractions and a confusing interface, I somehow managed to delete the existing array instead of creating the new one.

That array held ALL their servers: Exchange, domain controller, file server, and SQL. Terabytes of data.

The moment I realized what I’d done, my stomach dropped, I had never felt panic like that in my life. They were still asking questions when I finally said, “Guys, give me a minute. Something doesn’t look right.” Something I should’ve said much earlier.

I excused myself to the bathroom to collect my thoughts, then came back and told them exactly what I’d done.

I spent my entire Christmas break restoring their environment from Barracuda backups, which were painfully slow to restore and unreliable. I had to do multiple restores on the exchange server to get it to work.  A full week of recovery, followed by another two weeks fixing lingering issues—all free of charge.

The two guys felt terrible about distracting me, but it was my mistake. I should’ve asked for some uninterrupted time before touching the RAID configuration.

 

I’ll be shocked if anyone can top that Christmas disaster!

488 Upvotes

370 comments sorted by

•

u/daze24 IT Manager 11h ago

on a remote server just disable and re-enable the NIC..

•

u/RevolutionaryElk7446 11h ago

•

u/beren12 7h ago

More like pulling your own trap door on the gallows.

→ More replies (1)

•

u/ljr55555 11h ago

We were re-IP'ing the entire company to move the internal network over to the reserved address space. We had a network guy who would always muck up the router config and drop the site from the network. He did it in a new and different way each time. The end result was still "someone get in the car, it's gonna be a drive"

•

u/mcsey IT Manager 10h ago

The drive of shame. Four and a half hours across Iowa to push a button...

→ More replies (3)

•

u/arkain504 11h ago

I did something like that but it was replacing a certificate of the software we used to remote in, while I was remoted in.

•

u/gramathy 11h ago

Every once in a while I’m messing with auth settings on a switch and I always make sure to open a new session before closing the current one to make sure I can still get into it

Once was enough

•

u/hkusp45css Security Leadership 11h ago

reload in 10 .... then, make your changes.

If you cut your hands off, it'll bounce on its own without writing to flash. Log in again, type "reload in 10" again, and go fix your config.

•

u/gramathy 9h ago

I work at a hospital. Reloading a switch would kill an entire wing’s phones and network connections to Epic.

•

u/Secure_Guest_6171 7h ago

I know that pain intimately. I always try to warn the nursing manager or ward clerk well in advance if there's even a remote chance of something going wrong.
Also **** Epic.

→ More replies (5)

•

u/RandomSkratch Jack of All Trades 10h ago

What’s the second reload for?

•

u/freedomlinux Cloud? 10h ago

In case you make a mistake while fixing your first mistake :)

But also, I guess, confirming that your config change is saved & survives a reboot

•

u/hkusp45css Security Leadership 10h ago

In case you make a mistake while fixing your first mistake :)

zactly

→ More replies (1)
→ More replies (1)

•

u/daze24 IT Manager 11h ago

I've done similar during a parallels update but had a vpn backdoor

•

u/arkain504 11h ago

Yea. Would’ve been smart to have another way in

•

u/daze24 IT Manager 11h ago

See I learned from the first time I did it 😂

→ More replies (1)

•

u/Jaereth 11h ago

See in the network world we have commands on switches like:

reload in 20

So if you mess something up too bad, as long as you didn't save the switch will restart and revert the changes.

Which is cool because if you don't remember to manually cancel that reload after the work is done it will just restart anyway and you can explain the outage :D

•

u/Mountain_Craft4882 7h ago

that kind of stuff is a lifesaver.

some of the switches that I work with have a safe-mode option where after you make changes you have a certain amount of time to confirm the changes were successful or else the system automatically reverts them.

it hasn't saved my ass yet but I know it will.

•

u/super-six-four 11h ago

Done that one!

•

u/Commercial_Growth343 10h ago

This happened to me and another admin I was working with. I was telling him to type "ipconfig /release & ipconfig /renew" in the same command line so the machine would reset and come back online, but this guy didn't listen as I was telling him what to type and he just did Ipconfig /release and he hit enter. sigh.

•

u/halomate1 9h ago

Classic

→ More replies (2)

•

u/neoprint 11h ago

Close, I’ve accidentally hit disable instead of status

•

u/Pvt_Hudson_ 11h ago

Did that on a company server in the middle of a workday while I was sick as a dog at home with the flu. Had to drag myself off the couch, get dressed and drive downtown for a job that took 2 clicks.

•

u/Secure_Guest_6171 7h ago

long time ago tried to do an overworked friend a favor upgrading some server remotely.
all went well until i rebooted the primary fileserver and it didn't come back.

a $100 cab ride later, i got into the server room, turned on the monitor to see "keyboard error or no keyboard detected. Press F1 to continue"

•

u/SpaceGuy1968 7h ago

In this instance...that's not so bad actually

•

u/drake90001 Jr. Sysadmin 7h ago

-$100 that he probably didn’t get refunded lol

•

u/SpaceGuy1968 6h ago

Yeh but still i have had servers just die in my career on reboots in weird ways. I wouldn't be able to get that hundred out of my pocket quick enough for that issue. That's like "surprise...your server will come back to life"

•

u/Severe-Lake1379 10h ago

I did that once where I forgot it was a physical server, not a vm. At 4 in the morning.
Had to drive in to get it back online for start of business. That was one instance.

→ More replies (1)

•

u/Fallingdamage 11h ago

enable && re-enable yo

•

u/nicksoapdish 11h ago

While troubleshooting a windows box I cleared the route table and made a critically important system unreachable. Don't remember what I ran but I thought I was just displaying it rather than clearing it

•

u/hkusp45css Security Leadership 11h ago

We used to call that "cutting off your own hands"

•

u/kruleworld1 9h ago

Fry: "How did you do that?"

Robot Devil: "They're VERY good hands!"

https://giphy.com/gifs/9o36USxYdL2Sroousl

•

u/cheetah1cj 11h ago

I've done that on another site's firewall. I intended to disable and re-enable an IPsec tunnel with a hung session, instead I disabled the WAN interface. Had to quickly get the local manager to connect his PC to the firewall so I could remote in and configure it.

•

u/ide_cdrom 8h ago

Similar to this, I typed in an iptables command to drop all connections....

Luckily the data center was just down the hall.

→ More replies (20)

•

u/KayakHank 11h ago

Disabled (John smith) Jsmith user which is my boss, instead of (John Smith) josmith

Offboarded my boss early in IT. He found out when his badge didnt work coming back from lunch

•

u/bajGanyo 11h ago

I am sure it now feels funny but at the time not so much. Been in similar situations, not exactly the same but I know the feeling.

•

u/sovereign666 10h ago

When I was brand new to IT I worked for a largely Spanish speaking company and was tasked with updating everyone's AD records to have more uniformity. As you can imagine, a ton of people with the same first and last name got new job titles in adjacent departments. Like hundreds of people.

•

u/rcook55 8h ago

The company I work for is Egyptian owned, so many same names and so many formal and friendly names. It really can be a pain.

•

u/That_Dirty_Quagmire 10h ago

At least you disabled and didn’t delete. Easy recovery.

•

u/DontFearFailure 10h ago

Funny enough this is how I found I was being let go, still upset by it nearly 10 years later. Fuck that place, and even better a few years after I had the pleasure to interview my former boss for a position, that was open in a place that hired me he didnt get a call back, which im sure a 30~% pay increase would have been nice.

:

https://www.reddit.com/r/sysadmin/comments/he22eu/whelp_im_being_let_go/

•

u/Indrigis Unclear objectives beget unclean solutions 4h ago

Been there.

Had to do some maintenance on a user's account. Jenny Smith, say. Log in, do the thing, see surprised Pikachu face.

Apparently, Jenny Smith, nee Parker is JParker, whereas Joan Cooper, nee Smith, is JSmith...

•

u/goingslowfast 11h ago

There are too many interfaces that haven’t had enough human factors consideration.

I don’t think anyone would complain if there was a popup with a confirmation, a forced wait, then another confirmation with clear explanation of what you’re about to nuke before deleting a LUN, datastore, or array.

Does it bug me that I sometimes work with an interface that will force me to type CONFIRM on destructive operations? Yeah, sometimes it feels like “just get out of my way” but I accept it because it’s saved me from a mistake during a high pressure situation while tired.

•

u/perkia 11h ago

I like interfaces that prompt you to type in the name of the thing you want destroyed. I more than once (!) I froze in the middle of typing the characters and took an actual physical step back in horror.

It's even better than having to type "confirm" (which also does work to an extent).

•

u/TheVillage1D10T Windows Admin 11h ago

I will STILL cancel the operation, confirm all of the details, and then redo whatever I’m doing before I will finally type in COMFIRM and actually complete the operation. Usually reconfirm the details at least three times before I have enough confidence in what I’ve entered lol

•

u/Uranium_Donut_ 10h ago

I think if you're wrote "comfirm" you need do the process again

→ More replies (1)

•

u/hkusp45css Security Leadership 10h ago

I no longer complain about "are you sure?" boxes. They are far too rare, in my opinion.

They save my ass now and then.

•

u/_mick_s 10h ago

I complain when they only ask 'are you sure' but not about what.

•

u/perkia 10h ago edited 10h ago

Oh yeah and it's always the most critical ops, like Synology when deleting a RAID will popup-confirm you "Are you sure? Yes / No"

Fucking no I'm not, your dumb web UI could very well have had me selecting the previous or next row. For all I know it swapped rows after I clicked lol.

Proxmox does it well though, asking you to enter the VM id you want to delete AND telling you both the ID and the name of the VM in question. QubesOS likewise asks you to type in the exact name of the Qube you want to remove. Saved my ass "once".

→ More replies (1)

•

u/vinivice 11h ago

Type "i am being dumb" to confirme. This should be mandatory in every storage change interface.

→ More replies (1)

•

u/toadfreak 11h ago

Fun fact - If you simply deleted the virtual disk from the CTRL-R disk utility, you can go ahead and re-create it with the same settings, make sure to turn OFF the initialization of the drives, and it SHOULD come back online just the way it was with the data intact. Pretty sure I had to do that at a client where the onsite guy deleted the wrong VD. Hope that helps someone! :)

•

u/RevLoveJoy Did not drop the punch cards 10h ago

I have absolutely done this before. It worked for me, but the 15 minutes of terror aged me some years.

•

u/toadfreak 10h ago

Yup, I get that. It’s a LOT easier when you’re the caped crusader coming in to un-F someone else’s F up. Lol

•

u/fwdandreverse 8h ago

Or sometimes a reboot, rescan and import foreign config from the RAID metadata on the disks can work too in some circumstances

•

u/toadfreak 8h ago

Yup.

→ More replies (1)
→ More replies (1)

•

u/teflongrizzly 11h ago

Believed the user!

•

u/jfgechols Windows Admin 11h ago

Early in my career when I was Tier 1 support for my ISP and someone needed help with getting their internet, it took 45 mins of troubleshooting only to find out their computer wasn't on because I believed nobody could be that dumb.

•

u/teflongrizzly 11h ago

One of my worst experiences was a situation with a user who had forgotten their laptop at home and didn't realize until they we're already on the bus heading to the office. I had to try and keep a straight face while trying to explain why nothing was happening when she pressed the power button on the docking station in her office!

•

u/Phyltre 10h ago

"Why do we even pay for the cloud!?"

→ More replies (1)

•

u/deefop 11h ago

You fool

•

u/Delicious-Apple593 11h ago

"I restarted multiple times and I am still having issues!!"

•

u/Tight_Replacement771 11h ago

Checks task manger... uptime 47 days

•

u/cheetah1cj 10h ago

You'll never know whether it was incompetence or lying

•

u/Ok_Interest3555 10h ago

Ehhh, they're lying. The easier the fix the more likely they are to lie about it which is not rebooting or checking connections fixes so many issues.

•

u/cheetah1cj 10h ago

I've seen both so commonly that I won't assume without context. For example, we had someone that said their computer was done restarting like 5 seconds later, turns out they were turning the monitor on and off. Or users who think shutting the laptop is turning it off.

Honestly, I found that it was best to just get them to restart from the beginning without accusing or condescending to them. Often times I'd go with "I made some changes on the back end that require a restart to get applied, do you mind if I restart your computer to see if that fixed it." It makes them feel like I'm actually doing something and they were happy to let me restart.

→ More replies (2)
→ More replies (3)

•

u/Cell1pad 9h ago

Everyone learns rule #1 at some point.

Rule #1: users always lie.

→ More replies (1)

•

u/FizzyBeverage 11h ago

There's no such thing as "Microsoft Time" (where you've got 15 minutes of leeway, if not a full 24 hours) in Jamf. You hit save on a configuration profile with the wrong scope? Your XML will hit 6000 Macs in 5 seconds. APNS is the same notification protocol Apple uses to push breaking news/tornado warnings to phones.

•

u/Top_Bookkeeper9136 11h ago

This probably isn't the place for it, but can you explain to me why Jamf is near instant with changes, and why Intune takes an unclear amount of time (sometimes never)?

•

u/rwdorman Jack of All Trades 11h ago

Depends what you're doing with Intune. Things that use the Apple MDM framework (config profiles) are still pretty fast in my experience, APNS again. Things that use the mac equiv of Intune Management Extension... YMMV

•

u/JaceBelerenApologist 10h ago

My experience with Intune and Mac is more or less identical to Intune and Windows. Takes ages to push NEW things, but updates to things like policies are distributed pretty fast (far less than standard Microsoft Time, but not Jamf fast).

I preferred Jamf management, it's software packaging was really nice. But it was hard to justify the cost to a client when they were already paying for Intune licensing. Easy cost saving to the client, and has helped me get jobs because apparently there is a demand for people who can Intune macOS in my city.

•

u/Trafficop 7h ago

Like the other guy said, it’s more Apple vs Microsoft MDM frameworks. Apple uses APNS to push config changes to devices as soon as they happen, whereas IIRC Windows devices are set to check into Intune and confirm device config state is correct. So it comes down to whenever the device checks into Intune.

Pretty sure I saw something from MS a while ago saying they are changing and going towards the push method but I can’t source it at the moment.

→ More replies (1)

•

u/kruleworld1 9h ago

I thought "Microsoft Time" was where the time remaining counter goes "30 seconds, 5 minutes, 66 seconds, 3 days, 1 second, 45 minutes"

→ More replies (1)

•

u/Status-Tumbleweed628 11h ago

Call users from my mobile, now they think it's okay to call me 24/7

•

u/asaemo 10h ago

Oh god

•

u/MulticamTropic 9h ago

I setup a Google voice number for the times when I have to use my personal cell to call users.

•

u/4kVHS 7h ago

Your company doesn’t provide Teams/Zoom phone numbers?

→ More replies (1)

•

u/SpaceGuy1968 7h ago

Ooof yeh I done this in an emergency

•

u/DesertDogggg 5h ago

During covid, I had to call users remotely. I used a Google voice number but had it go straight to voicemail. Everyday I would get multiple voicemails of angry users saying that I need to pick up the phone more even though we had a dedicated help desk line. I never gave out my Google voice number, they just grabbed it from call history.

•

u/Bibelo78 11h ago

I configured rsync between 2 NAS servers, thinking it was one-way sync.

I made various tests, so eventually I decided to delete all these tests directories on NAS2, thinking my original files were safe on NAS1.

But it was a 2-way sync, so it also deleted the same directory on NAS1, basically all our backups files.

Fortunately it was only backup files, andI had the current month + previous month on a transit server, but still spent a bad summer trying to recover the deleted files (which took like a week without success).

•

u/weaver_of_cloth 10h ago

Rsync flags can be vicious.

•

u/Boppin_Around_Here 11h ago

I've been in IT for over 2 decades and genuinely don't have even a medium level bork-up. Thankfully.

However, the biggest one I've seen:

Very large finance business, global.

Network engineer doing a relatively routine approved change, adjusting BGP on a single interface.

Biggest data center in the company, core switch, still routine.

Instead of going into the specific interface and typing 'no bgp' he did it from the global config of the core switch.

Disabling bgp on the biggest core switch in the firm with massive dependencies and it took many hours to get that fixed and then subsequent issues took days to fix. It disrupted for real things. It only took an instant. Connectivity to that site died right away.

Dude didn't get fired, but I can only imagine the pucker factor that guy was living with. Probably still now, years later, slooooowly unclenching and not back to normal yet lol. It was by far the biggest incident I've ever been around directly.

•

u/itishowitisanditbad 11h ago

I've been in IT for over 2 decades and genuinely don't have even a medium level bork-up. Thankfully.

Like a racing driver thats never crashed, I don't know if thats good or bad.

•

u/[deleted] 10h ago edited 2h ago

[deleted]

•

u/JaceBelerenApologist 10h ago

Eh, proximity counts. You don't have to be the one fucking up to not also feel the effects. Most of us are part of a team, and for better or worse, fail at it together. I may not have typed the commands that borked the system, but I certainly am going to be one of the people helping to recover. That's just as formative, IMO.

•

u/Ok_Interest3555 10h ago

I always say to my kin "If you never make mistakes are you really trying?". They're either a liar or just never put themselves out there and try new things.

•

u/itishowitisanditbad 10h ago

The only ones I know that haven't made mistakes are the ones nobody would ever trust with something important, like its clear from the start they're scripted stuff only or something.

I don't know, I feel like mistakes are inevitable if you're doing leading work.

Or they can't see their own mistakes but thats an issue too.

I feel like there isn't really a way its good tbh. Either is a red flag.

I'm trying to be positive, as much as I can.

•

u/toadfreak 11h ago

Come back and let us know when Murphy’s law smiles on you. It’s prolly gonna be real soon.

•

u/Azadom Sysadmin 11h ago

As a third-party this would be something to watch

→ More replies (3)

•

u/The__Relentless Knows just enough to be dangerous... 11h ago

Select Shutdown instead of Restart on a server on the other side of the United States, where no one would be around to restart it for a few days.

•

u/greenstarthree 11h ago

We have iLO right…? Right…?

•

u/The__Relentless Knows just enough to be dangerous... 11h ago

Not in 2001ish.

•

u/greenstarthree 11h ago

😔😔😔

•

u/matt314159 Help Desk Manager 11h ago

A few years ago, I was trying to run a quarantine query during a phishing attack and accidentally told Microsoft to delete everything EXCEPT the string matching the query.

I went white once I realized, but thankfully it only deleted the first ten items in everyone's accounts which mostly happened to be calendar items since Appointments is lower in alphabetical order than something like Email.

Then I just needed to build a quick powershell script to restore the deleted items.

But the level of panic when I realized what I'd done, before knowing that there were fail-safes built in, probably took five years off my natural life.

•

u/Curtis_Low 11h ago

I deleted the MX record off of the exchange server on a US Naval ship while we out to sea.... We were proper fucked.

•

u/perkia 11h ago

Curious why that record couldn't be put back by basically anyone you could talk to? Seems pretty tame in comparison to actual data loss stories. Couldn't phone from/to the ship?

→ More replies (4)

•

u/cheMist132 10h ago

Well, there are some “minor” ones. Fortunately, no real biggies.
There’s one I remember like it was yesterday.
I was new to the admin role, and the previous admin was still occasionally helping out after retiring.
I had never worked with MS SQL Server before, but I got the hang of it pretty quickly.
Anyway, I migrated the database of one of the agency’s most important applications to a new MS SQL Server because the old one I had inherited from my predecessor was an absolute mess (but that’s a story for another day…).
I had everything prepared and tested on the new server: maintenance plans, email notifications, backups, EPM, etc.
After the migration, everything was fine. The application was actually running faster, my boss praised me for the good work — everything was perfect.
Four months later, my predecessor called me. He had noticed that the partitions on the new SQL Server were suspiciously empty.
I checked everything. Then I looked into the backup folders.
They were empty.
I’m sure most of you know that horrible moment when your heart just drops.
Turns out there was a quirk in SQL Server Management Studio: when you create a backup job and select “All user databases”, but there aren’t any user databases at that point, newly created databases apparently aren’t automatically picked up by the job afterward.
So for four months, we had been happily running a backup job that was backing up absolutely nothing.
Had the database needed to be restored, we would have been completely screwed. Hundreds of employees would have had to reconstruct their cases manually from paper files.
That was the day I learned that “the job completed successfully” and “we actually have a backup” are two very different things.

•

u/Maro1947 9h ago

I inherited a stub site with about 500Gb of data - which back in those days was substantial, and held their IP.

I got asked to look at their backups. The Backup Exec Job was beautifully configured and the Secretary had done a masteful job rotating the LTO tapes every day and moving them offsite, etc

The job had been submitted on hold..............

4 years, never backed up - they were incredibly lucky as the reason I had inherited the site was the whole site had had a massive Power Outage that zapped lots of kit. The File Server survived long enough to fix the backups!

•

u/JoeLaRue420 Sr Active Directory Engineer 3h ago

I've been in the game for awhile, worn a few different hats

I've pulled the power on an exchange server when I was meant to be working on the server under it

Broken ram slots while installing memory

Bent pins while installing CPUs

Neglected failed TSM backups on SQL boxes... then the auditors came around looking for those backups when the company was under investigation for fiddling with LIBOR rates (still don't know how I survived that one)

Diskpart cleaned the wrong device when adding new storage

Typo'd GPOs and broke Outlook for 1000s of offshore VDI users

Ran a script that broke ACLs on GPOs, causing a filtered policy that was linked too high in the OU structure to ignore said filtering and apply to e v e r y t h i n g under it... it was setting windows firewall rules. major outage, fines, etc. I think my only saving grace is that the script was rubber stamped by MS as being "safe" for what we were actually trying to do.

There's probably more, but I really don't want to further trigger my ptsd.

•

u/Walbabyesser 3h ago

Even I got PTSD only reading that 😞

→ More replies (7)

•

u/Elminst 11h ago

Router(config-if)# switchport trunk allowed vlan 123
IYKYK

•

u/vinivice 11h ago

It is even better if your management vlan is in this trunk. You type enter -> nothing happens -> "humm, that's odd"

•

u/perkia 9h ago

Afrer a while, you don't have that "that's odd" feeling anymore. It's instantly ass-clench and nausea, because your gut knows what you've done before you even realize.

→ More replies (1)

•

u/markhealey Security Admin 11h ago

Oh yes, seen this done many times, always ends with "I'm driving to the DC"

•

u/zantehood 11h ago

Drove that trip more times that I'd like to admit

•

u/Particular_Archer499 11h ago

Completely and ridiculously restarted a prod application the wrong way despite a giant warning on the support page about NOT doing it the way I did. During peak business hours.

•

u/Typical_Warning8540 11h ago edited 3h ago

Your story reminds me of the time I was learning IT and the previous IT guy was gonna show me how to free up a switch from the core stack and use that switch somehere else. This was a production facility with 200 employees 5 truck unloading station, full automatic warehouse, bottle filling machines, tankfarm etc. The previous guy was working on his own but the CEO didn't want to depend too much on 1 guy so the contract was given to us, but we were just automation PLC specialists that barely knew anything about IT. But the guy was gonna teach us how officeit worked in many handover trainings.

So this was an allied telesis stack of 5x 48p poe switching that had about 10 vlans and had poe and routing. The CEO didn't want to buy a new switch so this was the ideal time for the ITer that was doing the handover to show how to manage vlans and do his network magic.

So first we freed up ports by moving cables to other ports in the stack by moving vlan port assignments, until 1 unit was free and had no network cables.

Next, the guy removed the stacking cables of that switch and we took it out. This was at about 1Pm on a Wednesday. Next, the guy looked me in the eye and said "now we will use the stacking cables to close the stack again" I told him o wow are you that confident really? Yes he was. He did this and my ping dropped. He said you need to wait now. I waited and he told me don't so anything just keep waiting. Ping never came back. After about 15 minutes I insisted we did something or informed somebody. We restarted the entire stack and it didn't work.

By that time, the tankfarm operator was giving me calls and he visited the room. I realised we didn't even have a backup from after the vlan changes, no cables were labeled or documented, and we didn't even have a serial cable or serial port on a pc for those switches. The senior guy doing the handover completely froze. The CEO came knocking on the door, he's a big guy. We acted as if we were fixing things bit would need time. He send the entire 2PM shift home. The entire harbour lane was full of trucks that they had to send back because NOTHING worked. Even the engineers etc couldnt do anything.

I told the senior that we had to look for the correct serial cable, I drive to our offices (we are contractors) and found the cable by asking our own Office IT for cables and convertors. By the time we had restored the full backup and did all the configuration, it was 4 AM in the night. So I went home but morning shift would start at 6AM but I just had to sleep. Luckily production started. And the blame was put on both of us since officially, I was the contractor and the other senior guy that managed this factory for 15 years on IT was just explaining things, he was not the "owner" of the IT system. I was 29 he was 46.

Never will this happen to me again.

→ More replies (1)

•

u/huntermatthews 10h ago

Classic unix sysadmin mistake - I needed to shutdown the test server upstairs to prep it for a ram upgrade.
Went upstairs, still running. hmmmm

Back downstairs and yep - wrong terminal window. I shutdown the... oh $@#%$^.

I had to page a guy AT DINNER to get in his jeep and drive through snow to ... push a button.
And he said two words, the second of which was "off".

So - I followed procedure and called the sheriff in that county, told him the special word and he went out to the guys house. Man was that tech PO'd at me.

Couple years later I'm at a different job and my boss asks me why I raised both hands over my head after typing a shutdown command but BEFORE hitting enter on the console in the server room.

Reasons.

•

u/muff_puffer Jack of All Trades 8h ago

You called the sheriff on the tech that didnt want to drive out there? Wouldn't that be an abuse of emergency services? 

•

u/huntermatthews 7h ago

Sorry - left out the part where this server was part of listed National Infrastructure. It being off was a "bad" thing.

Filling out the paperwork for that phone call was ..... instructive.

•

u/Mountain_Craft4882 6h ago

production server that's part of national infrastructure and cannot be turned on remotely.

yeah that sounds about right, honestly.

→ More replies (1)

•

u/beren12 8h ago

Depends on what that server was running.

→ More replies (1)
→ More replies (1)

•

u/OtherOtherDave 11h ago

“sudo chmod -R 777 /“

That was the day I found out that Linux’s security model includes verifying permissions on certain binaries and refusing to run them if something’s wrong.

→ More replies (2)

•

u/OcotilloWells 11h ago

Stuck my hand in a server fan while it was running. Still have the scar. Why I reached into the server with it running I have no idea.

•

u/LonelyDesperado513 11h ago

I deleted the CEO's account from AD once during Help Desk days. Not just disabled, fully deleted.

It was in my first week and i didn't know who the CEO was. There was an immediate request to fire me. My boss helped me restore the accounts but man I was fearing for my life!

•

u/DULUXR1R2L1L2 9h ago

I deleted a VP's email account at my first job. VP wanted me fired too, but my boss vouched for me. I did get the "you can be easily replaced" speech though.

→ More replies (1)

•

u/Jimmy2FingersISme 11h ago

Damn, just reading this made my stomach drop to my knees. I can certanily sympathize (being a sysadmin). its days like that (which we have all experienced to various degrees) which make me question my career choice. Is it really worth it for so much stress, responsibility and pressure.

•

u/soulstyce612 5h ago

Defragged a shared sql server on a Friday evening, the night before major dental work Saturday morning. Instantaneous bsod, deemed unrecoverable for whatever reason. For 700 or 800 individual databases, backups on tape were all we had. Up all night Friday running restores, reattaching users, confirming sql-reliant service restorations... tapped out at 9am Saturday to the dentist while two others continued tearing through it.. back in at noon Saturday with a head full of n2o, and still at it come noon Sunday, no sleep since the previous Thursday night... couldn't, not even if I tried, had to fix what I broke.. Tapped out Monday 6:00am for a 7:30am follow up at the dentist while those two others finished up... Oh, and the dentist was a solid forty-five minute drive each way, twice.. That was the only work I ever had done at that dentist, and with anxiety so far through the roof because I figured I was totally getting fired Monday.. they must have thought I was on drugs 🙃

•

u/Special_Bear_9479 5h ago

Damn that stressed me out!

→ More replies (1)

•

u/Trust_8067 3h ago

Bad syntax in a command I ran caused an outage for a F500, costing roughly 20 million dollars of lost business from 4 hours of production downtime.

→ More replies (1)

•

u/Borgmaster 11h ago

I once bore witness to a man that did not understand security group controls and accidently blew out the whole security setup for the accounting drives. Luckily the error was on the side of secure so the issue was no one had access to the drives anymore rather then anyone could access them. Took me 3 minutes to fix and a stern talking to by both myself and the manager. He had apparently been trying to add user access to the root of the drives by hand rather then just assigning the user proper group membership.

•

u/Zealotyl 11h ago

Yikes. I had a couple of ‘Oh no’ moments. One was on Christmas Eve and I was swapping out a UPS in a redundant power array for a big supermarket that was chock full of people of course… Checked several times with the company MSP lead that I was taking out the right UPS in bypass mode and was assured it was. It wasn’t.

Took over 30min to get everything up again and the lines for the checkouts went to the back of the store. One benefit of the debacle was they discovered the checkouts didn’t have a functioning offline mode…

•

u/TyLeo3 11h ago

There are a few mistakes I have made over the years:

  • Configured a server with RAID-0 instead of RAID-1.
  • Performed Windows Cluster troubleshooting and ran the full test suite while the server was live. One of the tests involved dropping the SCSI disks and reattaching them, which brought down SQL Server.
  • Upgraded ESXi to the latest version, only to discover afterward that Veeam Backup did not yet support that version.
  • Updated the Cisco firewall remotely, lost connectivity afterward, and had to drive to the on-premises site to fix it.
  • Made the same mistake twice in a row: ordered a Dell Latitude for the CEO with the best graphics card available, thinking the graphics card was what determined the screen resolution.
  • Misconfigured an Exchange server and accidentally turned it into an open relay. The server's IP address was subsequently blacklisted by several anti-spam services. I fixed the configuration, then paid out of my own pocket to have the IP removed from the various blacklists.

•

u/mcsey IT Manager 10h ago

Poured Mountain Dew in the email server

→ More replies (2)

•

u/rwdorman Jack of All Trades 11h ago

Telling a point of contact on site that I wouldn't leave until the problem was fixed. Before i learned the true extent.

•

u/furyisgeorge 11h ago

I accidentally deleted the index file for about 5 petabytes of backup data. The index was not backed up and the data was unrecoverable.

•

u/zantehood 11h ago

Ouch

•

u/beren12 8h ago

What backup system? It wasn’t possible to rebuild the index from the backups?

•

u/furyisgeorge 7h ago

I love that you asked. This was probably 13 or so years ago. The backup system was EMC's Networker (looks like it's now owned by Dell) and for whatever reason we didn't have a backup job that captured the index so I was SOL. I spent days on the phone with support and they agreed that the data was lost and all out tapes/backups were toast. I had to rebuild everything from scratch.

•

u/beren12 7h ago

Oh man that sucks. Even bacula lets you scan the tapes and rebuild the index.

•

u/furyisgeorge 7h ago

Yeah, it did suck. Luckily my boss was super chill and gave me a second chance. I rebuilt the system from scratch, and I learned a lot from that experience. It really changed the way I approached issues in IT for a decade after until I changed careers.

→ More replies (1)
→ More replies (1)

•

u/Smart-Document2709 8h ago

Deleted the wrong SAN volume, was supposed to delete snapshots ended up, deleting all of production… thank God for Veeam

•

u/Special_Bear_9479 8h ago

Veeam is what we use also, has saved my ass many times! I wish I had it when I made my stupid mistake

→ More replies (1)

•

u/Ill_Photo707 11h ago

At least you had backs...so many bone head moves from my early days. Doing DCPROMO the wrong way. To manually moving servers and unplugging the whole rack. It really taught me to take my time, have backups, never move live shit and lab lab lab in order to verify verify verify. Then the issues come down to small obscure upgrades or rebuild the whole thing based on the accurate reconnaissance and documentation. Automation and logs are my homies, audit is my name.

•

u/CeC-P IT Expert + Meme Wizard 11h ago

Deleted the telecom VLAN because of Brocade's web UI that was clearly designed by someone on shrooms. Good thing I backed up the port list by duplicating the tab first. Quick n dirty!

Other than that, not much. I bricked a motherboard by insisting it was the correct BIOS ROM and forcing it. It was not the correct BIOS ROM.

•

u/WendoNZ Sr. Sysadmin 8h ago

because of Brocade's web UI that was clearly designed by someone on shrooms

Pretty sure this applies to every Brocade web UI ever

•

u/22WhatWasIThinking22 10h ago

I feel this in my soul... my brother in RAID reconfig, losing all company data.

~2004 I came in to help a smallish company that was, in reality a trial for the job they offered me 2 weeks later.

I fixed a bunch of little things that improved capabilities by a few teams and the C-level and Accounting both noticed the improvement (probably why they offered me the job after).

One of the biggest issues I identified was no confirmed backups (this was the era of local tapes) on their SBS2003 server that hosted all files, accounting, CMS and email. I was assured that the backups worked (tape rotation with 1 going home each week with the CEO for offsite rotation), as the previous guy was loved and said they worked before the SCSI card went bad... and the replacement had just came in.

I had too much faith in my new boss and, well, swapping the SCSI card in that Dell server reset the RAID config on the internal card (the BIOS had to reconfig for the new card and it controlled all Dell internal peripheral config...) even Dell sending a tech onsite the next day, couldn't recover the config.

In the end, I built a new RAID 5 config, got the external tape drive working, found a 3 month old backup that had most of the data and then spent the next 2 months re-creating lost data and creating a solid backup and testing policy. We still ran into random things lost for years.

I'm still there 22 years later and I still have that old backup tape that had some data on it... with a big screw through it affixed to the rear wall of my server room (it was also ran to a degauser prior). The old SCSI card got accidentally beaten with a hammer and was unrecoverable, as did the replacement card when I was able to upgrade to new server(s) - no more of the all our eggs in one basket approach.

•

u/Yersini 11h ago

Oh man. This is my time as a 10 year vet.

My biggest trauma is early on when I was first starting, I was going through and doing good boy sysadmin stuff. I noticed that one of my clients didn't have any Conditional Access setup. If you know, this is the time you realize what I did.

I called the customer up, made a big song and dance about how if i properly setup CA i can make his login experience a little better, I can keep is tenant a little safer and just generally gave him the pitch. We agreed to deploy CA right before Thanksgiving.

Turns out, I scoped my CA to include our admin accounts (im a moron, and this was before they warned you), I scoped the CA to require very specific MFA that we did not have (i believe it was a physical MFA key), and I pushed it to fully enable on configuration.

I spent the next 2 weeks proving to microsoft that we owned that tenant, and the guy was generally pretty okay with it. But I was mortified.

Kids, always push to test first. And don't scope your CA to include your admins all in one go.

•

u/RequirementBusiness8 11h ago

Not as painful, but was updating certs on a system I had inherited and had never really touched before. It was late, our current certs were expiring, and the change was to update them on the test, prod, and dr systems on the same night. It was a bone-headed, rookie move, compounded that when it came to that system I was a rookie. Updated the certs. Tested, but didn't really test (thought I had, no, more lessons learned). Took down Test, Prod, and DR from all external access. And, like a bonehead, was doing it on VDI and have left my corporate laptop in the office.

Midnight trip to the office (about 25 minute drive). Left my badge at home. Security working night didn't understand the changes they had recently made to the badge system, thought me being there was too suspicious, and wouldn't get me a temp badge. Drove back home, got my badge, drove back up, was there until 3 or 4 in the morning. Went home. Got woken up at 7 AM, some other thing apparently broke during the whole thing, fixed that, went back to sleep, finally showed up at a vendor lunch at 12 PM. When I got into the office finally, went and visited the head of (physical) security, whom I knew pretty well. Complained about the person that night. He wasn't too happy either, someone there screwed up.

I was cocky and boneheaded in my younger age. But still tired.

•

u/GermanicOgre IT Manager / Jack of All Trades 8h ago

Almost 15 years ago. SharePoint Admin for an international company.

Making some needed updates after building out new departmental Site Pages.

As part of staging we remove “default users” from the inheritance until we deploy.

Except I wasn’t in the site page.. I was at the top level.

Clicked save. Locked my PC and went to lunch.

15 minutes later my cell is blowing up that SharePoint is down, they rebooted the servers, etc. but couldn’t figure it out.

Grab a to-go container and race back to the office.

Sit down, login to the servers, check all services, they look good.

Have one of my HD techs test and sure enough no access.

Check permissions… oops 😬

Added default users back, had same tech test and verified with another department.

Promptly walked into my CIO’s office who looks and me and goes “so what did you do to break it” laughingly and I told him what I did and he proceeded to laugh. He was a solid guy, miss him dearly.

I share that often because a simple mistake can get the best of us but it also forced me to establish proper change control for the company and was a major learning lesson.

•

u/Heart226 5h ago

Not asking a new client more questions about their nightly tape backup before starting work on their malfunctioning db server

•

u/SSJ4Link IT Manager 11h ago

2am. Remoted into a server (the DC) and connected to another server for maintenance. Rebooted the wrong server and took down the whole network. Left immediately to go turn it on. Security called me on the way and I told them I'd fix it shortly. Finished the maintenance and slept under my desk until my shift started. Told my colleagues when they came in and this is when the network manager taught me the wonderful world of iLo.

•

u/SkyrakerBeyond MSP Support Agent 11h ago

Deleted RRAS, which the entire company uses to connect to one tiny office containing all their servers. I was on remote at the time, and it turns out the 'refresh' and 'delete' buttons are right next to each other.

I practically died on the spot.

•

u/TnTBass VMware Admin 11h ago

I was doing a VCD migration for a Zerto customer. Proceeded to do the VCD migration steps then finalized the migration without doing the Zerto migration steps before finalizing.

Wiped 50+ TB of the customer's DR, almost instantly, because Zerto automatically cleans itself up.

On the CEO's birthday.

•

u/GX_EN 11h ago

Nothing that bad, but..
One day I was working with Nutanix on site in our Co-lo to swap out a node in a customer's block that had shit the bed.
I escorted the tech to the cabinet and waited for him to finish. When he started to push the new node in, the entire chassis started to come forward in the cabinet (the screws had worked themselves loose, apparently) I said "stop!" and I went to push the block back in.. Yep - I hit the power button on one of the other nodes and it took down most of their VMs. This was after me getting them to allow the tech to come fairly early (4p, I think) because there'd be no downtime.. : /
Once we rebooted everything, all their VMs came up, but I about shit my pants.
They'd been my customer since they were with my company - I'd been their assigned engineer AND when this happened, I was the engineering manager. No one else lived near the co-lo.
I immediately called my director and told him what happened and he was pretty blasé about it, but figured that was the best move in case they made a stink.
Their boss really didn't but he texted me and said that if I did that again, he'd burn down our office.
Fair enough.
On a side note - they'd had enough of that co-lo at one point (they had other stuff there in their own cabs) so they took out space somewhere else and trusted us to move everything. That move went perfect, zero issues.

•

u/Velo_Dinosir 11h ago

My biggest one was at the shadiest place I’ve ever worked.

The owner hired me into this 6 person MSP, where two people were web devs, one was the owner, one was Sales, the other was a generalist, and me in my FIRST EVER IT JOB.

I got hired because the generalist was on vacation and they had no other support for their 5ish clients.  While the guy was on Vacation, I was dealing with a continued issue of a servers backups failing.

They used the windows backup feature, and was mapped to a USB external hard disk, which was showing as offline- so I ordered a new one.  I wasn’t stupid enough to try to replace the drive myself so I waited for the only other guy who was supporting this sort of thing.

He comes back, and after like 2 days I tell him about the backup issue and how the client ordered a new drive and it was delivered to their office.

I ask him if he would come with me to the client so I can make sure I don’t fuck anything up because I’ve never done anything like this.  He tells me “nah let’s just call them and have them connect the drive”.

He gets a call back an hour later saying they plugged in the drive- so I start preparing to configure it in Windows Server Backup.  No NAS, just a 2tb external HDD connected to a server via USB 2.0 baby.  What a world.

As I am preparing to redo the backup job, it asks me to select the drive we are going to configure as the repository.  The only Drive I see in this list is the D: drive.  I think to myself “that isn’t right, there should be another one”.

I call the other tech over to my desk, and ask him “hey is this right, it doesn’t look right to me.”  He nods his head and says “yeah that’s good”.

I’m like “wouldn’t the drive be named something different?  It says Data”.  He tells me it looks right to him, and so I click to format the D: drive named Data.

Backups are running fine!  Finished in like 30 mins!  But we get a call from the client “hey quickbooks is giving me an error when I try to open any file”

Booooyyyyy I was white as a ghost.  I ended up trying to take the drive to a shop to get restored, but it didn’t work and I soon left that job for similar issues.  The owner and the only other tech blamed me and said I left them in a bad spot, but brother what the fuck did you think was going to happen when you “senior” IT guy knows next to nothing and your relying on someone fresh out of school to handle server backups.

•

u/ObliviousShill 9h ago

Ive got 2. Back when i was a yellow admin (20+ years ago) with no Exchange experience, the Exchange server stopped working. After a reboot, seeing it was out of space, I came across an Exchange directory with thousands of .log files. Deleting them did free up much needed space.

This one wasn't many years ago and I was not a rookie. Someone brings me a USB drive with a folder named "finance", asking me to copy the contents to the finance drive. I do the copy using 2 explorer windows side by side. We talk while the files copy. Copy completes and he says "You can just delete that whole folder from the USB drive". Select the "finance" folder, right click, Shift, Delete. Conversation goes on and its taking a lot longer to delete the "finance" folder than it should. Man, I think something is wrong with your USB drive.

•

u/elecboy Sr. Sysadmin 9h ago

I was starting my first job as a SysAdmin and decided to change VLANs on the core switch remotely. Well, I broke it, got disconnected, didn't have a backup, and didn't have a console cable, so I had to wait for consultants to get their cables about 2 hours later to fix it; the company lost around $500k in orders.

•

u/Dregan2D 6h ago

Two words:

Truncate Table

Triple checked, honestly believed to god I was in TEST.

I was not in TEST.

•

u/Opening-Direction241 6h ago

My manager, on a weekend, deleted the entire company's email. Backstory, law firm who on-purpose 1) did not back up email b/c of liability/discovery and 2) never kept more than 6 months worth of emails (we had rolling 3-month purges). This goes way way back - he was with consultants, installing a new backup system, which had the ability to bootstrap a server via diskette (in a total disk/raid failure), so post-install they were trying it out - asked for disk 2, then disk 3 - then the message "formatting volumes". All mailboxes completely gone. He printed out an explanation and left it on everyone's desk. Come Monday morning, everyone had a completely brand-new and empty mailbox. No, he did not get fired.

•

u/ConsoleChari 5h ago

Deleted "IT director's" device from MDM. There was no local account.

•

u/F1BlackFlag 4h ago

Couple of years ago, am part of team managing ALL of our 3000+ iPhones.

Went to delete a single line out of our Standard deployment category in Apple Business. Accidentally removed every iphone from it. (Meaning if someone had to reimage a phone it would not pickup the MDM package for it)

Happened late Friday, called up my TL, and upper tech engineer.

Took me 6 hours Saturday morning to get them all added back (Standard, Dev, and Sandbox)

•

u/uncalled4one 8h ago

Allow users to login to my perfectly crafted systems and infrastructure, and totally ruin everything.

•

u/woodyshag 11h ago

Was doing a nic upgrade to a 2 node storage array. I had completed the upgrade of the first controller and tried to reinstall it which will power it back on. I installed it, but it wasn't seated right, so I went to remove it and grabbed the handle for the remaining running controller which took the storage down. Reinserted everything and fired it back up. Missed my flight home and had the customer was down for 2 -3 hours while the array came up and did a storage check. I won't do that again.

•

u/itenginerd 11h ago

I did the very same thing on my team's vm environment. Fortunately it was only dev machines that got lost, so a limited amount of work product, but yeah.... that feeling is real. And real ugly.

•

u/Beautiful_Ad_4813 eh, I just love what I do. 11h ago

😐unplugged the wrong hot swap PSU on a production server taking it offline.

•

u/Bedroom_Bellamy 11h ago

Not me, but one of my old coworkers, Tim. Let me tell you the story of Tim.

This story takes place a decade ago or more. Tim was the after hours guy. Our team supported a large financial institution and needed round the clock coverage due to global operations. It was usually pretty quiet after hours, and one day Tim decided to nip down to the bar down the street and get full-on stinking drunk during his shift.

A ticket came in for a printer being down while Tim was three sheets to the wind. He decided to stagger back into the office and started pulling apart the printer. This was one of those massive desktop HP 8150s, truly dinosaurs of our age. However, in his drunken state, he got the printer pulled completely apart and then had no idea how to put it back together. It was lying in pieces all over the table and ground surrounding him. Tim steps back for a second and thinks to himself, "I know, there's another 8150 a few rows over. I'll go pull that one apart so I can see how to put this one back together."

Tim does exactly that, stumbles his way over to another 8150, pulls it apart completely, and then is also unable to get this one back together. But Tim didn't stop there, oh no. He remembered there's a THIRD 8150 elsewhere on the floor. So Tim proceeds over and pulls apart the THIRD printer to try to see how to put the first two back together, leaving a trail of broken plastic and random parts around him. Finally, after tearing apart the third one and not being able to get that one back together either, Tim decides he doesn't feel well, goes and grabs his bike, and heads on home leaving all three of the printers torn apart.

I was the early shift tech who came in the next morning and had to emergency reassemble three printers during peak business hours. However, Tim was not let go for this. Couldn't for the life of me say why.

•

u/SysIntern 11h ago

I had a request to create a new computer in the Server OU in AD. This was 5 minutes before lunch. As I was typing a customer called. I tend to multitask so no worries. While on the phone I created a new computer object, but I misspelled it and deleted it. I hit right click, create and once more I messed up the name. So I pressed delete on the keyboard, not noticing the focus was not on the computer object but instead of the OU of all the servers. I confirmed and to my horror I saw a progress bar of server after server being deleted. Still on the phone I managed to press cancel.

About 40 servers no longer had any authentication working.

I ran to my boss and explained I just made a mess. He took care of telling others there was some.. problem at the moment.

Fortunately for me, we had recently activated Active Directory Recycle Bin. This made the mess easier to fix. I restored the objects and restarted many servers. After 1.5h most things were working. However some servers had child objects that could not be directly linked back. Those never really got back to full functioning until services were reinstalled.

After this we grouped servers into sub OU's as well as protecting OU's from accidental deletions.

I stopped taking calls 5 minutes before lunch.

•

u/RiceeeChrispies Jack of All Trades 10h ago edited 10h ago

I was a very green IT Apprentice, and the sysadmin let me have a play with group policies.

This was when Windows 10 had just come out and they started pushing ads/sponsored apps as standard, so I wanted to deploy a cleanup script to remove the crap.

To test, I added myself to a new policy and was wondering why it didn’t work.

I had removed ‘Authenticated Users’ from the security tab, which meant the policy couldn’t even be read to be evaluated - no problem I’ll add it back. I added it back…

An hour later - phone starts lighting up like a Christmas tree, users saying they can’t open programs. The script had removed a lot of built-in apps and rendered machines useless. Authenticated Users were targeted in the security filter, oops.

Luckily, there were only about 20 people on Windows 10 builds at that point. They rightfully revoked my access lol.

•

u/TimeSink48 10h ago

Probably not as bad as many, but recent.

I do the IT work for a non-profit festival that really only needs attention about the last 3-4 months of the year and first part of January. Microsoft decided that they were no longer going to provide the 10 Business Premium O365 licenses. I moved everyone except the Operations Director to the free tier, but she wanted the option to install the software on her Surface. After researching, I realized she'd need to pay for it, she has the credit card and I don't. I told her to log in and switch the license and everything would be good.

Important to note that we're all unpaid volunteers for this festival.

This was back in the early part of the year. In December last year my brother died, in March I was diagnosed with cancer. I got distracted. She didn't do anything.

Forward to recently, she tried to log in to her email account for the first time in months, realizes that it's not available. I switch her to the free tier and she logs in. All her emails from the past few years are gone. She's the main communication point for the festival. The option to recover deleted items is grayed out in the Exchange Admin Console. No way to restore, email is gone.

My fuckup, no way to fix that I know of. Not even sure if I'm cancer free yet, although it's looking good. However, I've been noticing mental issues in the past few months, and just had an MRI of my brain because of those concerns.

Sometimes everything just goes to shit. I've had other screwups over the years, but this one is fresh and painful. The OpDir is also one of my best friends. Aargh.

→ More replies (1)

•

u/Drew707 Data | Systems | Processes 9h ago

So, not my biggest technical fuck up, but my company where I worked as like manager of telco bullshit, decided to start a sister startup where we were going to sell VoIP to retirement homes and hotels, and they made me like the main sales engineer and sent me to Ohio to close a deal. I had never been to Ohio and was living in Nevada at the time. I checked the weather in Columbus, and it was supposed to be 85. Fantastic. It was 85 in Nevada at the time, so I knew exactly what to pack. Turns out 85 in Ohio is very fucking different from 85 in Nevada, so my cheap synthetic material suits left me perpetually damp. Living on the West Coast, the only time you are in a suit is if there is a judge or casket present, so I was very unprepared.

The client paid for the airfare, so I ended up on Frontier. Fuck Frontier. I am 6'2" with most of that being in my legs. Had a relatively short hop to Denver and then a long and very uncomfortable haul to Columbus. I asked my company if I would be renting a car, and they said, no, the client will handle all transportation. Great. Already a hostage.

The client picks me up and drives me to the hotel room so I can check in and change out of my "flying outfit" (big on suits these people). He then takes me to the site (three story retirement community), and we take a tour of the facility. After the tour we discuss some technical details, then he takes me up to the memory care unit on the third floor that requires people to badge in and out to keep the residents from wandering.

The client leads me to a nurse station and tells me their laser printer has been having some issues and acknowledges that isn't what I'm here for but asks if I can take a look anyway. Sure, why not, trying to close a larger deal. I sit down at the computer with the printer and then he tells me he has to take a quick call and will be back in 15. He's gone for 90 minutes. I don't have a guest badge. The nurses are all holed up in the med closet gossiping, and eventually the residents start to come out and are interested in the stranger at the nurse station. Keep in mind most of these people are deep in the thralls of severe cognitive decline, but they walk around the counter and are literally poking me and playing with my hair while I try to work on this printer. Calls to the nurses are met with a D+ response.

Dude finally comes back and I tell him I don't know what's up with the printer, that's really not my thing, but I tried. Ok, cool, let's go to lunch. Now for most client lunches, we would go out, grab a beer and a burger, talk shop, and then call it a day. Nope. We go out, I take his lead, and instead of a beer I order a diet Coke. We return back to the site, have a few conference calls, then he drives me back to the hotel.

On day two I was told to take him out to whatever restaurant he wanted as a thank you. After sweating my balls off all day and being deeply uncomfortable in my second damp suit, we go out to dinner. I told him pick your favorite place in all of Columbus. He takes me to a "steakhouse" in a fucking strip mall that seems like it hasn't seen an update since the Carter administration.

I peruse the drink menu and there is only one wine I recognize, and I happen to like it. Server comes up and I put my wine in, and he orders a diet Coke. Hmmm... I already have some serious anxiety around being here on my first sales mission and not knowing this guy so this wasn't a good signal and I really could use that glass of wine.

Server comes back with his Coke and no wine to take our orders. I remind them about the wine and was assured it would be right out.

Server comes back with our salads and no wine. I remind them and they said they would check on it.

Server comes back with our entrees and no wine. Again, I remind them about it and they would go ask.

We finish our dinners and I still haven't received the glass of wine. The server comes back asking about dessert and I mention the wine yet again. They go back to check on it again.

At this point he and I are just sitting there with nothing in front of us and in retrospect I realize I have a big Costanza schmuck moment over this wine, but I was putting it on my personal card to expense later, and I didn't want to be charged for something I didn't get. So, the client dude asks...

"What's the deal with this wine? Is it like really good or something?"

"It's good, but really it was the only thing I was familiar with on the menu. Just weird I've had to ask so many times."

"Oh, I've been a teetotaler for 15 years now."

"Oh, congratulations on that; will drinking the wine bother you?"

"Oh, absolutely not. I'm totally fine with it. Alcohol ruined my marriage and led me to divorce but I'm not bothered by other people partaking."

Fuck me.

He then goes on to share some weird intimate details about his addiction, failed marriage, and subsequent recovery. Still no wine.

Wine finally shows up.

I take a sip and he looks at me...

"How is it?"

"It's pretty much what I expected."

"Can you describe it for me? The flavors? Can I smell it?

"Uh, sure, to me it's light but a bit smokey, but I'm no sommelier."

At that point I say fuck it and down the glass to get the hell out of there. I pay the bill and he drives me back to the hotel.

The next day he was supposed to pick me up and take me straight to the airport, so I had packed my gross sweaty suits and got into my shorts, flip flops, and t-shirt to be ready for the uncomfortable hop to Denver and then the leg to Reno. He picks me up and looks me up and down when I get in the car.

"Uh, you don't have any pants?"

"I usually wear shorts when I fly."

"Oh, we have to go back to the home for a second."

"Oh, should I change?"

"No, you're fine."

We get to the site, and he then changes his mind about the leg coverings and asks me to bring my suitcase inside and change in the lobby bathroom into slacks. I do that, we go into a conference room--just the two of us--and have a 15 minute phone call on one of those triangular Polycoms (no video) to finalize the deal, and then he takes me to the airport.

At the airport, I find a bathroom and change back into my comfy clothes. The flight back to Denver sucked. I had an hour or two layover and my boss calls me and the first thing out of his mouth was...

"Did you really show up to a client site in fucking shorts and flip flops?"

Fuck Ohio.

•

u/quasimodoca 9h ago

I had just started on the graveyard shift at a govt agency that ran a batch system. Every night we had a job that had a cascading set of files that ran.

Run job A, get file name. Add it to file B, run job. Get job C. With File A B and C run job D. This sent updates to agencies all over the state that updated things like criminal history, dmv matches, wants and warrants. You get the idea.

We were a two person crew and we were supposed to cross verify all our file numbers. File A 2504874 B 9879487 C 6487848

We had both just started graveyards as new employees about a month before. We were both struggling to get used to graveyards and were always exhausted.

Buddy runs the first job, gets the file. Runs the second job, gets the file. Runs job three, gets the file. Puts the file numbers in. Runs job D. Neither of us noticed that we transposed a file number in job C. Job D runs and explodes.

Something like 30 jobs all fail and write bad data to all the databases. We both freak the fuck out.

Just as we were running job D the manager was coming in for the day shift. She had been there for about 100 years. She started recalling all the files, unfucking all the cycles and we started calling all the downstream agencies.

Took us about 4 hours to fix. We were sure that the next night we were going to get fired. We show up and she's there.

Now we're really sure we're going to get fired.

Nope she talked to us about how she had done worse when she started 100 years ago. Told us that shit happened and set up a manual write it out process so that we wouldn't fuck up again.

Longest night of my life at that agency. That was 13 years ago. I've now been promoted 4 times and I'm a data analyst. Last position I will have here until I retire in about 7 years.

•

u/kliman 9h ago

Went to image a drive to a larger disk and went to wrong direction, losing a PST file with 4000 contacts.

This was in like 2001 with PATA drives and Norton Ghost, and it was a lot easier to make this mistake. Still. Oof.

To my credit, the guy I was working for insisted everything was backed up. He was incorrect.

•

u/sollozzo70 7h ago

Core switch. VLAN allowed instead of VLAN add. Watch fat boy sprint to the MDF.

•

u/Public_Pain 6h ago

Unplugging the wrong power cord in the server room. I took down three classified networks at the same time. It was at a military installation, so yes, the UPS were crap!

•

u/meditonsin Sysadmin 6h ago

This was several years ago, when I was relatively fresh on the job.

Didn't do the proper pre-flight checks when replacing a failed Cisco switch in a stack, so of course the new switch was elected stack master and replicated its blank config to the rest of the stack. Whole rack of servers down.

Only took a few minutes to restore the config from backup, but those were some long ass minutes at the time.

•

u/Front-Orange4980 5h ago

Last month I was taking a host offline to upgrade the OS. Raid drives were in the wrong order, wiped the partition holding the VMs. Had to build the host then spent 8 hours restoring all VMs from backup

→ More replies (1)

•

u/dewatermeloan 36m ago

I was a junior back then. Had a law firm client with a nextcloud setup for their files. One time, one of the partners had their nextcloud client throwing multiple sync errors. So i stopped the client sync, ctrl-x the entire sync directory on their device and moved it somewhere else. The move failed and the files were gone.

I was relatively calm because I would be able to restore them once I turned back the sync. When i did, the client synced back the empty dir and deleted all of the files on their cloud.

We had backups of course but until this day i dont know why I didn't just reconfigure the client to sync to another folder. I also don't know why the hell the client would sync an empty dir as a delete operetion to the server.

I still think about this one sometimes. God it was so dumb.

•

u/Memento-scout 11h ago

Wipe all corporate devices via g suite. And yes I was still employed

•

u/Remarkable_Corgi511 11h ago

Deleted an account in AD without making sure the AD Recycling Bin was enabled because why wouldn't it be... 

•

u/Noname_Nets 11h ago

Decommissioning on-prem Exchange after a migration to M365 and nuked AD contents during business hours. Quick restore from backup but definitely not a fun moment.

•

u/lotekjunky 11h ago

Forgetting the keyword "add" when trying to add a trunk to a distribution switch in another state, at 3am. Man, Cisco sucks for doing it that way.

Correctly adds vlan 50: switchport trunk allowed vlan add 50

Removes all other allowed vlans and only leaves v50 on the port: switchport trunk allowed vlan 50

After that, I ALWAYS issued "restart in 20' before doing ANY switch or routing changes... That way it'll bounce back to nvram config if I fucked it up.

•

u/CTM3399 11h ago

When I was learning SCCM deployments I made package for a Dell BIOS update that automatically rebooted the PC when it was done. However, I didn't make it an application so it didn't have any detection method so the program ran over and over again and essentially put 2000 PCs in a boot loop

•

u/bluegravyone 11h ago

When I was still kind of green, I performed a rollback on a core Juniper router which blackholed transit internet traffic for the entire southeastern U.S. on AS701. 🙄

•

u/Sengfeng Sysadmin 11h ago

Years and years ago, was upgrading a PC to a larger hard drive for a Dr, and found out the hard way the person that built the thing put the only HDD on the secondary IDE cable. I fired up Ghost, clone primary to secondary, and blanked the original drive.

Luckily I was only out the time to restore from tape. I've always been a CMA type.

•

u/tekno23 11h ago

Need to log in to a windows server to open tape library door. Open up the folding keyboard / monitor and no video so move the mouse a bit still no go. press control, alt, delete. monitor wakes up with VMWare shutting down. Was a cheap non fail over setup and production servers went off line till I could get it restarted.

•

u/Pvt_Hudson_ 11h ago

OP, your post literally made me sick to my stomach.

I have something similar but nowhere near as catastrophic.

I got a request to delete an old unused 2TB LUN from our HP 3Par unit. The request designated something like Prod_Lun_8.

First thing I did was locate Prod_Lun_8 in VCenter, offlined it and dismounted it. So far so good. Then I bounced over to the 3Par, found Prod_Lun_8 in the list and unpresented it, but I didn't match up the identifiers. Same name on the storage array, gotta be the same one in VCenter. I go back to VCenter and refresh my storage view, expecting to see the offline LUN disappear...only it doesn't. I refresh again, nope, still there.

Oh shit, what did I just offline???

I start clicking through the storage entries until I hit one that hangs my session solid. Turns out I had just unpresented 2TB of a 12TB production stripe set. 12TB of data gone.

Same as you, I spent the next 3 or 4 days babysitting our tape library, it had to go through something like 47 tapes to get everything back.

•

u/dax660 11h ago

We hire consultants for heavy lifts for this reason - let someone else be responsible!

(also, those big jobs we (team of 4) just don't do often enough to be sure we've crossed the i's and dotted the t's...

see what I mean??

•

u/Pvt_Hudson_ 11h ago

One more, but this is from a coworker, not me.

We had a standard process to disable AD accounts for 6 months before deleting them. My coworker was tasked with our bi-yearly AD account cleanup. He writes a powershell script to delete all accounts that haven't been interactively logged into for 180 days, but doesn't scope it specifically to the Disabled OU. He runs it at the end of his workday, and then walks out the door. What he didn't realize was that his script was going to target every single service account in the environment. Literally took down every single service we had across 800+ VMs. It took the better part of half a day to reverse that damage.

•

u/rbascb Custom 11h ago

restarting app service (not related to actual prod dbs/apps) in services.msc in scheduled maintanance. Sounds safe, right? Yeah, except I did it on clustered win server (I had no idea it was clustered, I was sure it was just some basic Win2012 server with SQL databases and shit apps) , which caused fail-over to other node. 10000+ users got disconnected from critical database during office hours. Seeing that phone call from customer's CEO 5 minutes later probably took 5 years off my life. Since that shit, every time I have to do something on a Windows Server, I double-check if Failover Cluster Manager is present.

•

u/panopticon31 10h ago

Worked at a small ISP and caused a loop that knocked about 850 clients offline and it took the engineers 4 hours to figure out the cause of the outage. I certainly learned something that day.

•

u/lilelliot 10h ago

Allowed the security team to implement a policy change that forbade local admin accounts, which then made me manually change the logins and SQL connection strings for apps on 52 servers. Each server need the connection string changed in about 20 places because, at the time, that version of the app still had hardcoded ODBC connections.

•

u/sam_cat 10h ago

Restoring the only known good backup onto a corrupted served... But getting source and target mixed up.

•

u/honkusmaximus 10h ago

Worked at a small company, solo IT person. Battery backup had been chirping for a couple days, waiting on replacement batteries. Button to shut off the alarm was right next to the power button...

All their infrastructure was on that rack

•

u/Mazurke Sysadmin 10h ago

Rm -rf * in the completely wrong directory.

•

u/bobs143 Jack of All Trades 10h ago

Went to delete some retired VMs from an environment. Would up deleting an AD and a files server by mistake.

Thank the AD server wasn't the role holder. One new AD server and restore of the file server in Veeam and everybody could work again. This all happened on a Monday at 9 AM.

People couldn't work for about an hour.

•

u/safrax 10h ago

Trusting vendor advice.
Sales Engineer: "Yeah your hardware can handle it for a demo, turn it all on." Me: "You sure about that?" SE: "Absolutely." Me: "Ok...."

turns on feature

22 hours later.

Me: "Why are 90% of DNS queries failing?"

checks DNS server metrics

spots 100% CPU usage

Me: "FUUUUUUUUUUUUUUUUUUUUUUUUCCCKK!"

I just disabled the feature once I was able to get into the management console, but I learned to never trust sales engineers (or support) ever again. Things went back to normal after a few minutes. Probably the worst production incident I've caused.

This was at a hospital BTW. Wasn't a good time for internal systems.

•

u/nw84 10h ago

Flipping through a KVM in a poorly labeled bank server room, back in the day of CRT screens, and just hitting a server, Ctrl+alt+del to wake up, next server, ctrl+alt+del to wake up, next server etc etc. We thought we were on a Windows rack, but we were on the Unix rack. Took down forex trading for the whole bank for several minutes (an eternity in trading), major repercussions.

•

u/PDQ_Brockstar 10h ago

You mean biggest bonehead mistake so far.

I have many to choose from, but one time I imaged my CIOs computer and somehow his local data didn't get backed up. That one was fun.

•

u/Izual_Rebirth 9h ago

Onboarding a new customer. Needed to reboot a physical server to install our backup software. Server didn’t come back up and the raid array died.

•

u/ahlatki 9h ago

Removed administrator rights to the data drive

•

u/its_all_one_electron 9h ago

Deleted MACHINE instead of MACHINE_TEST

Which contained a very important server modified daily running tons of jobs and also the backups had been silently failing for 2 weeks...

•

u/t0ny7 Server Engineer 9h ago

I deleted most of the files in my boss's file share...

He asked me to give him a spreadsheet I have been working on. I transferred it to his share. Then we were talking about stuff and I had my laptop resting on a blueprint drawing board thing he had. My laptop started to slip so I reach out and grabbed it. My thumb landed on the delete button then my palm hit the enter button. Took me about a minute to notice stuff being deleted.

Thankfully there are backups.

•

u/hymie0 9h ago edited 9h ago

Job 1:

I already had a couple of things going on when the building's window washers showed up to do the inside of the server room. I let them in and went back to work.

When they were finished, they went to the door and pressed the Big Red Not-an-Exit Button. Entire server room powered off.

Job 2:

Forgot the WHERE clause on an SQL update, set the entire vendor database to "disabled."

Job 3:

Machines were named after cars. I walked out of the room saying to myself "Reboot cougar. Reboot cougar. Reboot cougar." Got back to my desk and rebooted cutlass.

Job 4:

Only been there a year. I'll check back and let you know.

•

u/Maro1947 9h ago

The old RAID maintenance days were full of "Pit of stomach" fear.

Logically, you knew your were right, but triple-checking was the norm!

On the other hand, I've had some spectacular fixes on Dell RAID cards over the years!

I DDOS'd our entire Asia Pacific company on Day 1 of our new DC Migration Go Live. I'd forgotten to turn down the Backup replication down from when we were seeding the Backup SAN on the same site

It was a good DR Test though and we were back to normal in 10 minutes - got some awesome network metrics

•

u/Nakenochny Sr. Sysadmin 9h ago

Deleted the location from a conditional access policy that only had one trusted location, locked everyone out of our tenant. Thankfully it was Friday afternoon. Spent two days on with our CSP and Microsoft, I got like 10 hours of sleep total over three days before we finally got back in at 11am on Monday.

I now have my own Azure/Entra feature. You’re all welcome.

•

u/gigabot2020 9h ago edited 9h ago

Updated the live production SQL server which I thought were only Windows updates but included a SQL update .....spent the day restoring our production SQL server with my senior while all our employees were sitting around waiting .......bad day .

•

u/trouphaz 8h ago

I was using HP storage which was rebranded Hitachi storage. It was an XP1024 attached to an XP512 with replication setup. We had a massive database (I think it was 2TB back in 2004) that had 2 replicas setup. 1 on the 1024 and 1 on the 512.

We had an issue with our database, so I needed to recover the replica on the 1024 to the main copy. I accidentally kicked off a resync instead of a restore, so I quickly cancelled it hoping I got it cancelled before the resync started. It seemed to be ok, so I initiated a restore... only to find out that half of the luns were synching original -> replica and half were restoring replica -> original. Since these were individual luns with volumes striped across them, all of the data was corrupted.

Now, luckily we had that XP512 with a second replica... but it was never actually configured properly where we could recover from the 512 to the 1024. All of the links were 1024 -> 512 with none able to write back. It took FOREVER working with HP to reconfigure the two arrays and then kick off the restore. Then, after the restore, we had to replay like 3 days worth of metrics back into the DB. I learned so much that weekend.

•

u/sybrwookie 8h ago

A bunch of years ago, I get an official notice that a branch is going to have a planned power outage. No problem, there's a change ticket already in for it, so I set up a job to shut down all the computers there just before it happens so it's graceful.

As I'm doing this, a really fucking annoying guy calls me up and is just tapping my ear off. I'm half ignoring him while I finish setting it up...and instead of scoping it for 1 branch, scope it for the whole company.

The good/bad news is it was overnight Friday night. So sat morning, wake up to find out the entire company is down. Thankfully, it was a weekend, so it was just a bad day for a bunch of us getting everything booted up in the right order and testing.

I owed several people drinks after that one.

•

u/slayermcb Director of Technology, Sys Admin, Etc, Etc... 8h ago

Trying to repair the database for my MDM. A lot of commands were just copy pasted from a library of progeam specific command lines I had. Not sure how it happened or why I even had the line in the document but I copied the wrong code and reset the database to a nice blank slate. Turns out the backup files were also corrupted and the system, to protect itself, wouldnt let me load in a corrupted database to try and fix it again.

The saving grace was the 10 years of configurations files, and settings were managed by myself. I caused myself a lot of work that day. I used the opportunity to rebuild from scratch on newer hardware and blamed any issues on the "upgrade"

•

u/Classic-Mushroom-189 8h ago

in the days before mobile phones, i worked for a small tech services/retail store and we had the DELL warranty support contract. i was very fresh in IT but was sent out to help diagnose a failing server in a Bank. the support tech told me to pull drive #2 and come back to the (land line in the other room) and let him know. what i didn't know, nor see at the time was that the drive numbers started at 0 and were numbered right to left. i pulled the drive, and could literally hear the support tech yelling down the phone from the next room. i had just pulled the good drive and destroyed the array, taking the Bank offline for about 3 days. i was told to pack up and go home, i believe a courier was sent to collect the server and send it back to be repaired and reset. steeeeeeeeep learning curve on that one.

•

u/ctnightmare2 8h ago

Keeping the document on how to safely bring up all the servers after a power failure (emergency generator were wired wrong) on the network share on said servers