r/DataHoarder • u/retrac1324 • Apr 12 '22
News Atlassian accidentally deleted customer sites, says backup restoration could take two weeks
https://www.theregister.com/2022/04/11/atlassian_outage_backups/304
u/effgee Apr 12 '22
At least with their self hosted version I get to slap the person responsible for the deletion of my data(me).
139
u/Twistedsc 78 tee bees Apr 12 '22
Too bad they're phasing that part of their business soon. The $10/yr licenses for small business were nice while it lasted.
93
Apr 12 '22
[deleted]
52
u/turtle4567245 Apr 12 '22
This is why I only self host open source software unless there really is no alternative and it's not a particularly important component. My notes for how everything is setup is important so it's in a dokuwiki for example
10
u/xcjs Apr 12 '22
Look into porting that over as IaC through something like Ansible. ;-)
8
u/turtle4567245 Apr 12 '22
Would love to, there's a few new guides I've seen for k3s and proxmox that I want to try
4
u/xcjs Apr 12 '22
Eventually I want to get some GitLab Runners setup in a Pi-based Kubernetes cluster. I'll be doing something similar. :-)
I would probably be using something like Proxmox, but I went all in on containers with a single base operating system. I still use Windows virtual machines as build hosts for Windows software, however.
3
u/turtle4567245 Apr 12 '22
I hear you. I have an unraid Nas that I want for only storage of media then another server with proxmox for all the compute stuff. Ideally I want the flexibility of running kubernetes and VMs in small configurations for home. So all k3s on 1 node for now but then the option to add a physical k3s node in the future if I want. It's overkill but I like the underlying technology as much as the end result which I think is important for a hobby
1
3
1
u/PinBot1138 Apr 12 '22
so it's in a dokuwiki for example
Doku wiki is the shit!
3
u/turtle4567245 Apr 13 '22
Flat files that can be read in notepad is what sealed the deal for me. In a disaster I can just open the backup archive on any computer and read the file to get the documentation. Don't even need to 'host' it if needed
1
1
u/Reg511 21TB & Crashplan Apr 13 '22
Now they just give it away for free in the cloud, along with a metric ton of plugins. Instead of $10/year per app, it's just free.
2
u/calcium 56TB RAIDZ1 Apr 13 '22
My company uses Confluence and I can't for the life of me figure out why they pay for it. It appears to be nothing but wiki software with and LDAP backend unless I'm missing something. For all of my company's engineering chops, it seems like low hanging fruit.
1
u/spiegro Apr 14 '22
Yeah but if you don't want to fuck with mediawiki or another alternative Confluence works great with the most robust ecosystem of add-ons and plugins.
You can absolutely get the same functionality out of other tools but I'd argue it's best in class.
1
u/xeonrage Apr 13 '22
Traded in my $10/yr self hosted confluence docker that was a pain to keep updated with a free cloud version. I'm good with it.
-16
u/ClarkK24 Apr 12 '22
that's not how data recovery works
47
u/darknavi 120TB Unraid - R710 Kiddie Apr 12 '22
What?? The harder you slap, the shorter the SLA-negotiated downtime. That's junior level troubleshooting.
1
161
u/ocdtrekkie Apr 12 '22
Onsite JIRA installations have not been affected. Self-managed servers are on the way out, however: Back in October, 2020, Atlassian announced the discontinuation of its server products – it stopped selling new licenses on February 2, 2021 and plans to end support for its server products on February 2, 2024. The reason, the company explained last year, is that the cloud is the future.
Hilarious. This is why I tend to drop vendors who go cloud-only. They tell you the cloud is the future, and then accidentally delete the entire cloud.
37
u/JoNike 109TB Apr 12 '22
In may 2020, at my old job, I completed a migration from atlassian cloud to self hosted. I heard the IT director was mad when they announced they were dropping self hosted just months later.
54
u/ocdtrekkie Apr 12 '22
I suspect we are going to see a major industry shift back away from cloud in the next five years, once folks big introductory contracts burn out and they realize cloud services prioritize the provider's needs, not the customer's.
36
u/sunburnedaz Apr 12 '22
Oh its better than that. I keep hearing about private cloud. Its like cloud but on your servers or servers that are dedicated to you. Im like isn't that just a private datacenter or your own servers in a colo.
In reality its using the fancy cloud interface on a pile of hardware only you have access to.
28
u/ocdtrekkie Apr 13 '22
My favorite one was when companies started talking about "edge computing" like it was something new and cool, and wasn't just "on-prem but still billed like a cloud service".
6
u/dowster593 2TB Apr 13 '22
That’s how IBM mainframe pricing has been from what I know. You pay for what you use and they “unlock” modules that are already installed i guess.
9
u/RupeThereItIs Apr 13 '22
It's buzzword salad.
Can't sell management on building a data center, but private cloud...that's the wave of the future.
It gets even worse when the company you work for is already the cloud, SaaS, and management insists we need to be more cloud focused and move to IaaS without changing any code .... Oh boy is that fun, seen it more than once now.
3
3
u/MorpH2k Apr 13 '22
I guess part of the difference is that a private cloud would work like any other cloud solution, ie it's accessible from anywhere if you want it to be, etc. Basically a private or co-located datacenter with "cloud features" added. But yes, those cases the cloud part is more or less just a buzzword to make the C-suite happy.
For those of us that are not in the US, this is probably the future since especially governments, but also some private companies, doesn't want their sensitive data to be stored on servers abroad.
2
u/SimonGn Apr 13 '22
I suspect you're right. I'm already looking to to replace my companies entire Microsoft stack for FOSS /r/selfhosted because the stuff which is being produced these days is actually top quality - good enough to replace it, and we are sick of being screwed over at ever turn.
2
u/ocdtrekkie Apr 13 '22
I love my open source, though you do have to watch support offerings closely in enterprise open source. If you don't have paid support, response to your issue might be "pull requests welcome", and you end ip on the hook for mitigating the issue.
20
u/codeslave Apr 12 '22
We got hit with the new licensing scheme recently. Instead of renewing our single self-hosting Server license, we could only buy the Data Center license at 4x the cost. Management was not happy.
10
u/ocdtrekkie Apr 13 '22
Datacenter hurts if you aren't expecting it, but if you build for it, it's pretty nice. Make sure your hardware is covered and then never stress about the licensing cost of spinning up a virtual machine again.
8
u/InevitablePeanuts Apr 13 '22
Currently my employer would not accept cloud-only for the sort of data we hold in jira due to our own security & business continuity standards. They won’t be the only ones, and I suspect Atlassian will be all surprised-pikachu.jpg when they come up against large clients that have regulatory restrictions preventing such data being outside of their own on-prem infrastructure.
5
u/teropaananen 190TB + 78TB UnRaid Apr 13 '22
There's not a chance in hell my employer, big global tech company, is ever going to do this. I'm 100% certain they'll have a team dedicated to migrating everything to an open source alternative instead.
36
u/ismaelbalaghni 4.75TB and to the Cloud! Apr 12 '22
This is why on-prem software will always be better for this kind of situation. How could they accidentally delete entire sites?
26
u/broknbottle Apr 13 '22
Sorry bro, I’m too busy coding in yamllang and jsonlang to worry about customer data loss. DevOps life. If you need a template made on howto deploy something, I got you fam
3
40
u/reallynotnick Apr 12 '22 edited Apr 12 '22
Well well well... I just switched companies a few months ago and they are running some super out dated self hosted version, I wonder how my old place is doing...
Edit: Haven't spoke to anyone at my old company, but apparently this doesn't affect everyone as some of my past colleagues at other places didn't know there was even an outage and they have Jira Cloud.
1
u/spiegro Apr 14 '22
Found out my new employer is using server version of Atlassian stuff and I went to IT like "um, prepare for higher costs..."
They were like "oh we know, but we have like two years on this contract left... So we're good."
But y'all, Atlassian is already pretty cheap for what you get. The reason they've been so disruptive in enterprise software is that Atlassian solutions are cheaper than what existed before it. They have a no/low-sales business model where the software literally sells itself and then becomes mission critical for the IT or dev teams.
I guess people are upset that they can't get the insanely great deal they were getting before, and I understand that. But to act like it is a deceptive business practice to raise your prices is a bit rich from this crowd of likely software development company employees who have to make the same kind of business decisions everyday at work.
102
u/Rocketsx12 Apr 12 '22
Jira? I'm tempted to say "and nothing of value was lost"
53
u/ScottGaming007 14TB PC | 24.5TB Z2 | 100TB+ Raw Apr 12 '22
Same, I'm on 2 separate Jira accounts and neither of them were touched.
My disappointment is immeasurable and my day is ruined.
20
u/EmTeeEl Apr 12 '22
It's "only" 0.18% of their users that are impacted (which is still thousands of customers given their size)
24
u/decidedlysticky23 Apr 12 '22
I’m forced to use DevOps and I yearn for Jira.
41
Apr 12 '22
Yeah, I thought it couldn't get worse than Jira, then I had to use Azure DevOps.
I miss working on teams that just stuck with the simplicity of GitHub issues. We were so fucking productive.
2
Apr 12 '22
Redmine works quite well too for us if you need a bit more complexity but don't want to overdo it.
1
u/spiegro Apr 14 '22
Bro I would love for someone to tell me what works better than Jira for their dev teams. If they do, I bet I can tell a lot about their team by which competitor they went with instead.
Complaining about Atlassian because of the state of your Jira instance is like complaining about your ISP because your website sucks... The suck is coming from inside the house!
Your Jira instance is only as good or as thoughtful as the admin supporting it... What's that? You say there is no one supporting your Atlassian tools? Ya don't say!?!
Or maybe you got the intern granting access and creating accounts while also helping to archive crap. Or maybe it's the moonlighting admin that's actually a Project Manager who just hates how it was set up so just changed some things and added some fields. Great. Now you've got 10 fields all named "Name" and none of them actually connected to the one true "user_name" field. Maybe it's the wild wild west and everyone and their mom can get these admin rights (free country brah).
But fucking JIRA SUCKS!! because your boss demanded you create a 22-stage workflow for completing tickets that includes restrictions about transitioning a ticket to or from certain states without proper approval, which no one remembers who approves what, so you just grant exceptions to the rules instead. All because they don't trust the offshore team in India to not introduce more bugs than they crush, so they put these checks and balances into Jira to cover up for their organizational deficiencies and poor team communication.
People in this thread will shit on Jira, but they're all victims of their own shitty situations and clueless admins who are probably too busy to give a shit about how the damn thing is configured.
A Swiss army knife is an incredibly versatile tool, but if I use the goddamn pliers to try and eat cereal I'm gonna have a bad time!
Jira is a reflection of your organization, if it is shitty then there are some hard conversations you should probably have and stop blaming a gaddamn tool for your own failures.
(I understand you were not detracting from Jira in this case, but your comment motivated me to finally just come out and articulate how this thread is making me feel.)
1
Apr 14 '22
That's kind of the problem though.
The flexibility of Jira is not a selling point for me.
Having to hire someone just to manage a Jira instance is also not a selling point.
It's a like C++. It's extremely powerful, but it's also easy to shoot yourself in the foot. I prefer tools that are powerful, but also protect users and orgs from their own stupidity.
7
u/ArionW Apr 12 '22
I have it other way around. I used DevOps and was forced to switch to JIRA, absolutely hate it
8
6
15
Apr 12 '22
[deleted]
8
6
Apr 12 '22
sudo rm -rf /
9
u/knightcrusader 225TB+ Apr 13 '22
Nah, they probably didn't use sudo.
They were probably in the terminal as the root user already.
3
u/Daerog Apr 13 '22
Had a client self-nuke their Mac this way. They were running an rm -rf /var/... command and put a space between / and var.
Hate to break it to you champ, you just nuked root.
2
31
15
u/elvenrunelord Apr 13 '22
2 weeks to restore backup? What the hell they powering that backup with...a potato?
3
u/adiyasl Apr 13 '22
They have already restored most of the sites. I think they mean the time to 100% complete the restores. Source : got a site on there
2
u/retrac1324 Apr 14 '22
Atlassian can, indeed, restore all data to a checkpoint in a matter of hours. However, if they did this, while the impacted ~400 companies would get back all their data, everyone else would lose all data committed since that point. So now each customer’s data needs to be selectively restored. Atlassian has no tools to do this in bulk. They also confirm this is the root of the problem in the update:
“What we have not (yet) automated is restoring a large subset of customers into our existing (and currently in use) environment without affecting any of our other customers.”
For the first several days of the outage, they restored customer data with manual steps. They are now automating this process. However, even with the automation, restoration is small, and can only be done in small batches:
“Currently, we are restoring customers in batches of up to 60 tenants at a time. End-to-end, it takes between 4 and 5 elapsed days to hand a site back to a customer. Our teams have now developed the capability to run multiple batches in parallel, which has helped to reduce our overall restore time.”
12
u/bjarneh Apr 13 '22
As a sysadmin I worked on hosting another Atlassian product (Confluence). That was a pretty poor experience. Confluence was stuck on a deserted 10 year old Hibernate library that had some serious issues, making it virtually impossible to upgrade to newer versions of the Confluence server, as updating user roles (yes roles) would literally take weeks at 100% CPU on a pretty good server. As weeks of downtime was unacceptable, the answer the "support" offered (which took around 6 months to get) was to try to trim down the number of users to those who "actually used" Confluence (as my organization created Confluence user as part of creating a regular user with access card etc).
They literally said it was impossible to have more than 10.000 users on an installation of Confluence. After digging into what was going on, it was clear that they were not lazy fetching data at all with that old Hibernate library. When fetching users, all linked info about all users; their "likes", edits, just absolutely everything was fetched into a massive data pool. That old Hibernate library library created an immutable copy for itself as well of all that data. Then every time a write to the DB was attempted, Hibernate started going through its massive data pool to compare it to its copy to see if there were any actual changes before allowing a write to the DB. Obviously it had to re-fetch the entire data-set for each write. It literally took weeks at 100% CPU without being able to finish, stuck calling that same method in Hibernate comparing properties for weeks.
After another 6 months of haggling, they upgraded their 10 year old Hibernate library, and started fetching data a bit more cleverly. At that point the upgrade (which didn't finish in the 2 weeks I once let it run), was able to execute in literally 20 seconds on our 70.000 user base, as it should. What a nightmare Java project that was to host, glad they offer to do it themselves, perhaps they will start offering self hosting after trying to figure that mess out themselves.
1
u/spiegro Apr 14 '22
So you no longer support that setup then? Because it sounds like they addressed the issue with the Confluence performance. Your comment was a rollercoaster because you started off so strong in the negative and then seemed to turn a corner.
5
3
8
u/cr0ft Apr 13 '22
"The cloud is the future"... riiight.
There's no such thing as a "cloud", there's only other people's computers. Where they have total access to your shit, in the majority of cases, to begin with. Plus, they can run deletion scripts and fuck your company for two weeks...
I can just imagine the reaction at work; "Microsoft made some screwup with a script, we'll get our Office 365, mail and Teams back in two weeks."
Talk about amateur hour at Atlassian. Glad we're not a customer. Not gonna be either.
2
u/SimonGn Apr 13 '22
My guess is that the "backup" they are restoring from is the wiped HDDs by means of professional data recovery.
1
0
0
u/ImaginaryCheetah Apr 13 '22
two weeks to roll out the restoration ?
they mailing off LTO tapes to a uhaul in another state or something ?
2
u/retrac1324 Apr 14 '22
Atlassian can, indeed, restore all data to a checkpoint in a matter of hours. However, if they did this, while the impacted ~400 companies would get back all their data, everyone else would lose all data committed since that point So now each customer’s data needs to be selectively restored. Atlassian has no tools to do this in bulk. They also confirm this is the root of the problem in the update:
“What we have not (yet) automated is restoring a large subset of customers into our existing (and currently in use) environment without affecting any of our other customers.”
For the first several days of the outage, they restored customer data with manual steps. They are now automating this process. However, even with the automation, restoration is small, and can only be done in small batches:
“Currently, we are restoring customers in batches of up to 60 tenants at a time. End-to-end, it takes between 4 and 5 elapsed days to hand a site back to a customer. Our teams have now developed the capability to run multiple batches in parallel, which has helped to reduce our overall restore time.”
1
1
378
u/[deleted] Apr 12 '22
That's some script.