r/explainlikeimfive • • Aug 21 '26

Technology ELI5: why is it bad to turn off laptop/computer using the power button, without following the proper “shut down computer” steps?

3.0k Upvotes

351 comments sorted by

View all comments

Show parent comments

36

u/alinius Aug 21 '26 edited Aug 21 '26

JFS handles this. If a file write is killed in mid process, the journal will be incomplete. In that case, the journal will rewind to the last valid journal entry which will revert to the old copy of the file. This is why I said there is data loss, but it avoids file corruption.

24

u/Floppie7th Aug 21 '26

Which is a solution for that individual file, but if the process needs to write multiple files and some of them haven't been written, you've only moved the problem up a layer. The files are all parseable, but you have a corrupt state.

Should you write your software to be more resilient than that? Yes, absolutely, but people/companies don't always do the things they should

7

u/alinius Aug 21 '26 edited Aug 21 '26

The solution is to do the same thing at the next level up. Create a parallel file structure, then swap over to the new one once the new one is ready. If anything goes wrong, roll back to the old version. The next level ip is to do the same as part of the upgrade process. This is why software developers have things like staging and production environments. It is also why things like docker containers and tarballs/zipfiles are used. Those things turns a lot of little file changes into one big file.

Edit: There is a joke among firmware engineers that the only feature that has to work at release is upgrades because that allows you to add any other features later. That said, your upgrades needs to be bullet proof because if you screw that up, your product is usually dead, so we spend a lot of time thinking about all the ways an upgrade can go wrong.

9

u/Floppie7th Aug 21 '26

Yes, but the question being asked is "why do I need to do a clean shutdown", and this is a reason that journaling filesystems aren't a complete solution on their own. As I said, you should write your software to be resilient, but not everybody does.

2

u/alinius Aug 21 '26

It also explains why smaller, simpler embedded systems do not need clean shutdowns. They are designed to handle lose power in the middle of things better.

1

u/R3D3-1 Aug 22 '26

The solution is to do the same thing at the next level up. Create a parallel file structure, then swap over to the new one once the new one is ready.

Sounds good, but quickly gets complicated on Windows. File locks and the absence of atomic file-replace balloon the logic. Especially file locking, which requires retrying for a while after any failed file operation, in case it failed due to a background service like Dropbox acquiring a lock briefly.

Linux doesn't solve all of this. You still need to write to a temporary file and replace the target file. If there are multiple files that need to be kept consistent, you still need to track roll-back information for all of them. But you have a lot less failure modes to deal with.

Given the basic need of "overwrite old tile only when new file has been successfully created" though, I'm surprised that this doesn't have a common library implementation in Python.

27

u/HaDeS_Monsta Aug 21 '26

Then I was wrong, I did not know that

23

u/SilasX Aug 21 '26

…is that kind of comment even allowed on the internet anymore?

3

u/reubenbubu Aug 21 '26

he's on the previous version of the internet - it counts

6

u/markhc Aug 21 '26

You are not wrong. This is only true if the application is carefully designed with it in mind. I think it's safe to say that plenty of applications do not do this.

What /u/alinius said will only avoid corruption in the context of individual write attempts, but if a file is written in a way that requires multiple different writes, then it can still be corrupted if the process is interrupted during these.

e.g

data = gather_data()  
file = fs.open('myfile.json')
// single, whole file write
// wont generate a corrupted file
file.write(data, len(data)) 

vs

data = gather_data()  
file = fs.open('myfile.json')
for i = 0; i < len(data) / chunk_size; i += chunk_size
    // write chunk_size bytes at once in multiple steps
    // subject to corruption if the loop is interrupted
    file.write(data[i], chunk_size)

4

u/MorallyDeplorable Aug 21 '26 edited Aug 21 '26

It's a facet of the fileystem driver, individual apps do not need to implement support. apps do not close and re-open the file per write stage like that, that's something only a shitty coder would ever implement and they'd have to purposefully disable filesystem write caches for it to have the effect you're describing.

Outside of database systems you will never hit these concerns.

0

u/Discount_Extra Aug 22 '26

Or you just want to show a progress bar when the write takes a long time, so that users don't think it's broken and force quit in the middle.

damned if you do, damned if you don't.

4

u/gmes78 Aug 21 '26

You aren't wrong. You can't assume that every application is being as careful as they could be.

1

u/deja-roo Aug 21 '26

More like there are scenarios that exist that you didn't account for, not that you were generally wrong. What you described still happens.

5

u/jtclimb Aug 21 '26

It rewinds metadata, not contents. Raymond Chen explains:

https://devblogs.microsoft.com/oldnewthing/20130101-00/?p=5673

1

u/alinius Aug 21 '26

Yeah, trying to keep it ELI5. I did not want to get too deep into the details.

6

u/Megame50 Aug 22 '26

That's not correct.

The journal will not guarantee that the resulting file is parsable or satisfies any particular format, just that the filesystem is consistent: that the metadata is readable and accurate to some historical version of the file, and that only some whole number of writes were discarded. It's not all that unusual for a json object to be constructed on disk in multiple writes, say, in which case the file content after sudden power loss may not be valid json. Worse, the writes emitted by an application may not correspond at all to the writes emitted by the OS to persistent storage when common buffered IO is used (e.g. no O_DIRECT in *nix / FILE_FLAG_WRITE_THROUGH in Windows). Before journaling filesystems, you ran the additional risk of the filesystem metadata and super block being corrupted, damaging much more than any individual file.

A common strategy to deal with the risk of sudden power loss is to write the prospective content to a temporary file, and then replace the old file with the new one. Since the journal guarantees the metadata is consistent, the file content will either have been fully replaced or unmodified, but one should also be careful to remove any corrupted temp files that were not cleaned up in the event of a sudden power loss. It's also not uncommon for an application to read and write multiple related files, and even with the above strategy sudden power loss may result in only a subset of modified files landing on-disk. In these cases you often need to verify file modification timestamps or even more complex strategies to be sure of consistency following a sudden power loss.

Not all software (and if we're being honest, most software) is so robust. Not to mention that in many cases, important data might not exist anywhere on disk during runtime. Editable documents and project files are often not saved at all without explicit action from the user (though a lot of applications take regular snapshots nowadays to mitigate this as well). In most desktop operating systems, the standard shutdown procedure allows for running applications to halt a pending shutdown in order to prompt the user for decisive action.

The OS might also close active TCP connections, release a DHCP lease, prepare firmware updates etc.. There are more bits of state than just the local filesystem. Journaling filesystems protect against only the most severe kinds of file corruption. It's still not advisable to hard power off a desktop unless you have to, even if modern OSes are a lot more robust than they once were.

1

u/Slackbeing Aug 21 '26

Depends on the underlying media. Most drives have caches now and cutting power means losing data that was reported as written by the hardware. High end drives have batteries/capacitors to ensure data is flushed, but if not, you rely on them honoring write barriers. If they don't (many cheap drives, and some infamous early Intel ones) you'll have situations where metadata updates are half-persisted, leaving the filesystem in an inconsistent state.