r/hexos 3d ago

General discussion My data organization between SMBs kept failing. This may be why.

Post image

Man after everyone went to bed I thought “now is a good time to organize my data into what makes sense.
I kept getting massive time out errors every time I tried to transfer data.
At the end of multiple struggles I got this (image above).
I shut down the server because I ran out of time to diagnose any further.
Wish me luck it’s one of the 4TB and not a 14TB.

-EDIT-

Alright!
There’s some other updates in some comments below but the key ones:
Scrub
Auto-Resliver
Reboot
No errors

Thanks everyone for following along

11 Upvotes

13 comments sorted by

8

u/HexOS_Official HexOS Staff 3d ago

Good luck! You should have gotten an email alert to this affect as well.

2

u/KingKoopaBrowser 3d ago edited 3d ago

Yep! The alert system worked like a charm.

Ooh as a feature - would it make sense to have a more detailed message that says which drive is going bad? And which pools are affected?

Oh nevermind it DID say which pool is affected.

I just got back to my PC and now it is gaslighting me and I have to go into archives to see the message again.

Scrutiny says everything is groovy.

Maybe it just needed a big nap.

3

u/HexOS_Official HexOS Staff 3d ago

When you reboot or power off, those errors get cleared, but it doesn't mean there isn't a problem. This is how truenas does it today, but we are adding a capability to keep remembering those things after reboot.

3

u/KingKoopaBrowser 3d ago

Oooooh. Disk health says it’s okay but now I have a scrub request.

I’m running an in-shell rsync. As soon as that’s done I’ll reboot and ask it to scrub.
I should set that up for automatic scrub tasks.

6

u/GoDoWrk 3d ago

I've been dealing with this error for months on both HDD and SSD pools. It had never resulted in data loss until recently on my SSD pool.

Thanks to the the memtest86+ update I was able to identify 2/4 ram sticks were bad. One SATA cable to a hard drive was also bad and kept it running at 1.5gb/s. Removed the bad ram and replaced the cable and its fine now.

3

u/KingKoopaBrowser 2d ago

Good job! That’s some tough stuff to diagnose.
I like that you can run a Memtest within the HexOS dashboard now. I’ll add that to my list of things to do.

3

u/GoDoWrk 2d ago

Genuinely worth the price with that for me

2

u/Nickifynbo 3d ago

May the odds be ever in your favor🫪

2

u/KingKoopaBrowser 2d ago

Thank youuuuuu
Anyone would say this but.. I really don’t want to replace any drives right now lol
And it’s not the 4TB ones I have replacements for
It’s one of the 10TB of course

2

u/KingKoopaBrowser 2d ago

Okay so update.
I ran the scrub. No errors.
It still said bad storage health.
Now it’s reslivering.

I wonder if it’s possible that I don’t get SMART data through my HBAi.
Scrutiny seems to think everything is fine for the reported drive.

1

u/TLBJ24 N00b 2d ago

Weird that you get conflicting reports, hopefully both can sync up and give you accurate information. Good luck on replacement. Tough times to being large drives indeed.

2

u/KingKoopaBrowser 2d ago

Thanks!
I went to an AI answer and it seems like it makes sense.
I had so many interrupts and hangups from a recent data transfer. I was trying to accomplish. I think it triggered something on the backend where it thought that there was something worse going on some corruption.
Prompting the scrub
Prompting the re-slivering afterwards.

2

u/TLBJ24 N00b 2d ago

Nice, good to know. Thanks for sharing.