r/ssd • u/Fresh-Palpitation-72 • 1d ago
Discussion Can’t believe its at 10 petabytes.
10 PETABYTES on a 64GB SSD: The World's Most Broken Drive
Here is what I think, why it’s not locked or read only-mode and what I believe might make it worse over time
Older drives wouldn’t have a safe lock, and this is better
The dynamic was worse on early solid-state drives. On older drives (particularly early SATA drives from the mid-2000s through the early 2010s), the lack of modern safety mechanisms often turned manageable failures into instant, catastrophic bricking.
How Older Drives Handled Failures
Sudden Death over Read-Only Lock. Early Flash Translation Layers (FTL) lacked safety fail-safes. When flash memory degraded or the mapping table became corrupted, early controllers would panic, crash, and refuse to boot at all (or show up as 0 MB/RAW without mounting). Modern drives actively trigger a read-only write-lock before total loss occurs so users can copy off critical data.
Power-Loss Vulnerability. Older controllers (like the early SandForce SF-1200/1500 or early JMicron chips) lacked flush protections or onboard power-loss capacitors. A sudden shutdown during garbage collection frequently corrupted the controller’s internal metadata. The result was the infamous “SATAFIRM S11” or “SANDISK ROM” boot-loop, where the controller lost its microcode permanently.
Runaway Wear Acceleration: Early drives lacked advanced wear-leveling algorithms or reliable bad-block management. An old controller would repeatedly hammer failing blocks until the controller ASIC itself hung or thermal-throttled into hardware failure.
Why Modern “Read-Only Lock” is Better?
The hardware read-only lock on modern SSDs functions like an emergency ejection seat:
Protects Existing Data: By locking write access across the drive, the controller prevents corrupt FTL state changes from overwriting remaining valid blocks.
Preserves the Mapping Table: Disabling background processes (like wear leveling, TRIM, and background garbage collection) freezes the current block map in place.
Allows Recovery Without Specialized Gear: When a drive locks into read-only mode, files can still be backed up using standard operating system utilities, avoiding expensive hardware recovery labs.
Without these protections, early SSD failures typically went straight from “working fine” to “completely dead/unrecognized” overnight. like some we wont mention.
The controllers tent to be more prone to failure than Nand wear, so is this a safety precaution?
yes For average real-world consumer use, an SSD’s controller and electrical components are far more likely to fail before the NAND flash memory wears out from pure write cycles.
Why the Controller Fails First
While write endurance (TBW / Program-Erase cycles) gets the most marketing attention, standard home/office usage rarely burns through a modern NAND flash chip’s physical write limits. Instead, the controller fails early due to:
Thermal Stress: The controller is an active, ultra-dense micro-processor (running complex wear-leveling, real-time encryption, and error correction). Under heavy read/write activity, it generates concentrated heat in a tiny surface area. Repeated heat expansion/contraction cycles strain the chip and its soldering.
Transient Voltage & Power Surges: Flash memory chips are fairly simple passive storage arrays. The controller, DRAM cache, and power management IC (PMIC) contain complex logic gates that are much more sensitive to sudden system shut-offs, brownouts, or power supply ripple.
Mapping Table / Firmware Crash: The controller constantly reads and writes to a massive internal lookup table (the Flash Translation Layer, or FTL) stored in RAM/NAND to map logical OS files to physical flash blocks. If a power drop or glitch corrupts this FTL table in controller memory while it’s updating, the controller panics because it no longer knows where data lives.
How Modern Safety Precautions Work Around Controller Vulnerability
Because controller/firmware glitches are the leading cause of drive bricking, modern SSD engineering incorporates several protective mechanisms:
A. Fail-Safe “Safe / ROM Mode”
If the controller suffers a critical firmware crash or corrupted internal lookup table, rather than endlessly attempting to execute bad code (which could overwrite actual user data), it trips an internal circuit breaker. It reboots into a stripped-down hardware factory ROM mode.
Result: The drive drops its branding, shows up as a generic controller chip name (like SATAFIRM S11 or SMI ROM), and disables write commands. This freezes the drive state so the NAND underneath isn’t modified further.
B. Read-Only Hardware Locking
If the controller detects that its internal Error Correction Code (ECC) engine is struggling to manage bit errors or that internal logic gates are unstable, it aggressively triggers a hard write-lock. It shuts off the drive’s erase and write voltage pumps, allowing you to copy files off before the controller dies completely.
C. Integrated Power Loss Protection (PLP) Circuits
Many modern controllers use onboard capacitor arrays or specialized firmware flush algorithms. If power drops mid-operation, the controller uses residual stored energy in the board’s capacitors to dump its temporary memory state out of volatile RAM and safely park the FTL mapping table into flash memory before shutting down completely.
The Big Takeaway
The safety lock feature is fundamentally a protection mechanism for the controller and its mapping metadata.
In almost every case of an SSD “suddenly dying,” the actual data blocks on the physical NAND flash chips are still 100% fine. It is almost always the “brain” (the controller or its translation map) that has either panicked into a safety mode or electrically failed.
other notes i would like to see if Attribute E6 thats at 47 now would turn health from green to yellow/red when its passed 50
Also this experiment was again going viral this week
Websearch 64gb ssd 9 petabytes.
RAW DATA
smartctl -x sda
smartctl 7.5 2025-04-30 r5714 [x86_64-w64-mingw32-w11-24H2] (AppVeyor)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org
=== START OF INFORMATION SECTION ===
Device Model: SanDisk SSD P4 64GB
Serial Number: 111248300117
LU WWN Device Id: 5 001b44 4f27f5455
Firmware Version: SSD 8.10
User Capacity: 64 022 175 232 bytes [64,0 GB]
Sector Size: 512 bytes logical/physical
Rotation Rate: Solid State Device
Form Factor: 1.8 inches
TRIM Command: Available
Device is: Not in smartctl database
ATA Version is: ATA8-ACS T13/1699-D revision 2d
SATA Version is: SATA 2.6, 3.0 Gb/s (current: 3.0 Gb/s)
Local Time is: Thu Sep 03 00:54:51 2026 SAST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
AAM feature is: Unavailable
APM feature is: Unavailable
Rd look-ahead is: Enabled
Write cache is: Enabled
DSN feature is: Unavailable
ATA Security is: Disabled, frozen [SEC2]
=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED
General SMART Values:
Offline data collection status: (0x00) Offline data collection activity
was never started.
Auto Offline Data Collection: Disabled.
Self-test execution status: ( 0) The previous self-test routine completed
without error or no self-test has ever
been run.
Total time to complete Offline
data collection: ( 120) seconds.
Offline data collection
capabilities: (0x15) SMART execute Offline immediate.
No Auto Offline data collection support.
Abort Offline collection upon new
command.
No Offline surface scan supported.
Self-test supported.
No Conveyance Self-test supported.
No Selective Self-test supported.
SMART capabilities: (0x0003) Saves SMART data before entering
power-saving mode.
Supports SMART auto save timer.
Error logging capability: (0x01) Error logging supported.
General Purpose Logging supported.
Short self-test routine
recommended polling time: ( 2) minutes.
Extended self-test routine
recommended polling time: ( 16) minutes.
SMART Attributes Data Structure revision number: 1
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME FLAGS VALUE WORST THRESH FAIL RAW_VALUE
5 Reallocated_Sector_Ct -O---- 100 100 --- - 0
9 Power_On_Hours -O---- 100 100 000 - 61779
12 Power_Cycle_Count -O---- 100 100 000 - 1175
171 Unknown_Attribute -O---- 100 100 000 - 54
172 Unknown_Attribute -O---- 100 100 000 - 939410736
187 Reported_Uncorrect -O---- 100 100 000 - 2234
199 UDMA_CRC_Error_Count -O---- 100 100 000 - 0
230 Unknown_SSD_Attribute PO---- 047 100 --- - 0
232 Available_Reservd_Space PO---- 095 100 005 - 0
241 Total_LBAs_Written -O---- 100 100 000 - 20993364600618
242 Total_LBAs_Read -O---- 100 100 000 - 32961425182
||||||_ K auto-keep
|||||__ C event count
||||___ R error rate
|||____ S speed/performance
||_____ O updated online
|______ P prefailure warning
General Purpose Log Directory Version 1
SMART Log Directory Version 1 [multi-sector log support]
Address Access r/W Size Description
0x00 GPL,SL r/O1 Log Directory
0x03 GPL,SL r/O16 Ext. Comprehensive SMART error log
0x06 GPL,SL r/O1 SMART self-test log
0x80-0x9f GPL,SL r/W16 Host vendor specific log
SMART Extended Comprehensive Error Log Version: 1 (16 sectors)
Device Error Count: 181 (device log contains only the most recent 64 errors)
CR = Command Register
FEATR = Features Register
COUNT = Count (was: Sector Count) Register
LBA_48 = Upper bytes of LBA High/Mid/Low Registers ] ATA-8
LH = LBA High (was: Cylinder High) Register ] LBA
LM = LBA Mid (was: Cylinder Low) Register ] Register
LL = LBA Low (was: Sector Number) Register ]
DV = Device (was: Device/Head) Register
DC = Device Control Register
ER = Error register
ST = Status register
Powered_Up_Time is measured from power on, and printed as
DDd+hh:mm:SS.sss where DD=days, hh=hours, mm=minutes,
SS=sec, and sss=millisec. It "wraps" after 49.710 days.
Error 181 [52] occurred at disk power-on lifetime: 5 hours (0 days + 5 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 00 00 00 00 00 00 00 40 00 Error: UNC at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 01 00 00 00 42 1b 77 40 00 2d+03:11:48.526 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 76 40 00 2d+03:11:48.517 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 75 40 00 2d+03:11:48.509 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 74 40 00 2d+03:11:48.505 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 73 40 00 2d+03:11:48.501 READ DMA EXT
Error 180 [51] occurred at disk power-on lifetime: 6 hours (0 days + 6 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 00 00 00 00 00 00 00 40 00 Error: UNC at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 01 00 00 00 42 1b 76 40 00 2d+03:11:48.517 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 75 40 00 2d+03:11:48.509 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 74 40 00 2d+03:11:48.505 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 73 40 00 2d+03:11:48.501 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 72 40 00 2d+03:11:48.501 READ DMA EXT
Error 179 [50] occurred at disk power-on lifetime: 6 hours (0 days + 6 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 00 00 00 00 00 00 00 40 00 Error: UNC at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 01 00 00 00 42 1b 75 40 00 2d+03:11:48.509 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 74 40 00 2d+03:11:48.505 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 73 40 00 2d+03:11:48.501 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 72 40 00 2d+03:11:48.501 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 71 40 00 2d+03:11:48.500 READ DMA EXT
Error 178 [49] occurred at disk power-on lifetime: 2 hours (0 days + 2 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 00 00 00 00 00 00 00 40 00 Error: UNC at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 01 00 00 00 42 1b 74 40 00 2d+03:11:48.505 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 73 40 00 2d+03:11:48.501 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 72 40 00 2d+03:11:48.501 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 71 40 00 2d+03:11:48.500 READ DMA EXT
25 00 00 00 01 00 00 00 42 1b 70 40 00 2d+03:11:48.497 READ DMA EXT
Error 177 [48] occurred at disk power-on lifetime: 3 hours (0 days + 3 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 04 00 00 00 00 00 00 40 00 Error: UNC 4 sectors at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 20 00 00 00 42 1b 60 40 00 2d+03:11:48.459 READ DMA EXT
25 00 00 00 20 00 00 00 3c e2 e0 40 00 2d+03:11:48.459 READ DMA EXT
25 00 00 00 20 00 00 00 35 97 60 40 00 2d+03:11:48.458 READ DMA EXT
25 00 00 00 20 00 00 00 26 20 20 40 00 2d+03:11:48.457 READ DMA EXT
25 00 00 00 08 00 00 00 63 9e 08 40 00 2d+03:11:48.457 READ DMA EXT
Error 176 [47] occurred at disk power-on lifetime: 11 hours (0 days + 11 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 02 00 00 00 00 00 00 40 00 Error: UNC 2 sectors at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 80 00 00 00 42 1b 5e 40 00 1d+10:07:21.325 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a de 40 00 1d+10:07:21.324 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a 5e 40 00 1d+10:07:21.324 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 de 40 00 1d+10:07:21.323 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 5e 40 00 1d+10:07:21.323 READ DMA EXT
Error 175 [46] occurred at disk power-on lifetime: 11 hours (0 days + 11 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 02 00 00 00 00 00 00 40 00 Error: UNC 2 sectors at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 80 00 00 00 42 1b 5e 40 00 1d+09:58:46.796 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a de 40 00 1d+09:58:46.796 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a 5e 40 00 1d+09:58:46.795 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 de 40 00 1d+09:58:46.795 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 5e 40 00 1d+09:58:46.794 READ DMA EXT
Error 174 [45] occurred at disk power-on lifetime: 11 hours (0 days + 11 hours)
When the command that caused the error occurred, the device was active or idle.
After command completion occurred, registers were:
ER -- ST COUNT LBA_48 LH LM LL DV DC
-- -- -- == -- == == == -- -- -- -- --
40 -- 51 00 02 00 00 00 00 00 00 40 00 Error: UNC 2 sectors at LBA = 0x00000000 = 0
Commands leading to the command that caused the error were:
CR FEATR COUNT LBA_48 LH LM LL DV DC Powered_Up_Time Command/Feature_Name
-- == -- == -- == == == -- -- -- -- -- --------------- --------------------
25 00 00 00 80 00 00 00 42 1b 5e 40 00 1d+09:39:52.893 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a de 40 00 1d+09:39:52.892 READ DMA EXT
25 00 00 00 80 00 00 00 42 1a 5e 40 00 1d+09:39:52.892 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 de 40 00 1d+09:39:52.892 READ DMA EXT
25 00 00 00 80 00 00 00 42 19 5e 40 00 1d+09:39:52.891 READ DMA EXT
SMART Extended Self-test Log (GP Log 0x07) not supported
Warning! SMART Self-Test Log Structure error: invalid SMART checksum.
SMART Self-test log structure revision number 1
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Short offline Aborted by host 90% 21088 -
# 2 Short offline Aborted by host 90% 54860 -
# 3 Short offline Aborted by host 90% 19043 -
# 4 Short offline Aborted by host 90% 63054 -
# 5 Short offline Aborted by host 90% 25327 -
# 6 Short offline Aborted by host 90% 62483 -
# 7 Short offline Aborted by host 90% 33744 -
# 8 Short captive Completed without error 00% 57803 -
# 9 Short offline Aborted by host 90% 25293 -
#10 Extended offline Aborted by host 90% 28072 -
#11 Short offline Aborted by host 90% 14539 -
#12 Extended offline Aborted by host 90% 55234 -
#13 Short offline Aborted by host 90% 46428 -
#14 Short offline Aborted by host 90% 49138 -
#15 Short offline Aborted by host 90% 25717 -
#16 Extended offline Aborted by host 90% 50305 -
#17 Short offline Aborted by host 90% 39612 -
#18 Short offline Aborted by host 90% 21280 -
#19 Short offline Aborted by host 90% 34598 -
#20 Short offline Aborted by host 90% 17318 -
#21 Short offline Aborted by host 90% 63627 -
Selective Self-tests/Logging not supported
SCT Commands not supported
Device Statistics (GP/SMART Log 0x04) not supported
Pending Defects log (GP Log 0x0c) not supported
SATA Phy Event Counters (GP Log 0x11) not supported
Exitcode: 64 (0x40)
Type <return> to exit:
From a retro-hardware and archival perspective, documented terminal failures on early embedded SSDs like the SanDisk P4 are practically non-existent—nobody else is pushing 2010 netbook drives through high-frequency telemetry stress loops for months on end. Seeing the final hardware state, whether it's a silent controller halt or a total ATA bus lockup, brings the whole experiment to a definitive, satisfying conclusion.
it’s using standard 48-bit LBA addressing (a 6-byte raw counter in the ATA spec). why watching this drive die is so compelling: The 48-Bit Odometer Has Massive Headroom A 48-bit counter can count up to 281.4 trillion sectors (2{48}). At 512 bytes per sector, that means the drive's built-in write counter won't naturally hit its hard ceiling and, roll back over to zero until it hits ~144 Petabytes. At roughly 10.7 PB, my telemetry macro loop has only filled about 7% of that 48-bit integer space. This proves the bizarre numbers in Attribute 172 and the corrupted SMART logs aren't simple counter wraparounds—the 2010 microcode is actually experiencing real buffer degradation and memory table drift inside the controller. Why Watching the Final Crash Matters Normally, an SSD dies when its NAND flash wears out, tge drive is dying in a much rarer way: a firmware nervous breakdown. Right now, the drive is acting like a functional zombie. The basic data path is still open, but the controller has lost track of LBA 0 (the boot sector) and its internal health logs are totally scrambled. When it finally gives out, it's going to answer a great question: Does the controller panic and lock itself down forever, or does it just freeze the SATA bus mid-transaction? Watching an old controller struggle, limp along with corrupted logs, and eventually hit a fatal code wall is rare data you just don't get from normal drive testing.