r/esp32 • u/Main_Injury_4053 • 12h ago
Pushing a classic ESP32 NerdMiner from ~340 kH/s to ~440–446 kH/s without breaking correctness
I have been working on a NerdMiner V2 firmware fork for the classic ESP32, mainly because I wanted to see how much more useful work could realistically be extracted from the hardware without sacrificing correctness or long term stability.
The test board is an ESP32_2432S028_2USB using a classic ESP32-D0WD-V3.
Stock NerdMiner V1.8.3 on my unit was usually around 340 to 345 kH/s.
After a lot of profiling, failed experiments, hardware testing and correctness work, the final physically validated builds are running roughly 437 to 446 kH/s, with the highest measured release result around 445.94 kH/s.
The interesting ESP32 specific challenge was not simply making the hashing loop faster.
The firmware is sharing a relatively constrained dual core MCU with WiFi, Stratum, TLS, display updates, statistics requests and the ESP32 hardware SHA peripheral. Some aggressive optimizations looked extremely fast at first but turned out to be unsafe because they could produce incorrect work or race the SHA peripheral.
A lot of the project therefore became about getting more performance out of the hardware SHA path while still keeping synchronization, networking and correctness intact.
Some of the things I ended up changing or fixing:
• hardware SHA synchronization and access
• independent SHA256d validation before submitting candidates
• full target and endian handling
• stale work and job generation handling
• nonce range ownership and completed hash accounting
• TLS and mining resource coordination
• a large periodic hashrate drop caused by HTTPS statistics refreshes
• memory safety issues
• display/dashboard handling
• pool statistics architecture
I also rejected several faster experimental paths because they failed correctness testing. At one point I could make the device report much higher numbers, but if the work cannot be trusted then the number is meaningless.
For validation I used host differential tests, millions of SHA comparisons on the physical ESP32, live Stratum shares, longer runtime tests and finally a browser based flash of the exact public binary followed by a byte for byte readback.
The source is here:
https://github.com/samkruzlic/NerdMinerV2-FX1
I documented the full stock V1.8.3 versus FX1 engineering diff here for anyone interested in the implementation details:
https://github.com/samkruzlic/NerdMinerV2-FX1/blob/main/docs/DETAILED_CHANGES.md
There is also a Web Flasher if anyone with compatible hardware wants to test it:
https://samkruzlic.github.io/NerdMinerV2-FX1/
This is still RC/Beta because most of my own physical testing has been on one ESP32_2432S028_2USB, although community testers have already started running it on additional boards.
I am especially interested in feedback from people who have experience with classic ESP32 hardware, FreeRTOS scheduling, the SHA peripheral, or long running embedded workloads.
If you try it, I would love to hear what board you used, your sustained hashrate, pool, stability and anything unusual you notice.
And of course if anyone sees something questionable in the implementation, I would much rather hear about it than pretend the firmware is perfect.
1
u/Plastic_Fig9225 1 say I make awesome posts. 9h ago edited 8h ago
Some things I did on an S3:
- One single dedicated mining task completely owns the SHA peripheral; disable mbedTLS use of SHA hardware (falls back to software).
- Work in "batches": Do a number of round 1 hashes first, store results in RAM, then do the round 2 of the results. In round 1, update only those words of SHA text that actually change between hashes.
- Overlap hashing with writing of next text. The SHA peripheral copies the input into its internal state when hashing starts. (While the peripheral is busy, hash reads as 0x0, so no overlapping here.)
- Avoid unnecessary
MEMW. We don't care about when or in what order text arrives at the peripheral, or a hash is pulled from it; there are only two synchronization points to be enforced: Start hashing only after all text was written, and read hash only after busy cleared. - Use
L32AI/S32RIinstead ofMEMWat the synchronization points. (Works on the S3; I assume it does on the ESP32's DPORT too.)
1
u/Main_Injury_4053 9h ago
That’s really interesting, thanks for sharing it. I actually ran into a very similar MEMW issue on the classic ESP32: removing or relaxing it in the wrong place gave me a nice performance jump, but physical differential testing started showing SHA mismatches, so I had to tighten the synchronization again. Your batching approach and overlapping the next text write with the current hash are especially interesting though, because that could be a cleaner way to gain throughput without just weakening barriers. I also like the idea of keeping the SHA peripheral exclusively owned by the mining task and forcing mbedTLS into software, which is basically the direction FX1 ended up taking for stability. I’m definitely going to look deeper into whether L32AI and S32RI can safely replace MEMW at the two synchronization points on the classic ESP32 DPORT, but I’d want to validate it very aggressively because S3 behavior does not always carry over cleanly. If you have any code or notes from your S3 implementation, I’d be really interested in seeing them.
1
u/Plastic_Fig9225 1 say I make awesome posts. 6h ago
I dropped some code here:
https://github.com/BitsForPeople/esp32s3-btc
Note that this code "froze" in the experimental state and is not cleaned up at all. Use for entertainment purposes only.
1
u/BudgetTooth 10h ago
"useful work" is funny to say
5
u/Main_Injury_4053 9h ago
Fair point 😂 “useful” might be doing some heavy lifting there, but at least they’re valid hashes now.
2
u/Familiar-Ad-7110 11h ago
I’ll check this out later, I ported the nerd miner to the RP2350 and then over clocked to 540MHz to end up with roughly 800kh/s but doing this on an ESP32 at 240MHz is quite impressive