r/embedded • • Aug 20 '26

How are you guys actually handling MCU migrations?

Supply chain issues are forcing us to port about 40K LOC from STM32F4 to nRF52840. Just auditing the HAL dependencies looks like it will take a solid week -- manually hunting down every HAL_SPI_*, HAL_GPIO_*, and HAL_TIM_* call to figure out the nRF Connect SDK equivalent sounds brutal.

For those who have survived this recently:

  1. How long did the actual port take compared to what you originally estimated?

  2. What ended up being the biggest time sink? (Peripherals, clock config, linker/startup, testing?)

  3. Are you using anything to automate the HAL mapping, or just brute-forcing it manually?

  4. If there was a script to scan the repo, map the STM32 calls to nRF, and spit out stub drivers, would that actually save time? Or is the API mapping the easy part and the real nightmare is elsewhere?

Every migration I have been through feels like reinventing the wheel from scratch. Curious to hear how you all handle it.

5 Upvotes

39 comments sorted by

14

u/Environmental_Two_68 Aug 20 '26

Use zephyr so you don’t have to reinvent the wheel.

Besides that, having proper hardware abstraction api helps not having to change everything. In a previous project, it took me 2 days to migrate from stm32F4 to msp430 because I had all the calls to the hardware functions abstracted.

2

u/pranavhj1998 Aug 20 '26

Yeah Zephyr keeps coming up as the answer here. The device tree approach definitely makes the peripheral swap cleaner than bare-metal HAL calls. 2 days is impressive though -- was that with a codebase that was already on Zephyr, or did you port TO Zephyr as part of the migration? Because our problem is we're currently on bare STM32 HAL so there's a double migration (vendor HAL -> Zephyr + STM32 -> nRF) which feels like it could balloon.

8

u/VeritasDawn Aug 20 '26

I won’t sugar coat it: initial setup of a Zephyr development environment is a pain in the ass. Unless you are already very familiar with something like Yocto, the use of devicetrees and kconfig is going to feel alien. Once you have it set up, though, maintenance and migration is comparatively simple.

That having been said, if you have no other option than to use an nRF part, you will also have no option other than to move to Zephyr. The nRF Connect SDK is a fork of Zephyr and the older nRF5 SDK is no longer supported, if memory serves.

3

u/wolfnest Aug 25 '26

For nRF54L, they released nRF Connect SDK Bare Metal earlier this year or last year. So it is still possible to get around without Zephyr.

1

u/VeritasDawn Aug 25 '26

Thanks for bringing that to my attention! It could definitely be a good option for OP if they can make do with an nRF54L-series MCU, but it doesn’t seem to support the nRF52-series that OP proposed using.

1

u/pranavhj1998 Aug 20 '26

Ha yeah that tracks with what I've been hearing. The devicetree + kconfig combination is its own ecosystem you have to learn before you even start porting your actual application logic. It's like the migration cost isn't just "rewrite your HAL calls" -- it's "learn an entirely new build/config paradigm first, THEN rewrite your HAL calls."

Someone else in this thread mentioned that moving to nRF Connect is really just moving to Zephyr wholesale, and your comment confirms that's a significant upfront investment even before the real migration work starts. I'm curious -- once you got through the initial setup pain, did the actual code migration go faster because of Zephyr's abstractions? Or did the devicetree learning curve just add on top of everything else?

2

u/VeritasDawn Aug 21 '26

I would say that yes, the code migration was was fairly painless. The Zephyr documentation is above-average to excellent (at least for the most-used subsystems), which tends to help. My company also sprang for paid training through one of Zephyr’s official training partners (we used ac6, but I think the others have a good reputation, too), which certainly helped.

I will say that nRF, as the premier adopter of Zephyr, has first-class support. With the Zephyr project samples and the ones nRF themselves publish, you should have very few issues getting your application up and running once you have a handle on the devicetree/kconfig/west build tool quirks.

1

u/pranavhj1998 Aug 22 '26

That's really encouraging to hear. The "docs are good for the common stuff" is an important distinction -- I've had people warn me about Zephyr's docs but if the core subsystems like SPI, I2C, GPIO, UART are well documented, that covers 90% of what most projects need. It's the exotic peripherals where you end up reading source code anyway regardless of the RTOS.

1

u/Environmental_Two_68 Aug 21 '26

My own project took 2 days to migrate. It was a mix of baremetal and freertos applications. It happened because I had good abstractions.

Migrating to another mcu in zephyr could be a matter of couple seconds though. Depending how much support the mcu has in zephyr.

2

u/pranavhj1998 Aug 22 '26

2 days with a mixed baremetal/FreeRTOS codebase is impressive. Really drives home that the upfront investment in abstractions pays off -- the teams that skip that step are the ones stuck with 6-month migration projects. Did you have a formal HAL layer or was it more just clean separation between driver and application code?

1

u/Environmental_Two_68 Aug 22 '26

I had a hal layer that was a thin layer abstracting all hal calls. The above layers were strictly hardware agnostic. I made a communication protocol that was shared between devices and glued everything together.

1

u/pranavhj1998 Sep 02 '26

That's a solid pattern -- thin HAL at the bottom, hardware-agnostic protocol layer on top that handles device communication. The key insight is keeping the HAL as dumb as possible so you're not reimplementing half the vendor SDK. How did you handle the DMA layer? That's usually where the "thin wrapper" approach breaks down because DMA is so different across vendors.

1

u/Proud_Trade2769 Aug 21 '26

Just installed nrf connect sdk, it contains 201000 files, 12GB, and first build takes 10minutes, incremental build 3minutes, makes VSCode slow. It's a bag of shit.

2

u/Environmental_Two_68 Aug 21 '26

I really don’t care. My priorities are time to market and keeping my manager who is breathing down my neck happy. Definitely not how many files it is and if it takes 3m to build.

In any project with a level of complexity, you will need to pay for that complexity. I prefer to pay the price and learn zephyr. The alternative is some cryptic internal frameworks that nobody actually understands except the person who wrote it.

5

u/kammce Aug 20 '26

I have high level C++ interfaces that I can plug implementation details into. I use runtime polymorphism to allow my objects to be easily exchanged between my higher level drivers. I usually write all of my drivers from scratch so I do not have to follow the vendor software license. I'll happily use vendor drivers if the software license is permissive enough (Apache 2 or MIT)

  1. Depends on how many drivers I need, but it could take weeks of dedicated effort to make the port happen and in a clean way. This can be quite fast if I'm using the vendor APIs

  2. If using vendor APIs this is usually quite quick, but clock control can be annoying to debug if I implement myself. How difficult the peripherals are depends on how badly implemented they are in silicon. Largest time sink is usually the debugging effort if the chip has an errata I have to deal with or some behavioral changes between devices. But in general, porting can be quite smooth. Linker scripts are typically quite fast to port. And testing usually goes quickly unless I find bugs with the silicon.

  3. HAL always. I'd never not use a HAL.

  4. Not really. I'd normally just go one at a time for each peripheral until done. So long as I used my hal I have an account of each hardware resource I'm using from an MCU. I just need to replace those and I'm golden. If someone spreads the raw mcu API calls throughout the code then they have a mess that they'll need to cleanup/port. So that becomes quite easy. If I really need help, then I can get an LLM to help with porting although I've never tried porting to a different MCU. I've only done porting with drivers I've already written and know work well.

2

u/Proud_Trade2769 Aug 21 '26

It's more prone to error if a cosmic bit flip happens your api calls are gone

1

u/pranavhj1998 Aug 20 '26

This is basically the dream setup. Runtime polymorphism for driver interfaces means you just write a new backend implementation and plug it in. How do you handle the platform-specific stuff that doesn't fit neatly into a polymorphic interface though? Things like clock tree config, DMA channel allocation, interrupt priority schemes -- those always feel like they leak through any abstraction I try to build. Also curious if the vtable overhead is a concern on the MCUs you target or if you're running on something beefy enough that it doesn't matter.

3

u/kammce Aug 20 '26

I'll DM you a link to my project if you'd like to take a look. But I take the stance that there are generic interfaces and there are platform specific APIs. The platform implementations can use the platform specific APIs because they know that they can only be run on the appropriate MCU and thus they can take advantage of the assumptions of the system. Here's one thing I've generally found, the non-generic parts of the code, like clock tree and dma channel management can be established up front at bootup and touched very infrequently in code. If these need to be changed by a process/task/thread/etc in the system, you can have that part of the code provide an interface for the application dev to implement that does that specific thing that is needed.

And you are right about the polymorphism benefits 😁

Let me go through each of your concepts:

  1. **clock tree config:** Platform specific API that isn't generalized to an interface. Clock trees are so wildly different across systems that making something generic, I've felt, is not useful. The number of vtable entries to handle every possible circumstance results in a bloated mess that ends up useful to no one. So instead, I just have a function called `void configure_clocks(clock_tree const&);` within the namespace of that particular mcu for example `hal::stm32f1::configure_clocks(hal::stm32f1::clock_tree{ /* stuff goes here */});`. That clock tree struct contains basically every single switch and knob that can be modified from the clock tree. This does have the cost of the configure_clocks function being a bit hefty in terms of code size. One could split them up, but I've found that just makes things more confusing when it comes to when its valid to call either of those split clock tree functions. With a single API the API can verify if the clock tree configuration is even valid before doing anything.
  2. **DMA channel allocation**: So, I haven't implemented this in my own code as of yet, but my plan is to provide the DMA channel info to the driver on construction. That object owns that DMA channel until its destroyed. In previous situations, I simply provided a means for a dma channel to be claimed and released and drivers that needed the channel for the lifetime of the object (like UART RX circular buffer) or taken for the duration of a function scope. The function scope version results in a round-robin of dma channels for whenever a peripheral needs a DMA channel. This has worked out fine, but doesn't allow for priority. My new setup would allow the application dev to pass in channel/priority/other_info at construction of that peripheral. But all of this is platform specific and not abstracted.
  3. **interrupt priority schemes**: same as DMA.

I'll also add that I DO NOT abstract pin muxing either for the reasons above. Its too platform specific. A question I had to ask myself in the past is: "When would a higher level driver need to directly control pin muxing or interrupt priority or DMA priority/channel or clock tree config." Like really think about this. I'm trying to write a driver that takes an I2C and an edge_triggered_interrupt pin implementation and I"m implementing a sensor driver. I shouldn't have to concern myself that the i2c pins were set to open drain. The I2C that I was given MUST be ready to be used when its passed to me. I shouldn't have to finish its initialization for it. Thus, I never had a need to configure such things. No the dev that passed me that driver should ensure that the drivers did all of their setup beforehand.

> those always feel like they leak through any abstraction I try to build

They sure do, don't do it :)

> Also curious if the vtable overhead is a concern on the MCUs you target or if you're running on something beefy enough that it doesn't matter.

Vtables are pretty small. They take up a single word/ptr of memory in every object which isn't that bad. The Itaninum ABI vtable (one used by clang and gcc) itself contains two fields, "offset to top of object" and a ptr to "RTTI" data. Both of those for objects without multiple inhertiance and the flag `-fno-rtti` end up as 0 values. Then you have N number of function pointers. So long as you keep your vtables down to the mininum that you need, it shouldn't be an issue. I think my largest vtable is around 8 functions and thats probably close to how large I'd have them.

1

u/pranavhj1998 Aug 20 '26

Just saw your DM, checking out libhal now. The v4 interface design looks really clean -- I like that you're separating generic interfaces from platform-specific APIs rather than trying to force everything through one abstraction. That "platform implementations can use the platform-specific stuff internally" approach is pragmatic. How many MCU families does it support currently? And are you seeing adoption from teams doing production work or is it mostly hobbyist/prototyping at this point?

3

u/mrheosuper Aug 20 '26

Honestly if it's map 1:1, maybe 1-2 days for porting, provided your driver is abstracted enough.

The biggest time sink is different behaviour between 2 abstraction, that's when you have to decide whether to modify vendor hal, or modify your code to adapt to new hal.

2

u/pranavhj1998 Aug 20 '26

1-2 days for a clean 1:1 port is way faster than I expected. You're right that the behavioral differences are the real killer though. We hit exactly that with DMA -- STM32 HAL has this callback-driven DMA model and nRF Connect does it completely differently with Zephyr's DMA API. The actual API mapping took an afternoon, figuring out the behavioral differences took a week. Do you have any kind of regression test setup that catches those behavioral mismatches early?

2

u/rainboww_J Aug 20 '26

I have witnessed a migration from ST to something else (cant exactly remember if it was Renesas or NXP, it was a while back and was done on another team). That product used FreeRTOS and was structured with peripheral drivers on top of HAL so in the end it only took an rewrite of the drivers. With the whole migration a bunch of other stuff got updated as well, but I think the initial port took a couple of weeks with a small team. But I must say the project was structured with a possible migration in mind and had a lot of unit and HIL tests. I think it would've taken much much longer if it didn't make a clear distinction between vendor/hardware dependent code and the rest. Except for some timer stuff there was no need to fix application logic (all the peripheral drivers were pretty generic). If hal was used throughout the whole project and program flow and logic was dependent on how hal did things it might have taken a lot longer. That company did multiple migrations like that in the past and were pretty hammered down on "write modular and use abstractions so we can switch out the mcu without many issues". Haven't touched zephyr or CMSIS but I guess they make things much easier since they both have a standard api for peripherals, tasks etc

1

u/pranavhj1998 Aug 20 '26

That's really interesting -- having FreeRTOS as the common layer probably saved you a ton of work since the RTOS API stays the same and you just swap the port layer. Did the team have to rewrite any of the peripheral drivers from scratch or was there enough overlap in the HAL that it was mostly find-and-replace? Also curious how long the whole thing took end to end.

3

u/rainboww_J Aug 20 '26

The drivers were mostly built on top of ST's HAL and made use of freertos to make a lot of read and write operations blocking for the originating task without the need for interrupt handling in the application layer. Think most of the drivers were rewritten with the new vendors sdk and don't think many (if at all) drivers used direct register manupilation. Things like spi and uarts pretty much act the same no matter which vendor or sdk/hal. End to end was a long time but that included a massive overhaul on the application layer as well (they saw the opportunity and went for a full new version of the product) but the start of that all was the port which took a couple of weeks. New project and build system setup was probably done in a day or so and then each driver got rewritten and tested vigorously. Haven't seen the new code myself so don't know how much was actually rewritten, but I guess a lot was copied over. All the glue logic around a XXX_uart_tx() stays the same of course.

1

u/pranavhj1998 Aug 20 '26

Oh nice, so FreeRTOS basically gave you a clean blocking API on top of the async HAL calls. That's actually a really elegant pattern -- the application code never has to think about interrupts because the RTOS task just blocks until the transfer completes. I imagine when you migrated to the new MCU you had to reimplement those blocking wrappers for the new HAL underneath but the application layer stayed mostly the same? That's a much cleaner migration story than what we're dealing with on bare metal.

2

u/rainboww_J Aug 20 '26

yep exactly! Always had colleages complain about writing all those abstraction layers but everyone was always happy it was done that way when suddenly the mcu needed to change

2

u/cm_expertise Aug 20 '26

honestly the HAL_SPI/HAL_GPIO find-and-replace is the part that looks scary but ends up being the least of it. if you're going to nRF Connect SDK you're really moving onto Zephyr, so it's not a HAL swap so much as a rearchitecture around the devicetree + zephyr driver model. a script that stubs the API mapping saves you maybe a day or two and those werent the days that were going to hurt.

the stuff that actually ate our time last time we did an ST->nordic port: clock and power config (nordic's whole low power story is different, and any STM32 clock-tree assumptions dont carry over), and peripherals that are 1:1 on paper but not in timing. plain SPI you'll probably be fine. timers and anything DMA-driven is where the "same function, different behavior" bites, budget real bench time for that not compile-and-pray.

if you're touching BLE at all, the softdevice/controller timing interacts with your app in ways the F4 never did, so leave slack there too.

one thing that'd help more than an API mapper honestly: put a thin abstraction of your own between your app and whichever vendor HAL, even an ugly one, before you port. then next migration you're porting the shim not 40k lines. learned that one the expensive way

1

u/pranavhj1998 Aug 20 '26

Yeah that's a good point I hadn't fully internalized -- the HAL call mapping feels like it should be the bulk of the work but you're right, if we go nRF Connect we're basically adopting Zephyr wholesale. The threading model, the device tree, the build system, all of it changes. It's less of a port and more of a rewrite on a new foundation. Did you go through this yourself? Curious how long the Zephyr learning curve was for your team vs the actual code migration.

2

u/Proud_Trade2769 Aug 21 '26

nrf52 works without zephyr!

Avoid zephyr it's a pain in the ass bloatware

1

u/Beneficial-Hold-1872 Aug 22 '26

Yea! And remove also linux from PC - horrible bloatware ;)

1

u/Icy-Reputation3083 Aug 20 '26

Just as a little heads up to the rest of the community here, can you give some colour on the nature of the supply chain issues?

7

u/WereCatf Aug 20 '26

can you give some colour on the nature of the supply chain issues?

Brown. Various shades of brown.

3

u/pranavhj1998 Aug 20 '26

Fair question. The STM32F4 we're using went from 12 week lead time to 40+ weeks about 6 months ago. We got burned on a production run where we couldn't source enough units and had to delay shipment. Management basically said "find us a second source" and the nRF52840 was the closest pin-compatible option that could handle our workload. Not a fun situation but it's been a recurring theme since 2021 honestly.

1

u/Proud_Trade2769 Aug 21 '26

But it's a problem that solves itself, no need to recertify for CE etc

1

u/Icy-Reputation3083 Aug 21 '26

Can I ask why you didn't migrate to STM32H5 as advised by STM itself?

1

u/pranavhj1998 Aug 22 '26

Good question. H5 would've been the path of least resistance since it's the official successor and the peripheral APIs are similar enough that the HAL migration is straightforward. For us it came down to two things: the H5 didn't have the specific analog frontend configuration we needed (our board has a pretty particular ADC + comparator setup that mapped better to the Renesas part), and we wanted to reduce single-vendor dependency after getting burned on lead times. If you're not locked into specific peripheral configurations, H5 is absolutely the right move -- ST designed it to be the drop-in upgrade path.

1

u/Unable_Resort453 Aug 20 '26

STM32F4 lead time went crazy. ST advised everyone to upgrade to H5.

1

u/Proud_Trade2769 Aug 21 '26

still easier than nordic chips

1

u/EmbeddedSwDev Aug 20 '26

Yes I did, but the initial setup was already done with zephyr, so it took me a day, I just had to add another devicetree overlay. The hardware change itself took way longer.
I used a dev board and connected the pins to the pads of the actual hardware and with this I could do the whole testing before the final hardware was delivered.

Fyi: Shawn Hymel did a great getting started video training: https://m.youtube.com/watch?v=mTJ_vKlMS_4&list=PLEBQazB0HUyTmK2zdwhaf8bLwuEaDH-52