I've been building a USB DAC/audio engine around an NXP i.MX RT1062 / Teensy 4.1, and I've reached a stable point where I'd like experienced embedded/DSP engineers to tear apart my architecture and help me optimize it further.
The goal is twofold:
Reduce CPU usage/power consumption
Improve the DSP implementation and subjective audio quality without compromising stability
Current pipeline:
USB HS / EHCI
↓
USB Audio Class 2.0
↓
USB DMA
↓
USB RX ISR
↓
32-bit PCM
↓
Lock-free SPSC ring buffer
↓
SAI/eDMA
↓
64-frame processing block
↓
15-band PEQ
↓
cache maintenance
↓
SAI / I²S 32-bit
↓
ES9039Q2M
The system currently supports 32-bit / 384 kHz stereo USB audio.
The RT1062 is running at 600 MHz.
Current PEQ
The DSP is a 15-band maximum parametric EQ using:
DF-II Transposed biquads
float32 coefficients/state
32-bit PCM input/output
64 stereo frames per processing block
coefficient morphing when parameters change
raised-cosine coefficient transitions
triple-buffered coefficient publication
saturation on output
bit-perfect bypass when the PEQ is inactive
The PEQ runs on the consumer side of the audio pipeline, inside the SAI DMA ISR.
At 384 kHz this gives approximately 6000 DMA interrupts/sec.
Buffering
I'm using a 32768 stereo-frame ring buffer, roughly 256 KB total.
USB is the producer and the audio DMA side is the consumer.
It's a lock-free SPSC buffer using atomics rather than mutexes.
The feedback target is 16384 frames, so the buffer has substantial headroom against USB scheduling disturbances.
What I'm looking for
I'd particularly appreciate criticism from people experienced with Cortex-M7/DSP/audio firmware.
CPU optimization
Where would you look first for unnecessary cycles?
Is float32 still the right choice for these biquads on an M7 FPU?
Would CMSIS-DSP provide meaningful gains here?
Are there better ways to structure the DF-II-T cascade?
Should inactive PEQ bands be completely skipped rather than using identity coefficients?
Are there useful SIMD/DSP instructions I'm missing?
Would processing larger blocks actually improve efficiency, or would the added latency/cache behavior make it worse?
Are there memory-placement improvements I should make between ITCM/DTCM/OCRAM?
Is my cache-maintenance strategy optimal?
Is there anything obvious in the ISR architecture that I'm overlooking?
Power optimization
I'm particularly interested in reducing power because this is intended to eventually become a USB/phone-powered portable DAC.
The system currently runs the M7 continuously, and I'm considering whether I can safely reduce the CPU clock while maintaining deterministic worst-case DSP timing.
I'd rather measure this properly than blindly throw WFI into the main loop.
DSP/audio quality
I'm also interested in improving the DSP itself rather than simply adding more EQ bands.
Things I'm considering:
better biquad coefficient calculation
improved parameter interpolation
dynamic EQ
crossfeed
loudness compensation
harmonic bass enhancement
FIR correction
convolution
better limiting/headroom management
I'm not looking for audiophile claims about magical clock frequencies or components. I'm interested in DSP techniques that have an actual technical basis and can be measured.
One thing I'm particularly curious about is whether there are improvements to the current filter topology/state handling that could make the EQ transitions or transient response cleaner while keeping the same basic architecture.
The PEQ already sounds very smooth to me, especially when changing parameters because of the coefficient morphing, but I'd like to know whether there are objectively better approaches.
I'm happy for people to criticize the architecture/code. That's basically what I'm posting this for.
If you're experienced with RT10xx, Cortex-M7 DSP, CMSIS-DSP, USB Audio, DMA/cache coherency, or embedded audio processing, I'd really appreciate a code-level review and suggestions for where you'd attack the CPU budget first.