Voice-enabled devices in noisy environments face a persistent challenge: delivering clear speech input without relying on cloud processing or dedicated DSP hardware. embedUR deployed a lightweight neural network on the STM32N6 that filters background noise in real time using only on-chip computing and memory. This eliminates the cost and complexity of traditional signal processing components while keeping latency low enough for responsive, real-time interaction.
Application domains
- Smart home appliances (ovens, vacuums, washing machines)
- Consumer electronics (smart speakers, wearables, remote controls)
- Industrial and commercial intercoms or access control systems
- Healthcare monitoring devices
- Automotive infotainment or driver assistance systems
Approach
In consumer electronics and industrial voice interfaces, background noise from machinery, traffic, or environmental sounds significantly degrades speech recognition accuracy. Products such as smart appliances, intercoms, and wearables need clear audio input, but traditional solutions each carry significant trade-offs.
Dedicated DSP chips add hardware cost, power consumption, and integration complexity. Cloud-based speech enhancement requires internet connectivity, introduces latency, and raises privacy concerns. Traditional filtering techniques on MCUs offer limited effectiveness in dynamic, multi-source noise environments and generalize poorly to real-world conditions.
embedUR addressed this by deploying a compact neural network directly on the STM32N6 microcontroller. The model was quantized and optimized to run entirely within available on-chip memory, delivering real-time noise suppression without any additional hardware.
Key benefits of this edge AI approach:
- Reduced cost: Eliminates the need for a dedicated DSP chip or cloud subscription
- Ultra-low latency: Real-time frame-by-frame processing ensures responsive user interaction
- Power efficiency: Runs within the ultra-low-power profile of STM32N6, ideal for battery-powered or always-on devices
- Enhanced accuracy: Improves keyword detection rates across real-world noise conditions (factory, traffic, indoor conversation)
- Full offline capability: Ideal for privacy-focused or connectivity-limited environments
This solution demonstrates how embedded AI unlocks new capabilities for audio-centric devices, bringing smart, responsive, and cost-efficient noise suppression to even the smallest products.
Application overview
The audio denoising pipeline captures noisy signals, processes them through a neural network, and outputs clean audio in real time, all on-device.
1. Signal acquisition and transformation
Audio is captured from the microphone input using a ring buffer implementation that ensures continuous frame processing. Raw audio frames are processed through a spectral transformation pipeline where the time-domain signal is converted to frequency domain using Short-Time Fourier Transform (STFT).
2. Neural network noise identification
The frequency-domain representation is normalized and scaled to match the neural network input requirements. The prepared spectrogram is sent to the neural network running on the NPU, which identifies noise components in the audio and outputs a noise prediction mask.
3. Signal reconstruction and enhancement
The predicted noise is subtracted from the original signal in the frequency domain. This denoised spectral representation undergoes post-processing including phase reapplication and inverse transformation to reconstruct the time-domain signal. Additional enhancement filters improve voice clarity and intelligibility. The cleaned audio is then sent to the output device for real-time playback.
Technical details
The solution leverages deep learning together with traditional digital signal processing techniques. It implements a neural network-based spectral noise suppression system on STM32N6 hardware, utilizing the NPU for efficient inference.
- Real-time processing: Frame-by-frame denoising with minimal latency, suitable for live communications
- High noise reduction performance: Achieves noise suppression while preserving speech intelligibility and naturalness
- Low resource requirements: Optimized for embedded systems with limited memory and processing power
- Adaptive noise handling: Works effectively across stationary and non-stationary noise types
- Voice enhancement: Additional DSP techniques improve voice clarity beyond noise removal alone
- Energy efficiency: Optimized for low power consumption, extending battery life in portable applications
Sensor
Digital MEMS microphone (U13) featured on the STM32N6570-DK development board. The MEMS microphone captures audio input while the audio jack (CN15) provides denoised audio output for the real-time processing demonstration.
Model and results
Model:
- UNet-based audio denoising neural network
- Optimized and quantized for on-device deployment via ModelNova (embedUR)
- Total epochs: 50 (46 on NPU hardware, 4 on CPU software)
Results:
- Flash (octoFlash): 1.863 MB of 4 MB (46.58%) used for weights
- CPU RAM: 1 MB fully utilized (100%) for activations
- NPU RAM: 1.344 MB total with 1.322 MB used (98.36%) for activations
- Preprocessing time: 12 ms
- Neural network inference: 19 ms
- Postprocessing time: 5 ms
- Total processing time: 36 ms
Software tools
- STM32Cube AI Studio, for neural network conversion and deployment
- ST Edge AI Core command-line interface
- STM32CubeIDE, for application development and debugging
- STM32CubeProgrammer, for flashing the artifacts to the target
Documentation
Resources
Optimized with STM32Cube AI Studio
STM32Cube AI Studio is a desktop tool designed to evaluate, optimize and compile neural network models for STM32 microcontrollers. It fully supports compilation for the Neural-ART Accelerator neural processing unit (NPU). It replaces the X-CUBE-AI in the ST AI product offering to cover new STM32 devices.
Most suitable for STM32N6 Series
The STM32 family of 32-bit microcontrollers based on the Arm Cortex®-M processor is designed to offer new degrees of freedom to MCU users. It offers products combining very high performance, real-time capabilities, digital signal processing, low-power / low-voltage operation, and connectivity, while maintaining full integration and ease of development.