Designers of compact vision-based products increasingly need more than simple motion detection; systems must understand who and what is present in a scene while remaining cost-efficient, energy-efficient, and privacy-compliant. embedUR deployed an optimized instance segmentation model directly on the STM32N6, delivering real-time person detection and pixel-level segmentation at 14 FPS, entirely at the edge, with no cloud connectivity or external accelerators required.
Application domains:
- Smart home cameras
- Industrial safety monitors
- Retail analytics systems
- Robotics and drones
- Interactive consumer devices (mirrors, displays...)
Approach
A customer developing a compact AI vision product for indoor monitoring wanted to add scene context and visual intelligence to their system. The goal was to identify and isolate regions in a video feed, such as people, background, and objects, to enable smarter automation like targeted lighting, security triggers, or device interactions. The solution needed to run directly on a microcontroller due to cost and form factor constraints.
Traditional approaches each carry significant limitations. Server or cloud-based segmentation delivers high-resolution outputs but requires connectivity, introduces privacy risks, and incurs ongoing cloud costs. PIR or motion sensors consume very little power but lack visual understanding and cannot detect scene type or object class. High-power AI accelerators are fast and robust but too expensive and large for constrained embedded form factors.
By quantizing and optimizing a semantic segmentation model to run on STM32N6, embedUR achieved real-time visual understanding in a compact and power-efficient format. The model provides pixel-level classification directly from a low-power camera feed.
Key benefits of this edge AI approach:
- No external accelerator: The full segmentation pipeline runs on the STM32N6 alone, removing extra silicon from the bill of materials
- Real-time inference: Under 100 ms per frame on-device
- Pixel-level scene understanding: Detects the presence, position, and outline of each person, not just motion
- Power-efficient design: Fits the constraints of battery-operated and always-on products
- Compact integration: Enables camera modules in tight spaces
- Privacy by design: Video frames are processed locally and never leave the device
Application overview
The system captures live video frames from a camera sensor connected to the STM32N6 microcontroller and processes them through a real-time segmentation pipeline.
1. Frame capture and buffering
Each video frame is duplicated in memory. One copy feeds the display pipeline for real-time visualization using double-buffered rendering to ensure smooth, flicker-free output. The second copy enters the AI processing path.
2. Preprocessing
The frame is downscaled and preprocessed to 224×224 resolution, the format required by the neural network for inference.
3. Instance segmentation inference
The preprocessed frame is passed to the YOLACT neural network deployed on the device. The model performs two key operations: object detection (generating bounding boxes and confidence scores) and instance segmentation (computing a unique mask for each detected person using learned prototype masks and per-instance coefficients).
4. Post-processing and display
Non-maximum suppression (NMS) refines detections, and masks are combined and thresholded to isolate individuals from the background. The detected person count and segmented masks are overlaid on the original frame and rendered to the display in real time.
Technical details
The solution is based on YOLACT, a real-time instance segmentation neural network optimized for embedded deployment. Unlike traditional pixel-by-pixel segmentation networks, YOLACT separates detection and mask generation by using a set of global prototype masks combined with per-instance coefficients. This architecture significantly reduces computational complexity while preserving segmentation accuracy.
The image above illustrates the person detection and segmentation process. Once individuals are detected in the frame, their presence is validated, and a segmentation mask is generated for each one. The total number of persons identified in the frame is also displayed in real time.
Using ST Edge AI tools, the model is quantized and deployed on the STM32N6 series microcontroller, enabling the full AI pipeline to run locally.
- Real-time performance: Maintains accuracy even with multiple people present in the frame
- High segmentation accuracy: Lightweight prototype-based mask generation technique
- Low-latency processing: Runs entirely on-device without cloud data transfer
- Smooth display output: Double-buffered rendering ensures a seamless user experience
Sensor
The application uses the MB1854B camera module built around the IMX335 CMOS image sensor. For instance segmentation, image quality directly affects mask precision: consistent color fidelity and sharp edges help the network resolve clean boundaries between people and background, while stable exposure across frames keeps mask outlines steady when several individuals move through the scene. The IMX335 delivers this level of detail together with dependable low-light behavior, so the RGB frames feeding the model stay reliable across varied indoor lighting conditions.
Model and results
Model:
- YOLACT (instance segmentation)
- Input size: 224 x 224
- Optimized and quantized for STM32N6 deployment via ModelNova (embedUR)
Results:
- Inference time: 61 ms
- Post-processing time: 1 ms
- FPS: ~14
- Deployment: Fully on-device, STM32N6 series
Software tools
- STM32Cube AI Studio, for neural network conversion and deployment
- ST Edge AI Core command-line interface
- STM32CubeIDE, for application development and debugging
- STM32CubeProgrammer, for flashing the artifacts to the target
Resources
Optimized with STM32Cube AI Studio
STM32Cube AI Studio is a desktop tool designed to evaluate, optimize and compile neural network models for STM32 microcontrollers. It fully supports compilation for the Neural-ART Accelerator neural processing unit (NPU). It replaces the X-CUBE-AI in the ST AI product offering to cover new STM32 devices.
Most suitable for STM32N6 Series
The STM32 family of 32-bit microcontrollers based on the Arm Cortex®-M processor is designed to offer new degrees of freedom to MCU users. It offers products combining very high performance, real-time capabilities, digital signal processing, low-power / low-voltage operation, and connectivity, while maintaining full integration and ease of development.