Designers of compact vision-based products increasingly need more than simple motion detection; systems must understand who and what is present in a scene while remaining cost-efficient, energy-efficient, and privacy-compliant. embedUR deployed an optimized instance segmentation model directly on the STM32N6, delivering real-time person detection and pixel-level segmentation at 14 FPS, entirely at the edge, with no cloud connectivity or external accelerators required.

 

Application domains:

  • Smart home cameras
  • Industrial safety monitors
  • Retail analytics systems
  • Robotics and drones
  • Interactive consumer devices (mirrors, displays...)

Approach

A customer developing a compact AI vision product for indoor monitoring wanted to add scene context and visual intelligence to their system. The goal was to identify and isolate regions in a video feed, such as people, background, and objects, to enable smarter automation like targeted lighting, security triggers, or device interactions. The solution needed to run directly on a microcontroller due to cost and form factor constraints.

 

Traditional approaches each carry significant limitations. Server or cloud-based segmentation delivers high-resolution outputs but requires connectivity, introduces privacy risks, and incurs ongoing cloud costs. PIR or motion sensors consume very little power but lack visual understanding and cannot detect scene type or object class. High-power AI accelerators are fast and robust but too expensive and large for constrained embedded form factors.

 

By quantizing and optimizing a semantic segmentation model to run on STM32N6, embedUR achieved real-time visual understanding in a compact and power-efficient format. The model provides pixel-level classification directly from a low-power camera feed.

 

Key benefits of this edge AI approach:

 

  • No external accelerator: The full segmentation pipeline runs on the STM32N6 alone, removing extra silicon from the bill of materials
  • Real-time inference: Under 100 ms per frame on-device
  • Pixel-level scene understanding: Detects the presence, position, and outline of each person, not just motion
  • Power-efficient design: Fits the constraints of battery-operated and always-on products
  • Compact integration: Enables camera modules in tight spaces
  • Privacy by design: Video frames are processed locally and never leave the device

Application overview

The system captures live video frames from a camera sensor connected to the STM32N6 microcontroller and processes them through a real-time segmentation pipeline.

EmbedUR segmentation application flow EmbedUR segmentation application flow EmbedUR segmentation application flow

1. Frame capture and buffering

Each video frame is duplicated in memory. One copy feeds the display pipeline for real-time visualization using double-buffered rendering to ensure smooth, flicker-free output. The second copy enters the AI processing path.

2. Preprocessing

The frame is downscaled and preprocessed to 224×224 resolution, the format required by the neural network for inference.

3. Instance segmentation inference

The preprocessed frame is passed to the YOLACT neural network deployed on the device. The model performs two key operations: object detection (generating bounding boxes and confidence scores) and instance segmentation (computing a unique mask for each detected person using learned prototype masks and per-instance coefficients).

4. Post-processing and display

Non-maximum suppression (NMS) refines detections, and masks are combined and thresholded to isolate individuals from the background. The detected person count and segmented masks are overlaid on the original frame and rendered to the display in real time.

Technical details

The solution is based on YOLACT, a real-time instance segmentation neural network optimized for embedded deployment. Unlike traditional pixel-by-pixel segmentation networks, YOLACT separates detection and mask generation by using a set of global prototype masks combined with per-instance coefficients. This architecture significantly reduces computational complexity while preserving segmentation accuracy.

EmbedUR segmentation on STM32N6 EmbedUR segmentation on STM32N6 EmbedUR segmentation on STM32N6

The image above illustrates the person detection and segmentation process. Once individuals are detected in the frame, their presence is validated, and a segmentation mask is generated for each one. The total number of persons identified in the frame is also displayed in real time.

 

Using ST Edge AI tools, the model is quantized and deployed on the STM32N6 series microcontroller, enabling the full AI pipeline to run locally.

 

  • Real-time performance: Maintains accuracy even with multiple people present in the frame
  • High segmentation accuracy: Lightweight prototype-based mask generation technique
  • Low-latency processing: Runs entirely on-device without cloud data transfer
  • Smooth display output: Double-buffered rendering ensures a seamless user experience

Sensor

The application uses the MB1854B camera module built around the IMX335 CMOS image sensor. For instance segmentation, image quality directly affects mask precision: consistent color fidelity and sharp edges help the network resolve clean boundaries between people and background, while stable exposure across frames keeps mask outlines steady when several individuals move through the scene. The IMX335 delivers this level of detail together with dependable low-light behavior, so the RGB frames feeding the model stay reliable across varied indoor lighting conditions.

Model and results

Model:

  • YOLACT (instance segmentation)
  • Input size: 224 x 224 
  • Optimized and quantized for STM32N6 deployment via ModelNova (embedUR)

Results:

  • Inference time: 61 ms
  • Post-processing time: 1 ms
  • FPS: ~14
  • Deployment: Fully on-device, STM32N6 series

Software tools

Optimized with

STM32Cube AI Studio

STM32Cube AI Studio

Most suitable for

STM32N6 Series

Most suitable for

Resources

Optimized with STM32Cube AI Studio

STM32Cube AI Studio is a desktop tool designed to evaluate, optimize and compile neural network models for STM32 microcontrollers. It fully supports compilation for the Neural-ART Accelerator neural processing unit (NPU). It replaces the X-CUBE-AI in the ST AI product offering to cover new STM32 devices.

Optimized with STM32Cube AI Studio Optimized with STM32Cube AI Studio Optimized with STM32Cube AI Studio

Most suitable for STM32N6 Series

The STM32 family of 32-bit microcontrollers based on the Arm Cortex®-M processor is designed to offer new degrees of freedom to MCU users. It offers products combining very high performance, real-time capabilities, digital signal processing, low-power / low-voltage operation, and connectivity, while maintaining full integration and ease of development.

Most suitable for STM32N6 Series Most suitable for STM32N6 Series Most suitable for STM32N6 Series
You might also be interested by

Vision | Tutorial | GitHub | Appliances | Smart city | STM32Cube AI Studio | Object detection | STM32 AI MCU

How edge AI cameras are changing the future of smart retail

Detect and count fridge beverages fully on-device for smarter retail inventory, with Camthink's NE301

Vision | STM32Cube.AI | STM32 AI MCU | Partner | Video | Smart home | Smart building

How to personalize smart home with familiar face identification

embedUR's on-device face authentication embedded on STM32N6 with easy mobile enrollment

Vision | STM32Cube.AI | STM32 AI MCU | Partner | Smart home | Wearables | Microphone | Accelerometer | Tutorial

Handheld development platform for real-time vision, motion, and voice at the edge

All‑in‑one STM32N6‑based platform for on‑device edge AI with NPU acceleration