VinRobotics · 2025

Edge AI on NXP i.MX93

I took an embedded Linux board from bring-up to running a learned control policy, so the robotics team could try its models on hardware small enough to sit inside a robot.

Goal

The robotics team wanted to know whether a small Linux SoC could run its learned control policies next to the real-time control stack it already used. Answering that needed a platform first: a reproducible image with a real-time kernel, a camera and an inference runtime, and a toolchain to build models and applications against it.

What I built

Board bring-up

NXP's real-time distribution supported only the evaluation kit, so I backported the FRDM board's device tree into it to get the board booting. I then moved to NXP's newer BSP, which supports the board directly, and built it with the team's PREEMPT_RT kernel.

A camera in the kernel

The board routes its MIPI camera port to an image processor, and NXP's kernel had no driver for the Sony IMX708 the team wanted. I ported the driver from the Raspberry Pi kernel and wrote the device-tree changes that connect the sensor to the MIPI CSI-2 receiver.

Runtime and toolchain

Yocto recipes build NCNN with its INT8 and BF16 kernels and package it so the same CMake project builds natively inside Yocto or with the cross-compilation SDK I generated from the image.

From PyTorch to the board

I converted the team's LSTM control policy with pnnx and wrote a small C++ runner, compiled for the Cortex-A55, that times one inference step.

Layer diagram from silicon to model. Added in this work: the IMX708 driver port and device-tree change in the kernel, the recipes and SDK in the image, the NCNN shared library, the demo application and the converted policy. The platform underneath: NXP's i.MX93 and BSP, the Sony IMX708 sensor, and the team's real-time layer and patch.
What runs on the board, layer by layer. Green marks what this work added on top of NXP's BSP and the team's real-time layer.

Result

Six launches of the runner on the board:

# ./imx93_ncnn_demo
inference time: 1124 us
# ./imx93_ncnn_demo
inference time: 308 us
# ./imx93_ncnn_demo
inference time: 309 us
# ./imx93_ncnn_demo
inference time: 345 us
# ./imx93_ncnn_demo
inference time: 305 us
# ./imx93_ncnn_demo
inference time: 350 us

Once warm, a step took about a third of a millisecond on a single Cortex-A55 core; the first launch took about one millisecond while caches filled. The timer covers inference only and the inputs are zeros, so every run prints the same actions. With the build cache warm, a kernel change reached a bootable image in under ten minutes.

What I learned

  • Bringing up a board on embedded Linux: device trees, porting a media driver between kernel trees, and keeping kernel patches with devtool.
  • Building a Yocto layer others can use: recipes, bbappends, packages and an SDK.
  • Taking a PyTorch model to an embedded runtime and timing it on the target instead of a workstation.
  • Fitting inference into a platform built around real-time control, with PREEMPT_RT and EtherCAT on the same image.