Many teams get their models working during the algorithm validation phase, but hit a wall when it comes to productization. The model performs well on a server, but once moved to an embedded device, it becomes slow, laggy, or runs out of memory, and sometimes the operators aren't even supported. In reality, going from a trained model to an embedded vision product isn't just about copying weight files over. There's a clear path to follow, and the key is using the right approach.
The first step is not to rush into changing hardware, but to re-examine your model. The backbone used during training is often oversized and leaves too much accuracy margin. What you really need is a model that is just good enough, so the first thing to do is to streamline the structure. You can try replacing the original backbone with lightweight networks such as MobileNet or EfficientNet-Lite. If you prefer not to change the network structure, you can also compress the model through pruning and quantization. Here's a tip: prioritize INT8 quantization. Most NPUs on embedded chips support INT8 best, and the accuracy loss is usually manageable. If accuracy drops significantly after quantization, try channel pruning or knowledge distillation before quantizing again.
The second step is toolchain adaptation. Different chip platforms have their own conversion tools, such as Renesas e-AI, NVIDIA TensorRT, and Rockchip RKNN. Once you have the model, run it through the conversion tool to turn it into the target platform's format. This is where unsupported operator issues often appear. There are two solutions. One is to modify the training code and replace special operators with common ones supported by the platform, such as decomposing custom attention modules into standard convolutions and fully connected layers. The other is to leverage deep learning compilers like TVM or Glow, which can automatically perform graph optimization and code generation, reducing the workload of manual adaptation.
The third step is to emphasize the closed loop of deployment validation. Many developers think they're done after running simulations on a PC, but in reality, execution efficiency, memory bandwidth, and heat generation on real hardware all affect performance. It's recommended to flash the model onto the target board as early as possible and test it with images captured by a real camera, paying attention to frame rate, latency fluctuation, and CPU usage. If the speed isn't sufficient, don't just focus on optimizing the model. Also check the data pipeline: is there parallelism between camera capture, preprocessing, inference, and post-processing? Are multithreading or hardware codec units being used? Many times the bottleneck isn't the model itself, but seemingly simple steps like image resizing or color conversion.
Finally, here's a real-world case. In a quality inspection project, ResNet50 was originally used to process industrial camera images, achieving less than 5 FPS. After switching to EfficientNet-Lite and applying INT8 quantization, it ran at 30 FPS on the Renesas RZ/V2M, with only a 0.3% accuracy drop. The entire change took just two weeks, with most of the time spent on operator replacement and memory reuse. This example shows that quickly deploying an existing model is entirely feasible, but it requires a three-step approach: model lightweighting, toolchain adaptation, and software-hardware co-optimization. Don't aim for perfection in one go. First, get a minimal end-to-end system working, then gradually add features. Iterating on embedded products is inherently a process of continuous optimization.
