On-device inference
Models quantised and compiled for embedded SoCs and NPUs, profiled on the actual board rather than a desktop GPU.
What you get
- Measured latency and power on target hardware
- INT8 / FP16 paths with accuracy deltas documented
- Runtime packaged for the board you ship