What TensorFlow Lite Micro actually costs on real microcontrollers, per model and per layer.
Ground truth for mcufit. Every number here came off a physical chip.
Each chip is slow at a different operator, and nothing tells you which.
Throughput in MACs per clock cycle, measured layer by layer:
| operator | ESP32 (esp-nn) | Nano 33 BLE (CMSIS-NN) |
|---|---|---|
CONV_2D |
0.073 | 0.190 |
DEPTHWISE_CONV_2D |
0.044 | 0.063 |
FULLY_CONNECTED |
0.022 | 0.186 |
Relative to its own convolution, the ESP32 is 3.2x worse at fully-connected and 1.6x worse at depthwise. The nRF52840 is the other way round: 1.14x at fully-connected and 3.3x worse at depthwise.
esp-nn ships optimised convolution and depthwise for the Xtensa LX6 but no fully-connected kernel, so those layers fall back to reference C. CMSIS-NN covers fully-connected and is comparatively weak on depthwise. Neither toolchain warns you, and neither publishes this.
So a chip cannot be ranked without knowing the model, and a model cannot be costed without measuring the chip. The Nano beats the ESP32 by 2.7x on depthwise-heavy person detection and by 8.4x on the all-fully-connected anomaly detector.
Small layers also cost a flat fee. On the ESP32, SOFTMAX over 2 values
takes 385 µs and AVERAGE_POOL_2D 415 µs per call, independent of size.
Invisible to any estimate based on operation counts.
Median of 30 timed inferences after 3 warmups, Invoke() only.
| model | operators | MACs | ESP32 @ 240 MHz | Nano 33 BLE @ 64 MHz |
|---|---|---|---|---|
| person_detect | depthwise-heavy | 7.16 M | 474 ms | 657 ms |
| kws_ref_model | conv + depthwise | 2.66 M | 162 ms | 225 ms |
| pretrainedResnet_quant | pure conv + add | 12.5 M | 724 ms | 1242 ms |
| ad01_int8 | 10x fully-connected | 0.26 M | 49 ms | 22 ms |
Raw records in results/results.jsonl (whole model)
and results/layers.jsonl (per operator). Layer times
sum to the whole-model median within 0.2%.
Run-to-run spread is 41 µs out of 474 ms. Two ESP32 boards of different silicon revisions (D0WDQ6 rev v1.0 and D0WD-V3 rev v3.1) returned medians 0-1 µs apart on all four models.
One reading per configuration is enough. Repeats and per-device records buy nothing.
The limit: that holds within one binary. Adding profiler code elsewhere in
the firmware moved the ResNet result by 2.3% while the other three moved under
0.2%, so store the build alongside the measurement. Every CI artifact ships a
provenance.txt for that reason.
mcufit's own latency estimate was 3.2x optimistic and ranked the two chips backwards. It has since been rewritten to use these measurements, and to return nothing at all for boards that have none.
Espressif's published figures did not reproduce. Their README quotes
person_detect at 380 ms with esp-nn and 4084 ms without. We measured 474 ms and
600 ms, a 1.27x gap rather than 11x. CONFIG_NN_ANSI_C still uses esp-nn's own
C kernels rather than stock TFLM reference kernels, so their 4084 ms is likely
a third configuration.
Measuring the arena on a laptop over-reports it. A 64-bit host build needs more interpreter bookkeeping than a 32-bit chip:
| arena section | host, 64-bit | wasm32 | real device |
|---|---|---|---|
| activations | 55,296 | 55,296 | 55,296 |
| interpreter overhead | 33,952 | 29,132 | 27,004 |
| total | 89,248 (+8.4%) | 84,428 (+2.6%) | 82,300 |
Activations come from the model and are identical everywhere. Everything else follows pointer width. mcufit now uses the wasm32 build because of this.
Nothing to install. GitHub builds the firmware, esptool flashes it.
ESP32
- Push, or hit Run workflow on the
buildaction. - Download the
mcufit-bench-fastartifact. It holds one merged.binplus theprovenance.txtrecording exactly how it was built. - Flash to offset
0x0and read the output at 115200 after a reset.
pip install esptool
esptool --port /dev/cu.usbserial-0001 --chip esp32 --baud 460800 \
write-flash 0x0 mcufit-bench-fast.binOr flash from Chrome at https://espressif.github.io/esptool-js/, same offset, no Python needed.
Nano 33 BLE
arduino-cli core install arduino:mbed_nano
arduino-cli lib install "Chirale_TensorFLowLite"
arduino-cli compile --fqbn arduino:mbed_nano:nano33ble arduino/mcufit_bench_nano
arduino-cli upload -p /dev/cu.usbmodem114101 \
--fqbn arduino:mbed_nano:nano33ble arduino/mcufit_bench_nanoAppend the MCUFIT_RESULT and MCUFIT_LAYER lines it prints to the files in
results/.
Each model runs twice. Once plain, timing only Invoke() with
esp_timer_get_time() (or micros() on Arduino), and once with a
TagProfiler attached.
The profiler implements tflite::MicroProfilerInterface. TFLM wraps every
operator's invoke in a ScopedMicroProfiler tagged with the operator name, so
accumulating per tag gives the cost of each layer type in about 40 lines. The
hook is compiled out by -DTF_LITE_STRIP_ERROR_STRINGS, which esp-tflite-micro
never defines.
scripts/macs_by_op.py counts MACs per operator on the host using mcufit's own
parser, and scripts/layer_throughput.py joins the two into the MACs/cycle
table above.
Two ESP32 build variants differ only in the kernel library, with 240 MHz and
-O2 pinned in sdkconfig.defaults so nothing else can vary:
sdkconfig.fastsetsCONFIG_NN_OPTIMIZED=ysdkconfig.slowsetsCONFIG_NN_ANSI_C=y
Three boards is not enough, and the gaps are the point of the finding above: a
chip cannot be ranked without measuring it. If you own a microcontroller that
is not in results/, it takes about fifteen minutes and needs only
arduino-cli and a USB cable.
arduino-cli lib install "Chirale_TensorFlowLite" # or tflm_esp32 for ESP32s
python3 scripts/prepare_arduino.py --target rp2040 --out /tmp/mcufit_bench
arduino-cli compile --fqbn <FQBN> /tmp/mcufit_bench
arduino-cli upload -p <PORT> --fqbn <FQBN> /tmp/mcufit_bench
arduino-cli monitor -p <PORT> --config baudrate=115200 | tee /tmp/capture.txt
python3 scripts/ingest.py /tmp/capture.txt --writeingest.py refuses anything that would poison the database: an unrecognised
chip, a clock of zero, a model whose bytes do not match models/. Then open a
PR with the changed files in results/.
Whole families are still blank, and one board answers for its whole family: Cortex-M7, Xtensa LX7, the RISC-V ESP32s, Cortex-M0+, Cortex-M33 and Cortex-M3. The one we want most is the ESP32-S3, where esp-nn ships hand-written assembly rather than generic C, so it is the board most likely to break the current model rather than confirm it.
Full instructions, and why simulator numbers are not accepted, in CONTRIBUTING.md.
Notes that cost time
esp-idf-ci-actionbuilds inside Docker as root, sobuild/is not writable by the runner afterwards. Write artifacts to the repo root, and readsdkconfigfrom the project root rather thanbuild/sdkconfig.- Native-USB boards print into the void. The Nano's first run was lost
because
setup()waited only 5 s for Serial. Usewhile (!Serial) {}and read back without toggling DTR/RTS. - A run ID fetched immediately after
git pushis the previous run. Matchgh run view --json headShaagainst the commit before downloading, or you will flash a stale binary. - Models are committed as 16-byte-aligned C arrays. TFLM reads the model in
place from flash and CMake's
EMBED_FILESdoes not guarantee alignment. - The ANSI-C build takes seconds per inference, so the task watchdog needs
raising and the loop needs a
vTaskDelayyield. - Wokwi cannot be used for this. It caps the simulated CPU frequency, which corrupts the exact quantity being measured.
Regenerating the model arrays
for m in person_detect kws ic_resnet ad; do
python3 scripts/gen_model_array.py models/$m.tflite main/model_$m.cc $m
cp main/model_$m.cc arduino/mcufit_bench_nano/model_$m.cpp
doneRedistributed so the benchmark is reproducible without fetching anything. All Apache 2.0.
| file | source |
|---|---|
models/person_detect.tflite |
tflite-micro, visual wake words reference model |
models/kws.tflite |
MLPerf Tiny, keyword spotting |
models/ic_resnet.tflite |
MLPerf Tiny, image classification, ResNet-8 on CIFAR-10 |
models/ad.tflite |
MLPerf Tiny, anomaly detection |
MIT, see LICENSE. The models keep their own Apache 2.0 terms.