A framework for running generative AI models on Snapdragon NPUs.
Qualcomm AI Runtime (QAIRT) is a suite of tools for developing, running, and optimizing AI models on Qualcomm hardware. It has the best hardware-aware design to get the metal performance.
This plugin, as one of the backends of geniex, uses QAIRT to support various generative AI models.
Executables and geniex_core (shared library) are placed under the build tree; see each platform below. The HTP runtime libs are copied to <build>/bin/htp-files/ automatically.
| Option | Default | Description |
|---|---|---|
GENIEX_BUILD_VLM |
OFF |
Build Vision-Language models (e.g. Qwen2.5-VL). |
GENIEX_BUILD_EXAMPLES |
OFF |
Build per-model example executables. |
GENIEX_BUILD_TESTS |
OFF |
Register CTest entries for LLM/VLM pipeline tests. Requires a Snapdragon NPU host. See tests/README.md. |
GENIEX_DEBUG |
OFF |
Verbose logging with file/line/func info; also compiles in the tensor-IO dump path (opt-in at runtime, see below). |
Built with -DGENIEX_DEBUG=ON, Graph::execute() can dump every graph's input and
output tensors to .npy files (dtype + shape self-described in the file header,
loadable with numpy.load()), one directory per graph. Disabled by default even
in a GENIEX_DEBUG build -- opt in at runtime via:
| Env var | Meaning |
|---|---|
GENIEX_DUMP_TENSOR_IO=<dir> |
Enables dumping and sets the output root directory. Unset = disabled. |
GENIEX_DUMP_TENSOR_IO_MAX_CALLS=<N> |
Caps how many execute() calls per graph get dumped (default 10; 0 = unlimited). |
Layout: <dir>/<graph_name>/<call_idx:03d>_{in,out}_<tensor_name>.npy. Float and
quantized tensors are dequantized to float32; integer tensors (ids, masks) keep
their native dtype.
HTP optrace profiling is available through ModelConfig:
geniex::ModelConfig model_cfg = geniex::modelConfigFromDirectory(model_dir);
model_cfg.enable_optrace = true;
model_cfg.optrace_output_path = "profile/qnn-profiling-data.log";The generic auto_llm example exposes the same setting as
--optrace [output-path].
The output log can be opened directly with qnn-profile-viewer when the context
binaries contain linked schematic data, as current GenieX model assets do.
Optrace capture requires a QAIRT 2.48 or newer runtime because it uses QNN
System profile serialization. Set QnnRuntimeConfig::htp_dir or
GENIEX_QAIRT_LIB to that runtime; the bundled QAIRT 2.45 runtime continues to
work when optrace is off. The QAIRT 2.45 viewer can read logs captured with the
newer runtime.
Enabling optrace uses detailed QNN profiling and serializes context-load and graph-execution events. The output file is replaced when the model initializes, then subsequent execution records are appended to it.
Prerequisites: Visual Studio 2022 with the MSVC ARM64 workload, CMake ≥ 3.17, Rust (with aarch64-pc-windows-msvc target — needed for the tokenizer).
# Configure
cmake -B build -A ARM64
# Build everything
cmake --build build --config Release -j32
# Build a specific model target
cmake --build build --config Release --target qwen3_4b -j32
# VLM build (Qwen2.5-VL)
cmake -B build -A ARM64 -DGENIEX_BUILD_VLM=ON
cmake --build build --config Release -j32Output: build/bin/Release/*.exe and geniex_core.dll.
Prerequisites: Android NDK (r25+ recommended), CMake, Rust with aarch64-linux-android target (rustup target add aarch64-linux-android).
export ANDROID_NDK_ROOT=/path/to/android-ndk
./build_android.sh # arm64-v8a Release, all examples
./build_android.sh --target qwen3_4b # build a single target
./build_android.sh --vlm --target qwen2_5_vl_7b # enable GENIEX_BUILD_VLM
./build_android.sh --debug --debug-log # Debug + verbose logging
./build_android.sh --help # full flag listOutput: build-android/bin/* (no extension) and libgeniex_core.so.
The script auto-detects the NDK host tag (
linux-x86_64vsdarwin-x86_64). It does not support building on Windows hosts.
Prerequisites: gcc ≥ 11.2 (matching the bundled runtime in third-party/linux-gcc11.2/), CMake, Rust.
# Configure
cmake -B build -DCMAKE_BUILD_TYPE=Release
# Build
cmake --build build -j$(nproc)
# Build a specific model target
cmake --build build --target qwen3_4b -j$(nproc)Output: build/bin/* and libgeniex_core.so.
| Hardware | SoC | HTP Arch | SoC Model |
|---|---|---|---|
| Snapdragon X Elite / Plus | SC8380 | v73 | 60 |
| IQ-9075 | QCS9075 | v73 | 60 |
| Snapdragon 8 Elite | SM8750 | v79 | 69 |
| Snapdragon 8 Elite Gen5 | SM8850 | v81 | 88 |
The bundled HTP runtime libs in
third-party/(windows,android,linux-gcc11.2) are QAIRT v2.45.0.260326 (single source of truth:GENIEX_QAIRT_VERSIONincore/include/version.h; consumers read it at runtime viageniex_qairt_version()). Runtime version is backward compatible with compile version, so all models compiled with v2.45 or earlier will run correctly.That is the version of the libs. What decides whether a runtime loads is the C API in
qnn-api/include/— see Using a different QAIRT runtime.
By default the plugin compiles against the single header set in qnn-api/include/, deliberately the lowest QNN C API we support (2.27). To compile against a different header set instead, set QAIRT_QNN_HEADERS to a directory containing QnnCommon.h, HTP/, and System/:
cmake -B build -DQAIRT_QNN_HEADERS=/path/to/qairt/includeqnn-api/include/ (this plugin's own MmappedFile.hpp/MmappedReader.hpp helpers) stays on the include path regardless, since an external SDK won't ship those. The load-time floor below stays at 2.27 whichever headers you compile against.
The build bundles the QAIRT runtime above and copies it to htp-files/ next to geniex_core, so the default path needs no configuration — no SDK download, no paths to set.
To run against a different QAIRT version instead, set GENIEX_QAIRT_LIB to a directory holding the runtime libraries:
# Windows
set GENIEX_QAIRT_LIB=C:\path\to\qairt-libs
# Linux / Android
export GENIEX_QAIRT_LIB=/path/to/qairt-libsOne build drives many runtimes: the plugin reaches QNN only through the versioned C interface, which negotiates at load time.
What sets the general runtime floor is the C API version (kMinApiMinor in
QnnApi.cpp, 2.27), not the bundled-lib release (GENIEX_QAIRT_VERSION, 2.45)
and not the headers compiled against. Optional features can require newer API
tails; per-operator profiling requires QNN System API 1.12 from QAIRT 2.48+.
| QAIRT SDK | QNN C API | Loads? | Optrace? |
|---|---|---|---|
| 2.36 (header floor) | 2.27 | ✅ floor | ❌ |
| 2.45 (bundled) | 2.34 | ✅ verified | ❌ |
| 2.48 | 2.37 | ✅ verified | ✅ |
| 2.49 | 2.38 | ✅ verified | ✅ |
| 2.50 (Workbench compiles against) | 2.39 | ✅ verified | ✅ |
| older than 2.36 | < 2.27 | ❌ rejected at load | ❌ |
Either layout works. A flat folder with the host libraries and their arch stubs together — the same shape as the bundled htp-files/:
qairt-libs/
├── QnnHtp.dll (libQnnHtp.so)
├── QnnSystem.dll (libQnnSystem.so)
├── QnnHtpNetRunExtensions.dll (libQnnHtpNetRunExtensions.so)
└── QnnHtpV73Stub.dll, ... arch stubs and skels
…or a stock QAIRT SDK root, as unpacked from the Qualcomm Software Center:
qairt/2.XX.0/
└── lib/
├── aarch64-windows-msvc/ host libraries (or aarch64-android,
│ aarch64-oe-linux-gcc11.2)
└── hexagon-v73/unsigned/ skels, one folder per arch
hexagon-v81/unsigned/
Host libraries come from lib/<target-triple>/; every lib/hexagon-v*/ folder goes on ADSP_LIBRARY_PATH so FastRPC matches the device's arch. An unrecognised triple is found by scanning lib/, so a renamed one (the Linux gcc suffix moves between releases) still resolves. The INFO log names the folder the host libraries actually came from, which for an SDK root isn't the path you passed.
Highest precedence first:
| Rung | Source |
|---|---|
| 1 | QnnRuntimeConfig::backend_path / system_lib_path / extensions_path, when all three are set |
| 2 | QnnRuntimeConfig::htp_dir |
| 3 | GENIEX_QAIRT_LIB |
| 4 | bundled htp-files/ next to geniex_core |
The chosen directory and rung are logged at INFO. Check that line before trusting a run against a non-bundled runtime: a mismatched runtime can load and generate at full speed while producing wrong output, so confirming which libraries loaded is the only reliable check.
├── models/ # Model specs (.h) and example executables (.cpp)
│ ├── falcon3/
│ ├── llama3/
│ ├── llama3_1/
│ ├── llama3_2/
│ ├── llama3_2_ssd/
│ ├── phi3_5/
│ ├── qwen2_5/
│ ├── qwen2_5_vl/ # VLM (requires GENIEX_BUILD_VLM=ON)
│ └── qwen3/
├── core/ # geniex_core framework (LLM model, graph, KV cache, RoPE)
├── modelfiles/ # Tokenizer and config files per model
├── qnn-api/ # QNN SDK integration layer (headers + API wrappers)
├── third-party/ # HTP runtime libs + geniex-proc submodule (tokenizer, preprocessing)
└── docs/ # Documentation
For security-sensitive reports, see SECURITY.md.
Contributions are welcome! Please read CONTRIBUTING.md for the branching model, pull-request workflow, and DCO sign-off requirement, and CODE-OF-CONDUCT.md for community expectations.
GenieX-QAIRT-plugin is licensed under the BSD 3-Clause License. See LICENSE.txt for the full license text.
This project also ships vendored third-party components (the QAIRT SDK
files under qnn-api/ and the prebuilt runtime libraries under
third-party/) that are governed by separate licenses. See
THIRD_PARTY_NOTICES.md for details.