Back to Nexa Sdk

Platforms & runtimes

docs/en/get-started/platforms.mdx

0.3.197.3 KB
Original Source

import Feedback from "/snippets/page-feedback.mdx";

GenieX runs exclusively on Qualcomm Snapdragon — no x86 or non-Snapdragon ARM build. To get going, pick the Snapdragon platform you'll run on, then pick a runtime that matches the model you want to run.

Snapdragon platforms

GenieX is supported across three Snapdragon families — compute, mobile, and IoT — covering Windows ARM64, Android, and Linux ARM64.

Supported chipsets

These are the chipsets GenieX is validated on. Each row lists the SoC identifier you'd see from auto-detection, and the AI Hub chipset id used when pulling Qualcomm AI Hub Models.

FamilyChipsetSoC idAI Hub chipset id
Compute (Windows ARM64 / Copilot+ PC)Snapdragon® X EliteX1E*qualcomm-snapdragon-x-elite
Snapdragon® X2 EliteX2E*qualcomm-snapdragon-x2-elite
Mobile (Android)Snapdragon® 8 EliteSM8750qualcomm-snapdragon-8-elite
Snapdragon® 8 Elite Gen 5SM8850resolved from SM8850 (see note)
IoT (Linux ARM64 — cameras, robotics, industrial)Qualcomm® Dragonwing™ IQ-9075QCS9075qualcomm-qcs9075
Qualcomm® Dragonwing™ IQ-8275QCS8275qualcomm-qcs8275
<Note> **Chipset vs. SoC id.** Snapdragon X-series parts are identified by their Oryon CPU SKU (`X1E80100`, `X2E80100`, …); every part within a generation shares one NPU architecture, so **all X Elite SKUs map to the same AI Hub asset** — including the X Plus and X2 Plus parts. Android reports its SoC through `ro.soc.model`, and Dragonwing boards through the device tree.

On Android, GenieX passes the SoC id (e.g. SM8850) straight to Qualcomm AI Hub, which resolves it through its own alias table — GenieX deliberately keeps no second mapping. Pass the SoC id, not an AI Hub chipset name, and the right asset is selected. Variant suffixes are not exposed by ro.soc.model: a Galaxy S25 reports SM8750, not SM8750-AC (the alias for qualcomm-snapdragon-8-elite-for-galaxy), so pass the chipset explicitly if you need the variant asset. </Note>

GenieX auto-detects the chipset on Windows on Snapdragon, Dragonwing Linux, and Android. Check what it found, or set it explicitly:

bash
geniex config get chipset          # show the detected (or configured) chipset
geniex config set chipset          # launch an interactive picker

<Warning>Android requires an explicit chipset for Qualcomm AI Hub pulls — auto-detect via ro.soc.model covers the CLI/SDK path, but the Android SDK needs ModelPullInput.chipset set to "SM8750" or "SM8850". See Android API reference.</Warning>

Interfaces per OS

OSInterfaces
Windows ARM64 (Compute / Copilot+ PC)<ul><li>CLI</li><li>Python SDK</li><li>Local server</li></ul>
Android (Mobile)<ul><li>Android SDK (Kotlin, Maven Central)</li></ul>
Linux ARM64 (Dragonwing IoT)<ul><li>Native install</li><li>Docker</li></ul>

The chipsets above are the validated set. GenieX may run on other Snapdragon parts within the same families — for everything Qualcomm AI Hub can compile for, see the Qualcomm AI Hub device list.

<Tip>No device on hand? Sign in to Qualcomm Developer Cloud (QDC) for remote sessions on Snapdragon X Elite / X2 Elite, Snapdragon 8 Elite / 8 Elite Gen 5, and Dragonwing IQ-9075. See the QDC walkthrough in the FAQ.</Tip>

GenieX runtimes

GenieX ships with two runtimes so you get both broad model coverage and peak Snapdragon performance in one stack:

  • llama_cpp — any GGUF model on Hugging Face, running on Hexagon NPU, Adreno GPU, or CPU through Qualcomm's GGML Hexagon backend. The widest model selection.
  • qairt (Qualcomm® AI Engine Direct) — pre-compiled bundles from Qualcomm AI Hub, compiled and quantized per chipset and pinned to the Hexagon NPU. The fastest path when your model is on Qualcomm AI Hub.

<Note>Qualcomm AI Engine Direct (also known as the Qualcomm AI Engine Direct SDK, Qualcomm AI Runtime, and historically QAIRT) is the official name. Throughout these docs we use the official name.</Note>

llama.cppQualcomm AI Engine Direct
Model formatGGUF (any community model)Qualcomm AI Hub pre-compiled bundles
Compute unitsNPU / GPU / CPUNPU only
Precisions (Quantizations) picked byYou (Q4_0, Q8_0, F16, …)Pre-quantized in the bundle
Best forBringing your own GGUF from Hugging FaceHighest NPU performance on Qualcomm® AI Hub Models

Pick llama_cpp for any GGUF from Hugging Face, or when you need CPU/GPU fallback (e.g. IoT devices without HTP). Pick qairt for the fastest NPU path on models published to Qualcomm AI Hub.

Defaults

If you don't pass a compute unit:

RuntimeDefault compute unit
llama_cppnpu (pinned to HTP0)
qairtnpu

For llama.cpp's HTP + CPU per-tensor scheduling (the faster path on Snapdragon), pass hybrid explicitly.


llama.cpp

The llama_cpp runtime executes any GGUF model through llama.cpp with Qualcomm's GGML Hexagon backend. Pull any community GGUF from Hugging Face and run it on Snapdragon NPU, Adreno GPU, or pure CPU.

Compute units

--compute maps to the underlying hardware as follows:

AliasEffect
npu (default)Pin to Hexagon NPU (HTP0). Best NPU-only path.
gpuAdreno GPU via OpenCL.
cpuPure CPU. Forces nGpuLayers = 0.
hybridEmpty device_id + n_gpu_layers=-1 (all layers) — llama.cpp's per-tensor HTP+CPU scheduler. The fast path on Snapdragon.

The precision you pick at geniex pull time also determines where the model lands — see Precisions (Quantizations) Supported.


Qualcomm AI Engine Direct

The qairt runtime executes pre-compiled bundles from Qualcomm AI Hub through Qualcomm® AI Engine Direct. NPU-only, with the bundle compiled and quantized for a specific Snapdragon chipset — typically the fastest NPU path when your model is on Qualcomm AI Hub.

Compute units

Qualcomm AI Engine Direct is NPU only.

AliasEffect
npu (default)Pin HTP0 — the only supported path.
cpu / gpuCoerced to npu with a warning. Never an error.

Runtime constraints

The bundle has its precision, context length, and KV cache size baked in — none can be changed at runtime. On Android, nGpuLayers != 0 and nCtx != 0 are rejected with PARAM_NOT_SUPPORTED; leave both at defaults and tune max_tokens / enable_thinking only. To change precision or context length, get a different bundle from Qualcomm AI Hub.

<Feedback/>