Rate this Page
★ ★ ★ ★ ★

Quick Start Pathway#

This pathway is for engineers who want to get a model running on a device as quickly as possible. It assumes you are familiar with PyTorch model development and have some prior exposure to mobile or edge deployment concepts. Steps are kept concise and link directly to the most actionable documentation.

Estimated time to first inference: 15–30 minutes.


Choose Your Scenario#

Select the scenario that most closely matches what you are trying to accomplish right now.

🚀 I have a PyTorch model and want to run it on device

Fastest path: Export → Run

  1. Install: pip install executorch

  2. Export with Getting Started with ExecuTorch (Exporting section)

  3. Run with Python runtime or deploy to Android / iOS

Time: ~15 min

📦 I want to use a pre-exported model

Fastest path: Download → Run

Selected target-specific .pte files are available from the ExecuTorch Community on Hugging Face. Match the artifact’s model configuration, precision, and backend to the runtime your application links.

Skip export entirely and go directly to the runtime section of Getting Started with ExecuTorch.

Time: ~10 min

🤗 I have a Hugging Face model

Fastest path: Choose an exporter

Use the experimental Transformers exporter for programmatic XNNPACK or CUDA graph export. Use Exporting LLMs with Hugging Face’s Optimum ExecuTorch for tested task workflows, quantization, and higher-level model wrappers.

Time: ~20 min

🦙 I want to run Llama on my phone

Fastest path: Llama on ExecuTorch

Follow the Llama on ExecuTorch guide for the complete Llama export and deployment workflow, including quantization and platform-specific setup.

Time: ~45 min (model download included)


The 5-Minute Setup#

If you have not yet installed ExecuTorch, run the following in a Python 3.10–3.14 virtual environment:

pip install executorch

Then verify export, XNNPACK lowering, serialization, and host execution:

import torch
from executorch.backends.xnnpack.partition.xnnpack_partitioner import XnnpackPartitioner
from executorch.exir import to_edge_transform_and_lower
from executorch.runtime import Runtime

class Add(torch.nn.Module):
    def forward(self, x, y):
        return x + y

model = Add().eval()
sample_inputs = (torch.ones(1), torch.ones(1))

et_program = to_edge_transform_and_lower(
    torch.export.export(model, sample_inputs),
    partitioner=[XnnpackPartitioner()]
).to_executorch()

with open("add.pte", "wb") as f:
    f.write(et_program.buffer)

runtime = Runtime.get()
runtime_program = runtime.load_program("add.pte")
method = runtime_program.load_method("forward")
output = method.execute(sample_inputs)[0]

torch.testing.assert_close(output, model(*sample_inputs))
print("Output:", output)

Expected output: Output: tensor([2.]).

This check runs through the host wheel’s XNNPACK backend. It verifies the exported program’s result, but it does not measure performance or prove that a different backend is available on the target device. Repeat accuracy and performance validation in the target application.


Quick Reference: Export Cheat Sheet#

Task

Code / Command

Install ExecuTorch

pip install executorch

Export with XNNPACK (mobile CPU)

to_edge_transform_and_lower(torch.export.export(model, inputs), partitioner=[XnnpackPartitioner()])

Export with Core ML (iOS)

Replace XnnpackPartitioner with CoreMLPartitioner; see Core ML Backend

Export with Qualcomm (Android NPU)

See Qualcomm AI Engine Backend for QNN SDK setup and partitioner usage

Run from Python

Load a Program, retain it while its Method is in use, then call method.execute(inputs) with one sequence containing all method inputs

Run from C++

See Running an ExecuTorch Model Using the Module Extension in C++ for the high-level Module API

Export an LLM

python -m executorch.extension.llm.export.export_llm --config path/to/config.yaml; see Exporting LLMs


Platform Quick Start Guides#

Jump directly to the platform-specific setup guide for your target.

Android Quick Start

Gradle dependency, experimental Java/Kotlin Module API, and XNNPACK / Vulkan / Qualcomm backend selection for Android.

Android
iOS Quick Start

Swift Package Manager setup, Swift/Objective-C Module APIs, C++, and Core ML / XNNPACK backend selection for iOS.

iOS
Desktop / Linux / macOS

Python runtime, C++ CMake integration, and XNNPACK / Core ML backends for desktop platforms.

Desktop & Laptop Platforms
Embedded Systems

Bare-metal and RTOS deployment, Arm Ethos-U, Cadence, NXP, and other embedded backends.

Embedded Systems

Backend Selection Guide#

Choosing the right backend has the largest impact on performance. Use this table to select the appropriate backend for your hardware.

Backend Selection by Platform and Hardware#

Platform

Hardware Target

Backend

Documentation

Android

CPU (Arm/x86)

XNNPACK

XNNPACK Backend

Android

GPU (Vulkan)

Vulkan

Vulkan Backend

Android

Qualcomm NPU/DSP

QNN

Qualcomm AI Engine Backend

Android

MediaTek APU

MediaTek

MediaTek Backend

iOS / macOS

Neural Engine / GPU

Core ML

Core ML Backend

iOS / macOS

CPU (Arm)

XNNPACK

XNNPACK Backend

Desktop

Intel CPU/GPU/NPU

OpenVINO

OpenVINO Backend

Desktop

Apple Silicon

Core ML

Core ML Backend

Embedded

Arm Cortex-M / Ethos-U

Arm Ethos-U

Arm Ethos-U Backend

Embedded

Cadence DSP

Cadence

Cadence Xtensa Backend

Embedded

NXP eIQ Neutron

NXP

NXP eIQ Neutron Backend


Troubleshooting Quick Fixes#

Symptom

Quick Fix

ImportError: No module named executorch

Run pip install executorch in your active virtual environment

Export fails with torch._dynamo error

Ensure your model is export-compatible; see Exporting to ExecuTorch

.pte file runs but produces wrong output

Use Developer Tools Usage Tutorials to compare intermediate activations

Android Gradle sync fails

Check executorchVersion in build.gradle.kts matches the release you intend to use

iOS build fails with missing xcframework

Verify the Swift PM branch name matches your ExecuTorch version (format: swiftpm-X.Y.Z)


Going Deeper#

Once your model is running, explore these topics to optimize performance and expand capabilities.