Rate this Page
★ ★ ★ ★ ★

MLX Backend#

The MLX delegate is the ExecuTorch backend for Apple Silicon GPUs via the MLX framework. It compiles PyTorch models into a custom FlatBuffer bytecode format at export time and executes them using MLX GPU primitives at runtime.

Note

The MLX delegate is experimental and under active development.

Features#

  • GPU acceleration on Apple Silicon (M1 and later) via MLX.

  • INT2/INT4/INT8 weight quantization via TorchAO.

  • Dynamic shape support.

  • Mutable buffers for persistent state across inference calls (e.g., KV cache).

  • Zero-copy constant loading on unified memory.

Target Requirements#

One of:

  • macOS >= 14.0 on an Apple Silicon Mac (M1 or later)

  • iOS or iPadOS >= 17.0 on a real device (experimental). The backend is built and shipped for iOS, but running a model on a physical device is not yet covered by CI. The iOS simulator has no Metal device MLX can use, so the backend reports itself unavailable there and a model delegated to it will not load.

Development Requirements#

  • macOS on Apple Silicon (M1 or later)

  • Xcode (full installation, not just Command Line Tools)

  • The Metal Toolchain component, installed separately with xcodebuild -downloadComponent MetalToolchain

Verify the Metal compiler is available:

xcrun -sdk macosx metal --version

If this prints version information, you’re set. If it errors, install Xcode from developer.apple.com, select it as the active developer directory, and download the Metal Toolchain:

sudo xcode-select -s /Applications/Xcode.app/Contents/Developer
xcodebuild -downloadComponent MetalToolchain

Using the MLX Backend#

To target the MLX backend during export and lowering, pass an instance of MLXPartitioner to to_edge_transform_and_lower. The MLX backend also provides a set of graph optimization passes via get_default_passes() that should be passed as transform_passes. The example below demonstrates this process using MobileNet V2:

import torch
import torchvision.models as models
from torchvision.models.mobilenetv2 import MobileNet_V2_Weights
from executorch.backends.mlx import MLXPartitioner
from executorch.backends.mlx.passes import get_default_passes
from executorch.exir import EdgeCompileConfig, to_edge_transform_and_lower

mobilenet_v2 = models.mobilenetv2.mobilenet_v2(weights=MobileNet_V2_Weights.DEFAULT).eval()
sample_inputs = (torch.randn(1, 3, 224, 224), )

et_program = to_edge_transform_and_lower(
    torch.export.export(mobilenet_v2, sample_inputs),
    transform_passes=get_default_passes(),
    partitioner=[MLXPartitioner()],
    compile_config=EdgeCompileConfig(
        _check_ir_validity=False,
        _skip_dim_order=True,
    ),
).to_executorch()

with open("mv2_mlx.pte", "wb") as file:
    et_program.write_to_file(file)

get_default_passes() includes RMSNorm fusion, consecutive view/permute/dtype-cast collapsing, no-op removal, and common subexpression elimination. These are recommended for all models and required for optimal LLM performance. The accompanying EdgeCompileConfig uses the settings in the MLX export examples. The _check_ir_validity and _skip_dim_order fields are internal and may change between ExecuTorch versions.

Note

The MLX backend is primarily designed for LLM and generative AI workloads on Apple Silicon. The MobileNet V2 example above is shown for simplicity, but in practice you would use this backend for models like Llama, Whisper, and other transformer-based architectures. See LLM example for a more representative use case.

See Partitioner API for a reference on available partitioner options.


Quantization#

The MLX backend supports INT4, INT8, and NVFP4 weight quantization via TorchAO for both linear and embedding layers. This is particularly useful for LLM inference. See MLX Quantization for details.


Runtime Integration#

Python (pybindings)#

The simplest way to get started is to install ExecuTorch with Python bindings. From the repo root:

python install_executorch.py

On Apple Silicon, when the Metal compiler is available, the MLX backend is automatically included. You can then export models in Python using the MLX partitioner and run them via the ExecuTorch Python API.

C++ (CMake preset)#

To build the C++ runtime with the MLX delegate, use the mlx-release CMake workflow preset from the repo root:

cmake --workflow --preset mlx-release

This configures and builds a Release build of the ExecuTorch runtime with the MLX delegate and installs artifacts into cmake-out/. The preset enables the MLX delegate along with commonly needed extensions (module, data loader, flat tensor, LLM runner, etc.).

Downstream C++ apps can then find_package(executorch) and link against mlxdelegate and mlx. The executorch_target_link_options_shared_lib utility handles whole-archive linkage (required for static initializer registration) cross-platform, and executorch_target_copy_mlx_metallib copies the Metal kernel library next to the binary so MLX can find it at runtime:

# CMakeLists.txt
find_package(executorch REQUIRED)

# Link MLX delegate (with whole-archive for static initializer registration)
target_link_libraries(my_target PRIVATE mlxdelegate mlx)
executorch_target_link_options_shared_lib(mlxdelegate)

# Copy mlx.metallib next to the binary for runtime
executorch_target_copy_mlx_metallib(my_target)

No additional steps are necessary to use the backend beyond linking the target. An MLX-delegated .pte file will automatically run on the registered backend.

There is also an mlx-debug preset useful during development:

cmake --workflow --preset mlx-debug

Runtime Options#

The MLX backend reads optional per-model runtime specs, set through a LoadBackendOptionsMap keyed by the backend id MLXBackend. All are optional and off by default.

eval_threshold_bytes (int)#

MLX is lazy: dispatching an instruction only builds a graph node, and nothing is materialized until the method’s outputs are evaluated. For a long instruction chain that means every intermediate in the method is live at the same instant, so peak memory tracks the size of the whole graph rather than the working set. Whisper-small’s 495-instruction encode peaks at 1105 MB of MLX allocation against 95 MB of steady-state active memory.

Set this key to N to evaluate the live per-execution tensors once the intermediates produced since the last evaluation exceed N bytes. Each evaluation costs a GPU sync, so the cost tracks the number of evaluations, and budgeting bytes rather than instructions puts them only in the methods that actually allocate.

0 (the default) disables the mechanism entirely, preserving the previous behaviour with no accounting overhead of any kind.

#include <executorch/backends/mlx/runtime/backend_options.h>
#include <executorch/runtime/backend/options.h>

executorch::runtime::BackendOptions<1> opts;
opts.set_option(executorch::backends::mlx::kEvalThresholdBytesKey,
                512 * 1024 * 1024);

Measured on an iPhone 16, whisper-small int8, full pipeline, medians of interleaved rounds:

setting

peak MB

peak while loaded

pipeline ms

0 (disabled)

1194.4

763.7

885.2

512 MB

692.8

261.0

831.3

This is a threshold, not a hard memory limit. It is best-effort evaluation scheduling and peak footprint can exceed the value:

  • A long SCAN or IF branch accumulates across its whole body and is only checked once control returns to the enclosing chain, so it can overshoot by the size of that body.

  • The per-instruction estimate is the largest tensor the instruction touches, which can overcount (an op that only reads a large tensor is charged for it) and so can evaluate earlier than the true pending bytes warrant.

  • Ops that evaluate internally reduce the real pending work without reducing the running estimate.

Tune it against measurements rather than expecting the value to bound RSS.

Reference#

→Troubleshooting — Debug common issues.

→Partitioner API — Partitioner options.

→Quantization — Supported quantization schemes.

→Op Support — Supported operators.