Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

PennyLane Metal

PennyLane Metal is a PennyLane quantum simulator for Apple silicon, built exclusively on Metal 4. It provides the metal.qubit device, where state evolution, noise channels, measurements, sampling, and adjoint gradients all run on the GPU.

  • Pure circuits use a state vector; noisy circuits use a density-matrix execution model, selected automatically.
  • Native Metal adjoint Jacobians and VJPs for the parameterized gate set.
  • Broad native gate, measurement, and noise-channel coverage — see supported operations.
  • Device-aware memory admission with a queryable resource planner — see memory and resources.

Requirements

  • macOS 26.0 or newer on Apple silicon
  • A Metal 4-capable Apple GPU (M1 or newer)
  • Python 3.10–3.14

Installation

pip install pennylane-metal

Binary wheels are published for macOS 26+ (arm64) and CPython 3.10–3.14.

Building from the sdist or a git checkout additionally requires Xcode 26 with the Metal Toolchain component (xcodebuild -downloadComponent MetalToolchain) and CMake 3.24+; then pip install .. See docs/DEVELOPMENT.md for the full workflow.

Quickstart

import pennylane as qml

dev = qml.device("metal.qubit", wires=2)

@qml.qnode(dev)
def bell_state():
    qml.H(0)
    qml.CNOT([0, 1])
    return qml.probs(wires=[0, 1])

print(bell_state())  # [0.5, 0.0, 0.0, 0.5]

Gradients use the native Metal adjoint method:

@qml.qnode(dev, diff_method="adjoint")
def expectation(angle):
    qml.RX(angle, 0)
    qml.CNOT([0, 1])
    return qml.expval(qml.Z(0))

angle = qml.numpy.array(0.37, requires_grad=True)
print(qml.grad(expectation)(angle))  # -sin(0.37)

Performance

A 22-qubit, depth-10 RY/CNOT circuit with an analytic expectation value (Apple M1 Max, macOS 26, pennylane-metal 0.2.0):

Device Time per execution
default.qubit 5808 ms
lightning.qubit 973 ms
metal.qubit 98 ms

Each execution carries a fixed ~1 ms GPU dispatch cost, so CPU simulators remain faster for very small circuits; the GPU advantage grows with qubit count and depth. Reproduce with benchmarks/benchmark_metal.py — see docs/DEVELOPMENT.md.

Capabilities and limits

Area Summary
Gates Native one/two/multi-qubit set incl. excitation families, arbitrary unitaries
Measurements State, probs, expval, var, purity, entropies (analytic); shots via Metal sampling
Noise Standard channels plus arbitrary multi-wire Kraus maps, all on GPU
Gradients Native adjoint/VJP (pure states); parameter-shift for noisy circuits
Precision complex64; capacity is device-memory-bound (2**n state vector, 4**n density matrix)
Not supported Mid-circuit measurements, reset, postselection, classical conditionals

Full details: docs/SUPPORTED_OPERATIONS.md and docs/RESOURCES.md.

Documentation

Development

python -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pytest

See docs/DEVELOPMENT.md for CMake iteration builds, Metal API validation runs, and the release process.

License

Apache 2.0

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages