by Zutfen LLC
InferSwarm is an open-source heterogeneous inference fabric for turning disparate compute and memory resources into one logical inference platform.
Many machines. One model.
Turn the hardware you already own into distributed inference capacity.
Status: Research / Proof of Concept
InferSwarm is an experimental Apache-2.0 project intended to let heterogeneous resources cooperate on inference without requiring every device or machine to look like the same kind of worker.
There is no released production InferSwarm runtime today. The repository is the canonical home for architecture decisions, the normative Fabric Doctrine, benchmark/evidence records, and the current evidence-gated roadmap. Runtime experiments continue in the Zutfen FreeToken fork.
Repository precedence is:
ADRs decide; the Fabric Doctrine specifies; ARCHITECTURE.md explains; ROADMAP.md sequences.
ADR 0008 adopts the current resource/residency/planning model after the completed Wayfinder (#37, decisions #38-#46).
The architecture is doctrine-shaped, API-unfrozen: current implementation must preserve the doctrine's semantics, but final public planner/strategy type names, plugins, wire protocols, and storage schemas are deliberately deferred until real implementations prove the seam.
| Capability | Evidence scope | Demonstrated result |
|---|---|---|
| Automatic planning | Physical | Strategy-constrained selection among legal local plans using applicable evidence and operator policy. Evidence; acceptance. |
| Distributed serving | Physical | A normal request drives planning, multi-Node realization, and backend-native execution on the tested Qwen path. Evidence; acceptance. |
| Plan epochs and recovery | Physical | Resource changes and recovery preserve plan-epoch authority and reject retired work on the tested serving path. Evidence; acceptance. |
| Separate Coordinator | Physical | A CPU-only Coordinator handles ingress, planning, epoch coordination, and committed output without hosting model execution. Evidence; acceptance. |
| Selective artifact acquisition | CPU fixture | Participants acquire only required model artifacts in the model-independent acquisition proof. Evidence; acceptance. |
| Peer reuse and replacement deltas | CPU fixture | Verified inventory, source selection, peer publication, and replacement-plan acquisition work across participants. Evidence; acceptance. |
| Locality-aware transition planning | CPU fixture | Verified artifact locality affects transition ranking without changing technical feasibility or execution ranking. Evidence; acceptance. |
| Dense Gemma numerical qualification | Physical | The frozen Gemma subject passed V5 numerical and semantic qualification; applicability remains specific to that subject. Evidence; acceptance. |
- Research / proof of concept; no released production runtime.
- Physical results apply to their tested model, backend, hardware, and topology. CPU fixture proofs do not establish physical integration.
- Public planner/strategy APIs, wire formats, and storage schemas remain unfrozen; broad vendor support remains an objective.
- Historical Phase 1 NO-GO and R6 failure remain unchanged. GLM-5.3-Flash / #13 is a later falsifier.
Earlier work established selective loading, accelerator residency without an
unexplained persistent host mirror, and local/multi-Node execution on Qwen.
Canonical Phase 1 retained a scoped NO-GO performance verdict; subsequent
Phase1R experiments established topology-dependent performance and capacity
tradeoffs. See ROADMAP.md and the
historical Phase1R record
for the exact experiments and immutable results.
Issue #117 — R6 successor dense full integration
Integrate V5-qualified dense Gemma with automatic planning, participant-exact artifact acquisition, selective materialization, and ordinary fenced serving.
-
Arm B — canonical cold acquisition + realization observation:
ISSUE117_ARM_B_COLD_REALIZATION_PASS. -
Maintainer acceptance: pending maintainer acceptance.
-
Recorded execution authorization: blocked — Arm C — ordinary external-Coordinator serving (blocked on Arm-B maintainer acceptance). Authority.
-
Arm B is observed ISSUE117_ARM_B_COLD_REALIZATION_PASS pending maintainer acceptance; Arm C must not begin before acceptance.
-
Arm A is accepted at merge 6774474941d7ce2a0252c8c1e148f8bce61a8d6d; the Arm-B observation executed from accepted main 5179c41232051e7455b778ddb8876a6539f4cb04.
-
Do not rerun accepted Arm A, the accepted physical preflight, or the observed Arm-B campaign.
-
Use the exact producer, checkpoint, Compute Unit, cold-root, Source, and zero-invariant requirements specified in issue #117.
-
No consumed h109 holdout material may be used as new evidence.
InferSwarm aims to make resources such as:
- NVIDIA GPUs;
- AMD GPUs;
- Intel GPUs;
- CPUs;
- GPU VRAM / HBM;
- system RAM;
- multiple GPUs with asymmetric local links;
- multiple machines connected over ordinary Ethernet;
- future useful backing/memory resources such as NVMe or CXL where evidence supports them;
available to one logical planning domain, with decisions driven by model semantics, measured capability, state requirements, workload demand, and operator policy rather than assumed hardware symmetry.
These are a concise overview; the Fabric Doctrine is normative.
- Heterogeneity is first-class. Vendor, generation, memory size, compute speed, bus topology, and network speed may differ.
- Resources do not have permanent plan roles. There is no canonical
primary/secondaryGPU or L0/L1/L2/L3 hierarchy. A GPU, CPU, RAM domain, or link participates according to the current plan. - System RAM and CPU remain first-class. They may provide residency, execution, staging, cache/replica value, or no active role depending on the plan; accelerators augment rather than deprecate them.
- State identity and physical copies are different things. Logical state, materializations, backing, residency, staging, cache, replica, execution location, and mutable authority are distinct.
- Accelerator residency does not imply a host mirror. Persistent host copies require an explicit purpose and accounting.
- Correctness and feasibility precede optimization. Slow-but-viable is still viable unless an explicit operator service requirement says otherwise.
- Measure hardware; do not stereotype it. Context-valid measurements drive economics; unknown is uncertainty; correctness failures quarantine rather than merely reduce a performance score.
- Model semantics stay behind a Model Execution Strategy. Strategies
define legal opaque state/execution units and boundaries; the generic
planner chooses among them without needing concepts such as
expert, Qwen, CUDA Graph, or NVFP4. - Granularity is measured and plan-relative. High-frequency/dependency- sensitive communication should stay on the lowest-cost measured locality practical, but coarse boundaries have costs too. Intra-node and inter-node granularities may differ.
- Elasticity works both directions. Better resources may be prepared and folded into active sessions at safe boundaries; resource loss should fall back to any correct feasible surviving plan—including slower GPUs or CPU/RAM—before declaring outage.
- InferSwarm can adapt to structural demand. Model/profile/Swarm/session history may inform future placement without requiring prompt/response retention or assigning human meanings to model parts.
- Commodity networking matters. Ordinary 1 Gigabit Ethernet remains the baseline network target. Faster networks are welcome optimizations, not a mandatory project dependency.
- Execution fabric and management plane are distinct. The open-source fabric must remain fully usable without a paid control plane.
- Keep the host integration seam narrow. FreeToken is the first proving vehicle, not the product boundary.
Host inference engine
|
v
Model Execution Strategy
|
| legal opaque units / state / demand /
| representations / correctness / economics
v
Generic InferSwarm planner
|
| Swarm resource graph + evidence + policy
v
Versioned Execution Plan / epoch
|
+-------------------+-------------------+
| | |
Compute Units Memory Resources Links/paths
GPU/CPU/NPU/... RAM/VRAM/HBM/... local/network
The resource graph describes what InferSwarm has. The Execution Plan describes what InferSwarm intends to do with it.
MoE expert execution remains the first strategy actually researched and is preserved by ADR 0004.
ADR 0007 remains accepted as the first network strategy/evidence direction: coarse contiguous model blocks over ordinary Ethernet. It is not a permanent rule that inter-node execution must use contiguous blocks. The current doctrine allows the Model Execution Strategy and planner to select another legal intra/inter-node granularity when measurements justify it.
The FreeToken fork is the initial runtime vehicle.
Its durable integration branch is inferswarm-research (established by #59).
Upstream-tracking main and immutable evidence branches have separate roles.
Execution uses the exact producer named by the current gate authority,
never an unreviewed branch tip. FreeToken is not the permanent product boundary.
See docs/integrations/freetoken.md.
.github/ issue templates, pull request template, CI
docs/ documentation map and reading rules
docs/adr/ architecture decision records
docs/architecture/ normative Fabric Doctrine and its supplements
docs/benchmarks/ benchmark methodology/results
docs/investigations/ research inputs and feasibility work
docs/implementation/ active/historical experiment plans and handoffs
docs/protocols/ semantic-boundary and transport design notes
docs/qualification/ heterogeneous correctness-qualification lane (v1-v5)
docs/integrations/ host-engine integration notes
scripts/ CPU-only, deterministic, fail-closed evidence tooling
tests/ tests that guard the retained evidence
ARCHITECTURE.md derived architecture overview
BENCHMARKING.md benchmark/evidence contract
ROADMAP.md evidence-gated successor roadmap
docs/README.md maps the whole documentation tree and states
the rules for reading retained evidence. scripts/README.md
and tests/README.md cover the tooling and how to verify the
repository locally.
Apache License 2.0. See LICENSE. Contributions are accepted under the same license; see CONTRIBUTING.md.