i-mrDed
Popular repositories Loading
-
-
-
weight-streaming
weight-streaming PublicRun LLMs larger than your RAM — out-of-core local inference on consumer hardware (MoE, GGUF, llama.cpp). NVMe-as-memory with honest telemetry: real tok/s, page faults, disk traffic.
Python
Repositories
Showing 3 of 3 repositories
- weight-streaming Public
Run LLMs larger than your RAM — out-of-core local inference on consumer hardware (MoE, GGUF, llama.cpp). NVMe-as-memory with honest telemetry: real tok/s, page faults, disk traffic.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…