Skip to content
@local-inference-lab

Local Inference Lab

Popular repositories Loading

  1. rtx6kpro rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    Python 1.1k 83

  2. b12x b12x Public

    Python 267 68

  3. llm-inference-bench llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    Python 105 17

  4. blackwell-llm-docker blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    Python 83 23

  5. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 41 32

  6. quant-toolkit quant-toolkit Public

    Python 17 9

Repositories

Showing 10 of 18 repositories
  • b12x Public
    local-inference-lab/b12x's past year of commit activity
    Python 267 Apache-2.0 68 59 73 Updated Sep 30, 2026
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    local-inference-lab/vllm's past year of commit activity
    Python 41 Apache-2.0 23,075 86 226 Updated Sep 29, 2026
  • blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    local-inference-lab/blackwell-llm-docker's past year of commit activity
    Python 83 23 3 8 Updated Sep 30, 2026
  • LMCache Public Forked from LMCache/LMCache

    LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

    local-inference-lab/LMCache's past year of commit activity
    Python 0 Apache-2.0 1,994 0 31 Updated Sep 29, 2026
  • fleet Public
    local-inference-lab/fleet's past year of commit activity
    Go 0 1 0 0 Updated Sep 29, 2026
  • lil Public

    Typed local and Spark/RDMA vLLM launcher for Local Inference Lab

    local-inference-lab/lil's past year of commit activity
    Go 3 Apache-2.0 3 0 2 Updated Sep 29, 2026
  • rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    local-inference-lab/rtx6kpro's past year of commit activity
    Python 1,114 83 59 19 Updated Sep 29, 2026
  • llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    local-inference-lab/llm-inference-bench's past year of commit activity
    Python 105 17 5 4 Updated Sep 28, 2026
  • InstantTensor Public Forked from voipmonitor/InstantTensor

    An ultra-fast, distributed Safetensors loader

    local-inference-lab/InstantTensor's past year of commit activity
    C++ 0 Apache-2.0 18 0 1 Updated Sep 14, 2026
  • flashinfer Public Forked from flashinfer-ai/flashinfer

    FlashInfer: Kernel Library for LLM Serving

    local-inference-lab/flashinfer's past year of commit activity
    Python 0 Apache-2.0 1,514 0 2 Updated Sep 14, 2026

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Most used topics

Loading…