Local Inference Lab
Popular repositories Loading
-
llm-inference-bench
llm-inference-bench PublicLLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
-
blackwell-llm-docker
blackwell-llm-docker PublicDocker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
-
Repositories
- blackwell-llm-docker Public
Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
- llm-inference-bench Public
LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
- InstantTensor Public Forked from voipmonitor/InstantTensor
An ultra-fast, distributed Safetensors loader
-
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Most used topics
Loading…