i like to engineer systems to make them run faster or to make them run efficiently on constrained hardware.
-
SRSWTI
- NYC
-
10:26
(UTC -12:00) - www.srswti.com
- @knowrohit07
- https://huggingface.co/srswti
Highlights
- Pro
Pinned Loading
-
SRSWTI/knivesysl
SRSWTI/knivesysl Publicbare-cuda fp6 inference engine for qwen3.8-27b on one rtx 5090 — sm120 tensor-core kernels, 256k context, near-lossless perf too!
Cuda
-
SRSWTI/shooting-brake
SRSWTI/shooting-brake Publicrun a 118b moe on one 5090 and two intel b70s. 88 tok/s. yes cuda and sycl both runtimes talking to each other. beating rtx pro 6000 on 120k ctx plus.
C++
-
SRSWTI/axe
SRSWTI/axe Publicaxe - a precision agentic coder. large codebases. zero bloat. terminal-native. precise retrieval. powerful inference.
-
-
SRSWTI/shadows
SRSWTI/shadows Publica fast and lightweight distributed background task processing framework with seamless scheduling.
Python 15
-
SRSWTI/bodega-inference-engine
SRSWTI/bodega-inference-engine Publicfastest runtime for apple silicon.
If the problem persists, check the GitHub status page or contact support.



