Fine-grained computation offload for off-the-shelf servers in tens of lines — paper (arXiv:2607.02630), code, and every measurement. Overlap accelerator offloads (GPU/HSM/inference) with other requests by rerouting through the server's own suspend/resume machinery, plus a zero-edit LD_PRELOAD fiber runtime.
nginx redis coroutines concurrency cuda fibers high-performance-computing low-latency gpu-computing research-paper ld-preload systems-research accelerator-offload latency-hiding
-
Updated
Aug 23, 2026 - C