Support public SparkGLM builds and long cold prefill - #29
Merged
Conversation
Add the measured SparkGLM managed adapter and immutable-image recipe materializer for EXL3 and NVFP4. Keep lifecycle, admission, worker-first readiness and routing in LLooM. Apply the existing extended HTTP dispatcher to chat forwarding so long cold prefill does not hit Undici's independent five-minute timeout. Provenance: - Adapted existing LLooM SparkGLM integration developed at local source 75e08ca, including its measured context settings. Existing Mia-derived MIT/Apache notices are retained in the launcher. No model weights or binaries are included. Verification: - Managed recipe test; delayed headers, streaming body and buffered transport regression; npm run check; npm test; format check; lint; interchange and package checks passed. Live SparkGLM uses this entrypoint hash and transport behavior; publishing source does not replace the installed gateway.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Support the public SparkGLM appliance and long cold prefill
Add the measured SparkGLM managed adapter and immutable-image recipe materializer for EXL3 and NVFP4. Keep lifecycle, admission, worker-first readiness and routing in LLooM. Apply the existing extended HTTP dispatcher to chat forwarding so long cold prefill does not hit Undici's independent five-minute timeout.
Provenance:
Verification: