Skip to content

Repository files navigation

Coloma

The Continuous Verification and Optimization layer for self-hosted LLMs

cover

license Discord X

Features

  • Tune vLLM deployments to maximize the metrics you care about, and prevent runtime crashes.
  • Compare your tuned models, and deploy them.
  • Drop-in OpenAI-compatible proxy: sits between your workload and your endpoint.
  • Live and saved traffic inspection.
  • Image optimization: optimizes base64 images before they hit your GPU, shrinking payload with no client changes.
  • Structured-output schema validation and custom Python validators.
  • Idle-compute verification: re-runs a heavier check on past predictions when the proxy is idle, no added latency.
  • Chat interface.
  • Runs 100% on-prem, no data collected / sent to third-party, except if you use an LLM provider instead of a self-hosted model.

Prerequisites

Quickstart

cp backend/.env.template backend/.env  # Then edit the file
uv sync
npm i
npm run build
npm run start  # Exposes on http://localhost:4173 and http://0.0.0.0:4173

Then head to Profile & deploy, after profiling and deploying, enjoy the Chat tab, proxy for your OpenAI client, and Traffic tab.

Architecture

architecture

Configuration

  • COLOMA_API_KEY protects both the dashboard and proxy
  • If you're using Coloma to deploy models, the API key field in Proxy settings is used to protect the OpenAI endpoint of your deployed model
  • If you're using Coloma to point to LLM providers (e.g. OpenAI API), set your OpenAI API key in Proxy settings

Roadmap

  • Better support for VLM profiling.
  • Reasoning models support, with higher reasoning effort on verification passes.
  • Latency-aware fidelity tuning: profile per-domain latency and accuracy, then hold the highest image fidelity that fits your end-to-end latency budget.
  • Feedback fusion: support human feedback, merge with schema validation and verification.
  • Track inference server metrics.

Community

Come hangout on our Discord!

Running open-weight inference in production? I want to hear what you're running! @tschillaciml thomas@colomalabs.ai

Leave a star if this is helpful ♥️

About

The Continuous Verification and Optimization layer for self-hosted LLMs

Topics

Resources

Stars

13 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages