Give JevAlt a situation, a question and answer options. Its three 4B language models choose between the options and return a probability for each, for tasks such as routing support tickets or checking a written policy. They run locally so these decisions can be inspected and tested on your own data.
Try an example or run a model below.
Requires Python 3.11 or newer. The Q4_K_M GGUF builds run on a CPU in about 3 GB of RAM; a GPU is optional.
pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serveThis downloads Deem-4B and starts a server at http://127.0.0.1:8000. In another terminal, send a support ticket:
curl -s http://127.0.0.1:8000/v1/systemone -H 'Content-Type: application/json' -d '{"state":"I was charged twice for March. Please refund the duplicate.","questions":{"team":{"type":"choice","instructions":"Which team should handle this ticket?","criteria":{"billing":"payments, refunds","technical":"bugs, outages","sales":"prices, contracts"}}}}'The server implements the Jev API (POST /v1/systemone), so existing Jev clients can use it. API and configuration docs cover question types, optional reasoning and the unknown answer.
Karar-4B is one of the best open Turkish decision models at 4B parameters, and each JevAlt model beats Kev-4B, a same-size baseline, by about ten points in its language. Accuracy on held-out decisions:
| Model | Language | Accuracy | Kev-4B |
|---|---|---|---|
| Deem-4B | English | 94.7% | 84.7% |
| Karar-4B | Turkish | 96.8% | 87.1% |
| Wähler-4B | German | 92.0% | 81.1% |
These tests come from JevAlt's own data pipeline and favour JevAlt. Results and evaluation records include calibration scores and comparisons with Laya and the starting checkpoint.
CPU downloads: Deem-4B-GGUF, Karar-4B-GGUF, Wähler-4B-GGUF.
All three are fine-tuned from Intern-Decision-4B. To serve Turkish:
jevalt serve --model mertkayacs/Karar-4B-GGUF --file Karar-4B-Q4_K_M.gguf- Confidence is calibrated on held-out project data, so the probabilities are fitted to match how often the model is right on that data. Refit calibration on your own data with JevOss before setting decision thresholds.
- Long, noisy text and date calculations remain weak spots. Kev-4B and Laya lose less accuracy under long padding; the starting checkpoint leads on JevBench-hard.
- The models still follow some hidden instructions planted in input text. Test this behaviour before using their answers to trigger actions.
See the reasoning study and JevOss failure analysis for methods and limits. Separate live tests cover 390 requests across the three languages.
- Training data and benchmark data.
- Training and export code for fine-tuning, evaluation and GGUF conversion.
- Emberwick: play a village simulation whose inhabitants' actions are chosen by these models.
- Project site for model pages and demonstrations.
Code and weights: Apache-2.0. JevAlt is independent of TypeSafe AI; Jev is a TypeSafe AI model.
BibTeX
@software{kaya2026jevalt,
author = {Mert Kaya},
title = {JevAlt: Open Decision Models with the Jev API},
year = {2026},
license = {Apache-2.0},
url = {https://github.com/mertkayacs/jevalt}
}
An Eschatia Labs project by Mert Kaya.


