FinVector-Market-4B is an MLX-native LoRA adapter for Qwen/Qwen3.5-4B, specialized for structured financial reasoning across macro interpretation, financial question answering, calculator routing, conditional causal transmission, and multi-asset scenario analysis.
This is the sanitized public companion repository. It contains documentation, a minimal inference example, and frozen aggregate evaluation results. The private training corpus, dataset-construction and augmentation code, prompts, benchmark examples, raw model outputs, and research artifacts are not included.
- Model and MLX-LoRA adapter
- Interactive demo and benchmark explorer
- Technical report
- Evaluation summary
- Machine-readable results
The study separates three effects that are often conflated in domain-model evaluations:
- whether a base model infers an implicit output contract;
- how much an explicit JSON schema improves the base model without training; and
- what domain supervised fine-tuning adds under the same frozen benchmark and decoding setup.
It also evaluates the adapted model under both implicit and explicit prompting. That final comparison matters: the schema prompt improved some metrics but reduced others, so prompt sensitivity is part of the result rather than something hidden by reporting only the best run.
The primary controlled comparison uses the same explicit-schema prompt contract for the base and adapted models. Evaluation used a private, frozen 600-example benchmark with 75 examples in each of eight task families. Decoding was greedy, thinking was disabled, the output ceiling was 700 tokens, and no output repair was applied.
| Metric | Qwen3.5-4B | FinVector-Market-4B |
|---|---|---|
| Strict JSON validity | 91.3% | 99.8% |
| Task-schema validity | 90.3% | 99.8% |
| Policy-tone accuracy | 80.0% | 84.0% |
| Policy-tone macro-F1 | 77.4% | 62.3% |
| FinQA answer exact match | 14.7% | 40.0% |
| Tool-selection accuracy | 100.0% | 100.0% |
| Correct calculator expression | 48.0% | 82.7% |
| Calculator-answer tolerance | 8.0% | 44.0% |
| Scenario-direction agreement | 20.1% | 89.5% |
| Implication-direction agreement | 52.4% | 87.2% |
Scenario and implication scores measure agreement with synthetic structural targets, not realized market outcomes. They are not evidence of live forecasting accuracy or probability calibration. The explicit-schema adapter also did not win every metric: its policy-tone macro-F1 was lower than the base control's, and the adapted model's results changed materially across prompt contracts.
The adapter was produced with one epoch of BF16 LoRA training on a private 22,000-example mixed real/synthetic financial-reasoning corpus. The release uses LoRA rank 16 across 32 decoder layers and requires the exact base-model revision:
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
The downloadable release is an MLX-LM adapter, not a Hugging Face PEFT adapter.
Install the two runtime dependencies on Apple Silicon:
python -m pip install mlx-lm huggingface_hubThen run:
python examples/inference_mlx.pyThe example downloads the public adapter and loads it over the pinned base revision. For real applications, define and validate a task-specific schema rather than treating unconstrained prose as a stable API.
- The benchmark is private and frozen; no benchmark examples are published here.
- Strict structured-output metrics use the raw generation with no Markdown stripping or repair.
- Causal-component F1 is exact normalized list-item overlap, not semantic similarity.
- Structural probabilities in the research corpus are not empirically calibrated forecasts.
- The model and demo are research artifacts, not financial advice or trading systems.
The companion code and released adapter are provided under the Apache License 2.0. The Qwen base model remains subject to its own license and terms.