A comprehensive collection of Large Language Models optimized for Apple Silicon using MLX framework.
This repository contains configuration files and documentation for running various state-of-the-art language models on Apple Silicon devices using the MLX framework. The models are optimized for efficient inference on M1/M2/M3 chips.
- llama-3-8b-instruct - Meta's Llama 3 8B Instruct model
- llama-3.1-8b - Meta's Llama 3.1 8B base model
- llama-3.1-8b-instruct - Meta's Llama 3.1 8B Instruct model
- mistral7b-v0_1 - Mistral 7B base model v0.1
- mistral7b-instruct-v0_3 - Mistral 7B Instruct v0.3
- ministral8b-instruct-2410 - Ministral 8B Instruct (October 2024)
- qwen2-7b-instruct - Alibaba's Qwen2 7B Instruct model
- qwen2.5-7b - Alibaba's Qwen2.5 7B base model
- qwen2.5-7b-instruct - Alibaba's Qwen2.5 7B Instruct model
Each model directory contains:
config.json- Model architecture and parametersgeneration_config.json- Text generation settings (temperature, top_p, max_length)tokenizer_config.json- Tokenizer configurationtokenizer.json- Main tokenizer vocabulary and rules
chat_template.jinja- Chat formatting template (Instruct models only)special_tokens_map.json- Special token definitionsadded_tokens.json- Additional tokens (Qwen models)vocab.json&merges.txt- BPE vocabulary (Qwen models)tokenizer.model- SentencePiece model (Mistral models)
| Model | Parameters | Context Length | Temperature | Top-P |
|---|---|---|---|---|
| Llama 3/3.1 8B | 8B | 8,192 tokens | 0.6 | 0.9 |
| Mistral 7B | 7B | Variable | Default | Default |
| Qwen2/2.5 7B | 7B | Variable | 0.7 | 0.8 |
| Ministral 8B | 8B | Variable | Default | Default |
# Install MLX
pip install mlx-lm
# Or install from source
git clone https://github.com/ml-explore/mlx-examples.git
cd mlx-examples/llms
pip install -e .from mlx_lm import load, generate
# Load a model
model, tokenizer = load("path/to/model")
# Generate text
response = generate(
model,
tokenizer,
prompt="Hello, how are you?",
temperature=0.7,
max_tokens=100
)# Adjust generation parameters
response = generate(
model,
tokenizer,
prompt="Explain quantum computing",
temperature=0.3, # Lower for more focused responses
top_p=0.9, # Nucleus sampling
max_tokens=500, # Response length
repetition_penalty=1.1
)- Recommended: Llama models with
temperature=0.1-0.3 - Alternative: Qwen models for multilingual code
- Recommended: Any Instruct model with
temperature=0.8-1.2 - Best: Llama 3.1 or Qwen2.5 for coherence
- Recommended: Instruct models with
temperature=0.2-0.5 - Best: Llama 3.1-8b-instruct for accuracy
- Recommended: Any Instruct model with
temperature=0.6-0.8 - Best: Models with chat templates
- 0.1-0.3: Conservative, accurate responses
- 0.6-0.8: Balanced creativity and accuracy
- 0.9-1.2: Creative, diverse outputs
- Llama 3/3.1: 8,192 tokens (6,000-7,000 words)
- Other models: 4,096-8,192 tokens
- Recommendation: Monitor usage, summarize long conversations
MLX/
βββ models/
β βββ llama-3-8b-instruct/
β β βββ config.json
β β βββ generation_config.json
β β βββ chat_template.jinja
β β βββ tokenizer files...
β βββ qwen2-7b-instruct/
β β βββ config.json
β β βββ added_tokens.json
β β βββ vocab.json
β β βββ tokenizer files...
β βββ ...
βββ README.md
βββ .gitignore
This repository contains configuration files only. The actual model weights (.safetensors files) are not included due to their large size (15-20GB per model).
To obtain the actual model files:
-
Hugging Face Hub:
# Install huggingface-hub pip install huggingface-hub # Download models huggingface-cli download meta-llama/Llama-3.1-8B-Instruct --local-dir ./models/llama-3.1-8b-instruct
-
MLX Community:
# Use MLX-optimized models huggingface-cli download mlx-community/Llama-3.1-8B-Instruct-4bit --local-dir ./models/llama-3.1-8b-instruct-4bit
- Apple Silicon: M1/M2/M3 Mac recommended
- Memory: 16GB+ RAM for 7-8B models
- Storage: 15-20GB per model
Model configurations follow their respective original licenses:
- Llama models: Custom Meta license
- Mistral models: Apache 2.0
- Qwen models: Tongyi Qianwen license
Feel free to:
- Add new model configurations
- Improve documentation
- Share optimization tips
- Report issues
Note: This repository provides model configurations and documentation. Download actual model weights separately from official sources.