You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Teaching Llama 3.2 1B emotional self-awareness via GRPO and a real emotion classifier. A reproducible, low-cost pipeline for calibrating an LLM's affective state and aligning its response tone.
About
Low-cost GRPO fine-tuning pipeline for valence-arousal state estimation and tone-aligned responses in Llama 3.2 1B.