← Back to news
Hugging Face

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps