metadata
pipeline_tag: text-generation
inference: true
license: apache-2.0
datasets:
- simplescaling/s1K-r1
Model Summary
s1 is a reasoning model finetuned from Qwen2.5-32B-Instruct on just 1,000 examples. It matches o1-preview & exhibits test-time scaling via budget forcing.
- Repository: simplescaling/s1
- Paper: https://arxiv.org/abs/2501.19393
This model is a successor of s1-32B with slightly better performance. Thanks to Ryan Marten for helping generate r1 traces for s1K.
Use
The model usage is documented here.