s1.1-32B / README.md
Muennighoff's picture
Update README.md
5c3b7dd verified
|
raw
history blame
733 Bytes
metadata
pipeline_tag: text-generation
inference: true
license: apache-2.0
datasets:
  - simplescaling/s1K-r1

Model Summary

s1 is a reasoning model finetuned from Qwen2.5-32B-Instruct on just 1,000 examples. It matches o1-preview & exhibits test-time scaling via budget forcing.

This model is a successor of s1-32B with slightly better performance. Thanks to Ryan Marten for helping generate r1 traces for s1K.

Use

The model usage is documented here.