SKU SFT-2310 · Sold by External
Nemotron-Post-Training-Dataset-v1
Product specifications
| SKU | SFT-2310 |
|---|---|
| Data type | SFT dataset |
| Volume | 25659642 rows |
| Size on disk | 203 GB |
| Format | Parquet (Hugging Face) |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open-source license |
| Quality score | — |
| License | CC-BY-4.0 |
NVIDIA's Nemotron-Post-Training-Dataset-v1 is a 25,659,642-row post-training corpus with per-split counts: chat 746,622; code 1,896,395; math 2,044,407; stem 20,662,167; tool_calling 310,051. Synthetic responses were generated by DeepSeek-R1-0528 and Qwen3-235B-A22B. Its standout feature is the tool_calling split (~310K rows): unlike the code-interpreter reasoning found elsewhere, these rows contain genuine structured function-calling traces — a tools specification plus assistant messages carrying tool_calls, covering single-turn, multi-turn and multi-step tool use. This is the rare authentic function-calling SFT data in this catalog.