Convex Markets / Datasets / Preference data / Skywork-Reward-Preference-80K
SKU PRF-3009 · Sold by External

Skywork-Reward-Preference-80K

Product specifications

SKUPRF-3009
Data typePreference data
Volume77016 preference pairs (chosen/rejected chat pairs)
Size on disk~209 MB (Parquet download) / ~416 MB uncompressed
FormatParquet
Access modelPUBLIC LICENSE
PricingFree · open dataset (HF)
Quality score
LicenseNo license published by source
Skywork-Reward-Preference-80K-v0.2 is a subset of ~80K human/model preference pairs assembled by Skywork (Kunlun Inc.) to train the Skywork-Reward-Gemma-2-27B-v0.2 and Skywork-Reward-Llama-3.1-8B-v0.2 reward models. Each record is a (chosen, rejected) pair of multi-turn chat message lists plus a source label indicating which public dataset it came from. The pairs are subsampled (with no other modification) from HelpSteer2, OffsetBias, WildGuard (adversarial), and the Magpie DPO series (Ultra, Pro Llama-3.1, Pro, Air); Magpie samples were selected by average ArmoRM score and the WildGuard subset was additionally filtered by a Skywork reward model so that the chosen response scores higher than the rejected one. Version v0.2 is the decontaminated release: 4,957 magpie-ultra-v0.1 pairs with significant n-gram overlap against RewardBench evaluation prompts were removed. The dataset is a single Parquet train split of 77,016 rows totaling ~416 MB uncompressed (~209 MB download).