Convex Markets / Datasets / SFT dataset

Dataset catalog

Filter by type, access, and pricing. Specs show before you open the product page.

39 results
U2
External · SFT-2004

UltraChat 200k

HuggingFaceH4's heavily filtered subset of UltraChat (~208k train_sft dialogues) used to train the Zephyr chat models.

Type SFT datasetVolume 207865 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
M
External · SFT-2203

MetaMathQA

395k math QA pairs bootstrapped from GSM8K and MATH via rephrasing, self-verification and backward (FOBAR) augmentation; text solutions only.

Type SFT datasetVolume 395k QA pairsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
M
External · SFT-2207

MathInstruct

262k math instruction examples combining chain-of-thought and program-of-thought (executable Python) rationales compiled from 13 datasets.

Type SFT datasetVolume 262k instruction-response examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
Ev
External · SFT-2302

Evol-CodeAlpaca v1

111K English code instruction/output pairs — an Evol-Instruct augmentation of CodeAlpaca-20k using GPT-4 across 10 evolution strategies.

Type SFT datasetVolume 111272 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
D
External · SFT-2108

Dolphin

Open Orca-style FLAN reproduction: ~892K FLANv2 completions from GPT-4 plus ~2.84M from GPT-3.5, filtered to remove refusals/alignment.

Type SFT datasetVolume 3731947 examplesFormat JSONL (Parquet auto-conversion available)Access PUBLIC LICENSE
Free
open-source license
View
E
External · SFT-2303

Evol-Instruct-Code-80k-v1

78K code instruction/output pairs — an open reproduction of WizardCoder's Evol-Instruct-Code (CodeAlpaca run through 3 evolution rounds).

Type SFT datasetVolume 78264 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2405

OpenOrca

An open collection of ~2.94M FLAN-Collection instructions paired with GPT-4/GPT-3.5 'Orca-style' reasoning-trace responses for supervised instruction fine-tuning.

Type SFT datasetVolume 2,942,029 records (instruction-response pairs, single train split)Format ParquetAccess PUBLIC LICENSE
Free
open dataset (HF)
View
WE
External · SFT-2102

WizardLM Evol-Instruct V2 196k

Complexity-evolved (Evol-Instruct) instruction conversations derived from Alpaca/ShareGPT for WizardLM SFT; HF split ships 143K rows.

Type SFT datasetVolume 143000 examplesFormat JSON (Parquet auto-conversion available)Access PUBLIC LICENSE
Free
open-source license
View
S
External · SFT-2402

Self-Instruct

Machine-generated instruction-following dataset created by bootstrapping GPT-3 from 175 seed tasks, released as ~82K prompt/completion demonstrations for instruction tuning.

Type SFT datasetVolume 82612 instances (prompt–completion pairs) in the self_instruct configFormat Text; HF configs served as prompt/completion string pairs (Parquet via datasets-server); original repo ships JSON/JSONLAccess PUBLIC LICENSE
Free
open dataset (HF)
View
L
External · SFT-2008

LIMA

GAIR's 1,030-example gated dataset of high-quality curated prompts and responses used in the 'Less Is More for Alignment' study.

Type SFT datasetVolume 1030 rowsFormat Parquet (Hugging Face)Access GATED
Free
open-source license
View
Oo
External · SFT-2006

OpenAssistant oasst1

Crowdsourced multilingual assistant conversation trees with human quality ratings; HF splits are 84.4k train / 4.4k validation messages.

Type SFT datasetVolume 84437 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
DD
External · SFT-2007

Databricks Dolly 15k

~15k human-authored instruction/response records across 8 task categories, written by Databricks employees for commercial-friendly instruction tuning.

Type SFT datasetVolume 15011 rowsFormat JSONLAccess PUBLIC LICENSE
Free
open-source license
View
C2
External · SFT-2304

CodeAlpaca 20K

20K single-function code instructions (instruction/input/output) generated via Self-Instruct with text-davinci-003 from 21 seed tasks.

Type SFT datasetVolume 20022 rowsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
SA
External · SFT-2101

Stanford Alpaca (and Alpaca-Cleaned)

52K single-turn English instruction-following examples generated from OpenAI text-davinci-003 via Self-Instruct; cleaned mirror fixes ~51.8K rows.

Type SFT datasetVolume 52002 examplesFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
FC
External · SFT-2106

FLAN Collection (Flan v2)

Community re-release of Google's Flan 2022 Collection: 1,836 tasks across Flan/T0/NIV2/CoT/Dialog in zero/few-shot × option/no-option formats.

Type SFT datasetVolume examplesFormat Gzip-compressed JSONL (Parquet auto-conversion available)Access PUBLIC LICENSE
Free
open-source license
View
S
External · SFT-2401

Super-NaturalInstructions

A benchmark of 1,616 diverse NLP tasks, each paired with an expert-written declarative instruction (Definition) plus positive/negative demonstration examples and many input/output instances, used for instruction-tuning and cross-task generalization.

Type SFT datasetVolume 1,616 NLP tasksFormat JSON task files (repo) / JSONL instance rows (HF mirror)Access PUBLIC LICENSE
Free
open-source license (Apache-2.0)
View
UI
External · SFT-2403

Unnatural Instructions

A large instruction-tuning dataset of ~68K instruction-input-output triplets (240K with paraphrases) generated almost entirely by OpenAI's text-davinci-002 from three seed examples, with virtually no human labor.

Type SFT datasetVolume 68,478 instruction-input-output triplets (core set; 240,670 examples in the full set with paraphrases)Format JSONLAccess PUBLIC LICENSE
Free
open-source license (MIT)
View
N
External · SFT-2202

NuminaMath-1.5

~896k competition-math problems (Chinese high-school to IMO level) with chain-of-thought solutions; improved successor to NuminaMath-CoT.

Type SFT datasetVolume 896k problemsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View
O
External · SFT-2201

OpenR1-Math-220k

220k competition math problems, each with 2-4 DeepSeek-R1 chain-of-thought traces verified by Math-Verify and Llama-3.3-70B; text reasoning, no code.

Type SFT datasetVolume 220k problemsFormat Parquet (Hugging Face)Access PUBLIC LICENSE
Free
open-source license
View