SKU SFT-2407 · Sold by External
Aya Dataset
Product specifications
| SKU | SFT-2407 |
|---|---|
| Data type | SFT dataset |
| Volume | 204,112 prompt-completion pairs (202,362 train + 1,750 test) |
| Size on disk | ~275 MB (Parquet download, default config; uncompressed ~256 MB) |
| Format | Parquet |
| Access model | PUBLIC LICENSE |
| Pricing | Free · open dataset (HF) |
| Quality score | — |
| License | Apache License 2.0 |
The Aya Dataset is a multilingual instruction fine-tuning (SFT) dataset of 204,112 human-annotated prompt-completion pairs curated by an open-science community through the Aya Annotation Platform from Cohere Labs (formerly Cohere For AI). It spans 65 languages (71 including dialects) and, unlike many instruction datasets, the completions are written and reviewed by native/fluent speakers rather than machine-generated. Each record contains an instruction (inputs), a reference completion (targets), the language name and its ISO code, an annotation_type flag distinguishing original annotations from re-annotations, and a hashed annotator user_id. The default config is split into a train split of 202,362 examples and a test split of 1,750 examples; a separate demographics config holds self-reported metadata for 1,456 annotators. It is released under the Apache License 2.0 and is intended to train, fine-tune, and evaluate multilingual LLMs.