Invent-A-Dataset:零数据集生成评测
Invent a Dataset: Measuring dataset generation abilities with zero seed
OpenAI等五家前沿模型API评测,Invent-A-Dataset在零数据场景下生成高质量多样化数据集效果显著。
研究人员推出Invent-A-Dataset系统,用于从描述生成大规模训练数据集。该系统在八种任务类型和高达20K样本的规模下,质量领先17%,多样性领先19%。随着数据集规模扩大,多样性优势从200样本时的持平提升至20K样本时的37%相对增益。使用该系统微调的模型在不同架构下均表现更优。
Invent a Dataset: Measuring dataset generation abilities with zero seed
Building datasets remains one of the most manual and brittle parts of AI development. In this technical report, we focus on the most extreme but also most prevalent setting real world practitioners face: a zero data regime. Here, practitioners don't have any data for the capability they want to learn. We introduce Invent-A-Dataset which is a prompt based system to go from dataset description to realistic and large scale post-training datasets. We evaluate Invent-A-Dataset against five frontier model APIs including Anthropic, Google, Open AI, DeepSeek, Zai. Across eight task types and dataset sizes up to 20K samples, Invent-A-Dataset significantly outperforms with both the highest quality (17% relative gains) while simultaneously producing the most diverse samples (19% relative gains). Its diversity advantage widens with scale of training dataset size (from parity at 200 samples to 37% relative gains at 20K samples). This translates into considerable downstream training gains, resulting in far more performant post-trained models. Invent-A-Dataset fine-tune consistently ranks higher compared to other generator fine-tunes across different post-trained model architectures.