论文

CCQ 数据集发布:覆盖美国 12 州 5.9 万条托育机构记录

CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health Research

精选理由

一个做表格或文本数据研究的朋友可以看看,5.9 万条托育记录,连代码一起开源了,拿来练跨州迁移正合适。

CCQ 数据集整合了 12 个美国州的 59,372 条托育机构记录,面向 AI 与儿童健康交叉研究。团队用 LLM 流水线对原始记录做匿名化、清洗和标准化,提供清洗后文本版和预处理表格版两个发布版本。基准测试显示,州内质量评级预测以表格版上的分类器表现最好;跨州零样本迁移接近随机,但少量目标州监督数据即可恢复大部分州内性能,且其他州预训练对微调语言模型有增益。

原文 · arXiv cs.AI

CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health Research

High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a large-scale, de-identified dataset for applied data science research at the intersection of AI and early childhood health. CCQ integrates 59,372 child care provider records across 12 U.S. states, covering diverse provider types as well as data schemas. To ensure research utility while protecting privacy, we implement an automated, LLM-based curation pipeline that anonymizes, cleans, and standardizes raw state records into two complementary releases: a cleaned textual release and a fully preprocessed tabular release. We also benchmark traditional machine learning models, tabular foundation models, and language models on quality rating prediction and important features analytics. Within a state, tabular classifiers on the preprocessed tables perform best. Across states, zero-shot transfer is near chance, but modest target-state supervision recovers most of the within-state performance, and pretraining on other states benefits finetuned language models. We release both datasets with all code to accelerate AI-driven research on child care quality and ultimately improve children's health and development.