Séminaire au DIC: «Data efficiency in children and language models» par Michael C. Frank
Séminaire ayant lieu dans le cadre du Doctorat en informatique cognitive, en collaboration avec le centre de recherche CRIA
TITRE : Data efficiency in children and language models
Michael C. FRANK
Jeudi le 1er octobre 2026 à 10h30
Local PK-5115 (Il est possible d'y assister en virtuel en vous inscrivant ici)
RÉSUMÉ
Large language models show intriguing emergent behaviors, yet they receive at least three to four — and sometimes as much as six — orders of magnitude more language data than human children. What accounts for this vast difference in sample efficiency? I will describe steps towards a paradigm in which we can address this question. In particular, I'll discuss the use of egocentric video ("baby headcam") data for model training, and the use of developmental data for model evaluation. This paradigm provides a model-based framework for exploring the nature of children's early development.
BIOGRAPHIE
Michael C. FRANK is Benjamin Scott Crocker Professor of Human Biology in the Department of Psychology at Stanford University and Director of the Symbolic Systems Program. He studies children's language learning and development, with a focus on the use of large-scale datasets to understand the variability and consistency of learning across cultures. He is a founder of the ManyBabies Consortium, and has led open-data projects including Wordbank and the ongoing LEVANTE project.
RÉFÉRENCES
Frank, M. C. (2026). Children, but not language models, show accelerating returns in word learning. https://arxiv.org/abs/2608.17120
Frank, M. C., & Goodman, N. D. (2025). Cognitive modeling using artificial intelligence. Annual Review of Psychology. doi:10.1146/annurev-psych-030625-040748.
Frank, M. C. (2023). Bridging the data gap between children and large language models. Trends in Cognitive Sciences. https://psyarxiv.com/qzbgx/

Date / heure
Lieu
Montréal (QC)
Prix
Renseignements
- Mylène Dagenais
- dic@uqam.ca
- https://www.dic.uqam.ca