With the aim of disseminating studies in Language Acquisition, LLAS makes three corpora on children's speech, writing, and linguistic environment available to the community of linguists, teachers, and researchers. These data aim to provide empirical support for scientific, academic, and educational studies on child language development.

The oral language database consists of two files containing the most frequent words in child speech (FI Corpus, 200,000 tokens) and in child-directed speech (FDC Corpus, 400,000 tokens), providing information on their segmental, syllabic, stress, and morphological composition, with a constantly expanding volume of words. The database on child writing, in turn, is composed of a set of digitized notebooks belonging to children from the 1st to the 8th grade of Elementary School. These materials are readily available upon completion of a registration form.

To access the corpora, click on the links above.