Research
Seven papers on the reading scale, four on speaking skills. Each paper says what its numbers rest on, even when that is synthetic test data.
Papers on speaking skills
All accuracy figures so far come from test material: synthetic voices, written test utterances and the text of rubrics. None come from real pupils.
What the numbers rest on: Synthetic data: 24 learner utterances in Spanish and French written for the study, graded by 45 language-model deployments; every figure is an upper bound.
Read the paper (pdf)What the numbers rest on: Synthetic voices: 70 clips, each as a clean version and as a twin with one planted sound error, scored by one commercial pronunciation service.
Read the paper (pdf)What the numbers rest on: Synthetic, well-formed verb forms (a grid of 1,005) plus 83 rule-generated learner errors; the taggers' figures are upper bounds.
Read the paper (pdf)What the numbers rest on: The wording of 166 criteria from 10 published speaking rubrics, not data from learners.
Read the paper (pdf)Papers on the reading scale
The calibration does not rest on texts labelled by people, and part of the validation uses synthetic voices. Each paper says what each figure rests on.
What the numbers rest on: A specification with two pilot experiments on 20 Dutch texts and model judgments; testing with readers is future work.
Read the paper (pdf)What the numbers rest on: Model judgments and text features, compared with the human-labelled English CLEAR corpus; Dutch validation beyond pilot scale is still open.
Read the paper (pdf)What the numbers rest on: A framework and a pre-registered study design, written before there were classroom data; the data come later.
Read the paper (pdf)What the numbers rest on: Simulations (1,080 runs) with noise taken from the scale's own judgment records; the first real recalibration round is still to come.
Read the paper (pdf)What the numbers rest on: Synthetic recordings with planted reading errors (144 items); recordings of real children are follow-up work.
Read the paper (pdf)What the numbers rest on: Simulations of readers (225 cells), not recordings of real readers.
Read the paper (pdf)What the numbers rest on: Synthetic voices reading texts with planted errors; the paper itself says this cannot show accuracy on children's speech.
Read the paper (pdf)