Research

Seven papers on the reading scale, four on speaking skills. Each paper says what its numbers rest on, even when that is synthetic test data.

Hablar

Papers on speaking skills

All accuracy figures so far come from test material: synthetic voices, written test utterances and the text of rubrics. None come from real pupils.

Agreement Alone Selects the Wrong Grader

v0.6Hablar

What the numbers rest on: Synthetic data: 24 learner utterances in Spanish and French written for the study, graded by 45 language-model deployments; every figure is an upper bound.

Read the paper (pdf)
Scored but Not Flagged in French

v0.5Hablar

What the numbers rest on: Synthetic voices: 70 clips, each as a clean version and as a twin with one planted sound error, scored by one commercial pronunciation service.

Read the paper (pdf)
Where Taggers Lose the Tense

v0.4Hablar

What the numbers rest on: Synthetic, well-formed verb forms (a grid of 1,005) plus 83 rule-generated learner errors; the taggers' figures are upper bounds.

Read the paper (pdf)
What a Machine Can Hear, Read and Judge

v0.4Hablar

What the numbers rest on: The wording of 166 criteria from 10 published speaking rubrics, not data from learners.

Read the paper (pdf)
FRS

Papers on the reading scale

The calibration does not rest on texts labelled by people, and part of the validation uses synthetic voices. Each paper says what each figure rests on.

The Free Reading Standard: Definition and Construction of an Open, Continuous Scale for Technical Reading Difficulty in Dutch and English

v0.2FRS

What the numbers rest on: A specification with two pilot experiments on 20 Dutch texts and model judgments; testing with readers is future work.

Read the paper (pdf)
Calibrating an Open Reading-Difficulty Scale without Human-Labeled Corpora: LLM Pairwise Judgments, Deterministic Features, and Their Composition

v0.1FRS

What the numbers rest on: Model judgments and text features, compared with the human-labelled English CLEAR corpus; Dutch validation beyond pilot scale is still open.

Read the paper (pdf)
Measuring Reading Growth: The Free Reading Standard and Microsoft Reading Progress in the Classroom

v0.1FRS

What the numbers rest on: A framework and a pre-registered study design, written before there were classroom data; the data come later.

Read the paper (pdf)
A Living Standard: Crowdsourced Comparative Judgment and Batched Recalibration of an Open Reading-Difficulty Scale

v0.2FRS

What the numbers rest on: Simulations (1,080 runs) with noise taken from the scale's own judgment records; the first real recalibration round is still to come.

Read the paper (pdf)
Measuring Oral Reading Fluency with a Dual-Pass ASR Architecture: Design, Instrument Characterization, and Synthetic Validation of the FRS Fluency Meter

v0.1FRS

What the numbers rest on: Synthetic recordings with planted reading errors (144 items); recordings of real children are follow-up work.

Read the paper (pdf)
Estimating Reader Ability on an Open Text-Difficulty Scale from Oral Reading Performance: A Conjoint Model with Simulated Uncertainty Behavior

v0.1FRS

What the numbers rest on: Simulations of readers (225 cells), not recordings of real readers.

Read the paper (pdf)
Synthetic Voices as Validation Instruments for Automated Oral-Reading-Fluency Scorers: Method, Demonstration, and Limits

v0.1FRS

What the numbers rest on: Synthetic voices reading texts with planted errors; the paper itself says this cannot show accuracy on children's speech.

Read the paper (pdf)