The methodology behind the assessment

A score should
show its work.

How is my child’s work actually graded?

RALE connects student work to evidence, feedback and a rubric score. A teacher reviews the assessment before it becomes a published report.

Follow the assessment
From a real report / Precision AuditView full excerpt ↗
Precision Audit: original student writing, its correction, and the grammar or lexical skill code with an explanation.
Read across the row

The original.
The correction.
The reason.

The student’s words stay beside the feedback. The skill code identifies what to work on next.

GRAM / AGREEMENT / SVA
  1. 01Student work
  2. 02Evidence
  3. 03Skill-tagged feedback
  4. 04Rubric score
  5. 05Teacher review
  6. 06Report

RALE · Rubric Assessment Learning Engine · Writing, speaking, reading and listening inside Skyonomy.

The assessment domains

Different work.
Different evidence.

The evidence ledger

What we can prove.

Writing · ELLIPSE

ρ .846

Score correlation

N=100. Evidence for writing on the Standard scale. Correlation is not a percentage of exact agreement.

Essay-level CEFR · ICNALE

94%

Within one band

37% exact. The band estimate has a margin of uncertainty; it is not an official exam result.

An excerpt from the Scale Trust Ledger. These writing findings do not establish speaking, reading or listening accuracy.

Read the full validation ledger
Source recorddocs/reference/SCALE_TRUST.md

Say these out loud

The caveats belong
beside the claims.

A

CEFR is ±1 band at essay level.

The corpora themselves overlap that much. C1/C2 thresholds are provisional, pending the 500-essay top-end run after recalibration.

B

Absolute grades can be harsh in a classroom.

ELLIPSE’s best essays (about B2) map to US letter “D+”. That is correct on the absolute scale. Schools wanting classroom-relative letters can use per-school custom_boundaries.

C

TOEFL conversions diverge at the extremes.

The rubric-mean path and ETS’s IELTS↔TOEFL population concordance disagree at the low and high ends. The mid-range (B2) agrees.

D

The taxonomy count was corrected.

The public impact field skills_per_scan was corrected from 100 to 50. A list of skills is not evidence that every skill is measured in every submission.

A shared language for feedback

Name the skill.
Keep the evidence.

The source defines 50 skills across these domains, including strengths and receptive comprehension. A code is a label, not proof that a diagnosis is correct.

Explore the complete taxonomy
Source recordmodules/evaluation/taxonomy_v1.yaml · version 3

From method to classroom

See what this looks like
for a learner.