Calibration
A fixed, mixed-difficulty opening samples the four dimensions and initializes the proficiency estimate used during delivery.
AILAT is an IRT-informed adaptive assessment designed as a formative diagnostic. This page separates the implemented measurement approach from empirical claims that the current evidence does not yet support.
Current evidence position: AILAT has not published an instrument-specific validation study or population norms.
Evidence ledger
What is established—and what is not
AILAT samples conceptual understanding, practical use, critical evaluation and creation, and ethics. These dimensions are informed by published AI-literacy frameworks, then operationalized through AILAT’s own item bank and scoring rules.
AI fundamentals, system behaviour, limits, and terminology.
Prompting, workflow integration, and practical use of AI tools.
Judging output quality, identifying limitations, and improving AI-assisted work.
Bias, privacy, accountability, transparency, and responsible adoption.
Only the middle phase uses adaptive item selection. Describing every question as adapting to every answer would overstate the design.
A fixed, mixed-difficulty opening samples the four dimensions and initializes the proficiency estimate used during delivery.
Eight multiple-choice items are selected using 2PL item-information concepts; two open-ended prompts sample reasoning that fixed choices can miss.
Four multiple-choice items and one open-ended capstone place the same literacy dimensions in the participant’s selected industry context.
During the adaptive phase, AILAT uses a two-parameter logistic model and Fisher information to choose informative multiple-choice items around the running proficiency estimate.
The present item parameters are part of the authored assessment design. They should not be described as empirically calibrated until a calibration study has been completed and published.
Multiple-choice correctness and fractional credit from rubric-scored open-ended responses produce dimension percentages. The overall score then weights those dimensions.
The current evidence supports cautious, formative use. The named levels describe score bands inside AILAT; they are not population percentiles or external certifications.
Publishing these analyses would allow future claims to become more specific. Until then, AILAT’s public language should remain explicit about the distinction between theoretical foundation and instrument-specific evidence.
Estimate difficulty and discrimination from a sufficiently large, representative response sample instead of treating authored parameters as empirical estimates.
Report conditional measurement precision, internal consistency where appropriate, and stability across repeated administrations.
Test whether the four-dimension structure fits observed data and how scores relate to relevant external measures or performance.
Evaluate differential item functioning and subgroup performance before making broad comparability claims.
Compare model-assigned partial credit with independent human ratings and monitor agreement when scoring models or rubrics change.
Citing an established framework does not validate a new assessment automatically. These sources help explain the design context; AILAT must generate its own empirical evidence.
Long & Magerko (2020)
A competency-oriented definition and design framework for AI literacy.
Open sourceCarolus et al. (2023)
A self-report AI-literacy scale grounded in competency models. AILAT is not MAILS and does not inherit its validation evidence.
Open sourceDing, Kim & Allday (2024)
Performance-oriented assessment research for a non-technical population. AILAT is a separate instrument.
Open sourceThis page is a product evidence statement, not a peer-reviewed publication. We will update the status when AILAT-specific evidence is available.