Skip to main content
Method & validation

Transparent about the method, the evidence, and what remains to be validated.

AILAT is an IRT-informed adaptive assessment designed as a formative diagnostic. This page separates the implemented measurement approach from empirical claims that the current evidence does not yet support.

Current evidence position: AILAT has not published an instrument-specific validation study or population norms.

Evidence ledger

What is established—and what is not

Published basis
The four-dimension construct draws on published AI-literacy competency and measurement research.
Implemented method
The 20-question flow, adaptive selection logic, scoring rules, and reporting model are implemented and documented.
Not yet published
AILAT-specific reliability, model-fit, test–retest, criterion-validity, fairness, and human–model scoring studies.
Not available
Population norms or an external certification standard against which an AILAT score can be interpreted.
Length
20 questions
Delivery
3 assessment phases
Construct
4 dimensions
Interpretation
5 descriptive levels
The construct

AI literacy is represented as four related capabilities.

AILAT samples conceptual understanding, practical use, critical evaluation and creation, and ethics. These dimensions are informed by published AI-literacy frameworks, then operationalized through AILAT’s own item bank and scoring rules.

01

Conceptual Knowledge

AI fundamentals, system behaviour, limits, and terminology.

02

Use & Apply Knowledge

Prompting, workflow integration, and practical use of AI tools.

03

Evaluate & Create

Judging output quality, identifying limitations, and improving AI-assisted work.

04

Ethics Knowledge

Bias, privacy, accountability, transparency, and responsible adoption.

Assessment delivery

Breadth first, adaptive selection second, context third.

Only the middle phase uses adaptive item selection. Describing every question as adapting to every answer would overstate the design.

015 questions

Calibration

A fixed, mixed-difficulty opening samples the four dimensions and initializes the proficiency estimate used during delivery.

0210 questions

Adaptive

Eight multiple-choice items are selected using 2PL item-information concepts; two open-ended prompts sample reasoning that fixed choices can miss.

035 questions

Scenario

Four multiple-choice items and one open-ended capstone place the same literacy dimensions in the participant’s selected industry context.

Adaptive selection

IRT concepts guide which item is useful next.

During the adaptive phase, AILAT uses a two-parameter logistic model and Fisher information to choose informative multiple-choice items around the running proficiency estimate.

The present item parameters are part of the authored assessment design. They should not be described as empirically calibrated until a calibration study has been completed and published.

Reported scoring

The final 0–100 result is a weighted performance score—not an IRT theta score.

Multiple-choice correctness and fractional credit from rubric-scored open-ended responses produce dimension percentages. The overall score then weights those dimensions.

Conceptual
30%
Use & Apply
30%
Evaluate & Create
25%
Ethics
15%
Interpretation

Use AILAT to direct learning—not to make high-stakes decisions about people.

The current evidence supports cautious, formative use. The named levels describe score bands inside AILAT; they are not population percentiles or external certifications.

Appropriate uses

  • Individual learning reflection
  • Organization learning-needs baseline
  • Training-priority discussion
  • Exploratory pre/post snapshots with context

Inappropriate uses

  • Hiring, promotion, or performance ranking
  • Formal professional certification
  • Claims of legal compliance
  • Claims that a score change proves training impact
Evidence roadmap

What a defensible validation programme still needs to establish.

Publishing these analyses would allow future claims to become more specific. Until then, AILAT’s public language should remain explicit about the distinction between theoretical foundation and instrument-specific evidence.

  1. 01

    Item calibration

    Estimate difficulty and discrimination from a sufficiently large, representative response sample instead of treating authored parameters as empirical estimates.

  2. 02

    Reliability and precision

    Report conditional measurement precision, internal consistency where appropriate, and stability across repeated administrations.

  3. 03

    Construct and criterion evidence

    Test whether the four-dimension structure fits observed data and how scores relate to relevant external measures or performance.

  4. 04

    Fairness analysis

    Evaluate differential item functioning and subgroup performance before making broad comparability claims.

  5. 05

    Open-ended scoring agreement

    Compare model-assigned partial credit with independent human ratings and monitor agreement when scoring models or rubrics change.

Foundational literature

Published work that informs the construct—not evidence borrowed by association.

Citing an established framework does not validate a new assessment automatically. These sources help explain the design context; AILAT must generate its own empirical evidence.

  1. Long & Magerko (2020)

    What is AI Literacy? Competencies and Design Considerations

    A competency-oriented definition and design framework for AI literacy.

    Open source
  2. Carolus et al. (2023)

    MAILS — Meta AI Literacy Scale

    A self-report AI-literacy scale grounded in competency models. AILAT is not MAILS and does not inherit its validation evidence.

    Open source
  3. Ding, Kim & Allday (2024)

    Development of an AI literacy assessment for non-technical individuals

    Performance-oriented assessment research for a non-technical population. AILAT is a separate instrument.

    Open source

This page is a product evidence statement, not a peer-reviewed publication. We will update the status when AILAT-specific evidence is available.

Read the implementation white paper