AI Adoption GuideEducationAssess
Adaptive Assessment Generation
An LLM generates unique question variants from an item bank, calibrated to each student's current mastery level and prior response patterns.
Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance
By Don, DoneThat’s AI coach · updated
A scored variant must cite its source item and mastery band
Adaptive assessment generation is usable when the output is a question variant that names the banked stem it came from and the mastery band it was written for. If the item bank has no eligible stem for that student at that band, the output stays empty. Faculty still decide what counts as a scored attempt. A model does not invent a standard, a cut score, or a mastery level the course has not already established.
The job is not unlimited new items. Keep the measured construct stable while changing surface details so a student does not sit the same stem twice, and so difficulty stays inside the band the course already uses. Prior response patterns inform the target band only when those patterns exist and have been interpreted under the course's own rules. Do not fill missing evidence with a guessed level.
This is not the same as writing lesson copy. automated content generation can produce explanations or practice that never enter the gradebook. A generated assessment variant that will be scored is an item. Treat it with the same custody you already apply to banked items: provenance, review, and an explicit decision that it may count.
Pull an eligible stem, then generate at the cited band
Start from the bank, not from a blank prompt. Identify the outcome the attempt is meant to measure. Select stems tagged to that outcome and eligible for the student's current band. Eligibility is a course rule: exam-reserved, practice-only, retired, or dependent on a figure or calculator policy the generator cannot recreate faithfully. If no stem meets those constraints, stop. Empty is the correct result.
The mastery band should come from evidence the course already trusts. If the program infers bands from a sequence of scored work, use that inference rather than a chat guess. competency mastery inference is the upstream step: it tells you which band to cite. Adaptive generation consumes that citation. It does not replace it. If the student has not sat the quizzes or checkpoints that feed the inference, do not invent a band from a missing quiz. Deliver a banked item at the default entry band the syllabus already names, or wait until evidence exists.
Once you have an eligible stem and a cited band, generate a variant that keeps the same construct, cognitive demand, and scoring intent. Change names, numbers, contexts, or distractors only as far as the original item's blueprint allows. The variant record should include the source item ID, the mastery band used, the outcome tag, and reviewer notes on what was held constant and what was allowed to change.
A p-value interpretation stem
A first-year statistics instructor has a banked multiple-choice stem that asks students to interpret a p-value for a two-sided test of a mean. The scenario, test statistic, and p-value are given. The construct is interpretation, not computation. A student has a documented approaching band on that outcome from homework and a checkpoint quiz. The generator is asked for one variant at that band, citing stem STAT-211-014.
A usable variant keeps the same decision: whether the p-value is evidence against the null under the stated alpha, and whether failing to reject is confused with proving the null. It may change context (dining wait times instead of bus arrivals), sample size, and the numeric p-value, if those numbers stay in the same interpretive region. It must not become a calculation item that asks the student to compute the test statistic. That is a construct shift, even if the topic still looks like hypothesis testing.
If the bank has no interpretation stem tagged for the approaching band, because remaining items are calculation stems or exam-reserved, the generator returns nothing. The instructor does not ask the model to write something similar anyway. That is how unofficial standards appear.
Review is what makes the attempt countable
Generation is a draft. Scoring is a faculty or assessment-office decision. Before a variant is attached to a quiz, exam, or mastery checkpoint, a human who owns the outcome reviews it against the source stem.
A variant that changes the construct fails review. If the original item measured interpretation and the variant now requires a procedure the student was not meant to perform at that band, the item measures something else. If a selected-response item adds units, formula recall, or a second outcome, it is no longer a variant of the source. Reject it or send it back with the construct named in the comment.
Scoring a generated item that was never reviewed also fails the process. It is tempting to drop a fresh variant into Canvas, Blackboard, Moodle, or an Anthology-hosted course because the wording looks clean. An unreviewed item is not a scored attempt, even if the LMS will collect a response. Keep generated drafts in a holding state the gradebook cannot see until review is recorded. If you reuse the variant in another section, the review travels with the item, or you review it again.
Inventing mastery from a missing quiz is the third failure. If prior response patterns are incomplete, the generator must not infer a band from silence, participation, or a quiz that was excused or never opened. Missing evidence is missing. Pair generation with the inference process you already trust, or do not generate an adaptive variant for that student. A static banked item at a published entry level is more defensible than a personalized item aimed at a band nobody measured.
For constructed-response variants, review includes the scoring guide. If the original stem uses a rubric, the variant needs a rubric that still matches the construct. rubric-based essay scoring is a separate step from writing the prompt. Do not assume a generated essay prompt inherits the original rubric if the task demands shifted.
Keep practice separate from the gradebook. formative feedback generation can comment on a student's working without creating a new scored item. Mixing those jobs is how unofficial items slip into the official record.
Course platforms hold delivery; they do not own the construct
Learning platforms and integrity tools are the delivery and custody layer, not the psychometric layer. Canvas, Blackboard, Moodle, and Anthology systems can host the quiz, randomize presentation, and store attempts. Turnitin can sit on written responses where the course already uses it for originality or similarity review. None of those products defines the mastery band, bank eligibility, or whether a generated variant is the same construct as the source stem. A successful import is not item validation.
Keep provenance out of the question text the student sees. Source item ID and band label can coach the answer or reveal adaptive logic. Staff-facing records should still show source item, band, reviewer, and date.
Adaptive item generation is not the same as sequencing the next module. adaptive learning path engine work decides what to study next. This page is only about the question asked at a given checkpoint. Collapsing both into one prompt hides which outputs were reviewed for scoring.
Generated items inherit existing constraints: item analysis, blueprints, and approved accommodations. A variant that cannot be delivered with the same accommodations as the source stem is not eligible, even if the wording is otherwise sound.
When the bank has nothing eligible, stop
An empty result is a process success when the constraints are working. The bank cannot support a unique variant at the cited band without inventing a stem, stretching a construct, or borrowing an exam-reserved item. The operational response is human: tag more stems, revise eligibility, offer a banked item that is allowed, or postpone the adaptive attempt until evidence and inventory exist.
Do not ask the model to write a new question at this level and call it a variant. That is new-item authoring and needs the institution's original-item workflow. Do not backfill a mastery band from attendance, time on task, or a quiz the student never took. Do not put an unreviewed draft in a slot that reports to the gradebook.
Faculty remain accountable for the scored attempt. The useful artifact is narrow: a variant that cites the source item and the mastery band used, or nothing. Keep that artifact small enough to review, and refuse to score it until review is done.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first