Skip to main content
DoneThat

AI Adoption GuideEducationAssess

Rubric-based Essay Scoring

An LLM grades open-ended submissions against decomposed rubric criteria and flags low-confidence grades for human review.

Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance

By Don, DoneThat’s AI coach · updated

Faculty posts the grade; the model only drafts the worksheet

A usable scoring run produces a draft score per rubric row, each tied to spans in the submission. If a criterion has no evidence in the text, that row stays empty. Nothing posts to the student or the grade book until a faculty member or assessment lead accepts, edits, or rejects the draft.

The risk in this workflow is treating a model output as a posted grade. LMS grade books in Canvas, Blackboard, and Moodle, plus integrity and originality layers such as Turnitin and Anthology, already own the official record. The scoring assistant should write into a review queue, not into those records.

If a draft score posts without faculty review, three things happen at once. The student sees a number that may not match the rubric as the instructor uses it. Empty or invented rows can silently inflate or deflate the total. The instructor loses the chance to catch a low-confidence call. Keep posting rights with people.

Break the published rubric into independently scorable rows

Do not send the model a paragraph-long holistic rubric and ask for a single letter grade. Split the published instrument into rows the model can treat as separate questions.

Each row needs four things before a run: the criterion name as students see it, the performance-level descriptors or points for that criterion only, what counts as evidence in the submission, and an explicit instruction that missing evidence means an empty score, not a zero guessed from other rows.

Holistic language such as "sophisticated, well-supported argument" mixes thesis quality, evidence, and style. If those are separate rows on the actual rubric, keep them separate in the prompt and in the output schema. Inventing a criterion that is not on the published rubric is a failure mode even when the invented comment is insightful. The draft must map onto the instrument students were given.

Weighting belongs in the faculty step, not in silent model arithmetic. If Organization is worth 10 points and Use of Evidence is worth 30, record those weights in the worksheet so a reviewer can see them. Do not let the model average filled rows and skip empty ones as if the missing criterion did not exist.

When the same assignment also feeds competency mastery inference, keep scoring and mastery as two artifacts. A rubric row is a local judgment on this submission. A competency claim is a pattern across tasks. Do not collapse them into one number in the scoring run.

Score each criterion from cited spans, or leave the row blank

For every criterion, the model should return a proposed level or points, the exact spans it used, and a short justification that only uses those spans. Cites should be quotations or tight locators a reviewer can jump to, such as a paragraph or sentence range. If the model cannot point to a span, the score field stays empty.

A first-year writing prompt asks for a 1,200-word argument on whether campus libraries should extend 24-hour access during exams. The published rubric has four rows: Thesis and claim, Use of evidence, Organization, and Source attribution. A submission opens with a clear claim in paragraph 1, uses two news articles in paragraphs 3 and 4, and has a recognizable intro-body-conclusion shape, but it never names those sources in-text or in a list.

A defensible draft fills Thesis, Evidence, and Organization with cites to those paragraphs, and leaves Source attribution empty. It must not invent a citation-quality score from the fact that articles were mentioned in passing, and it must not fill the empty row with a guessed midpoint so the total looks complete. Averaging the three filled rows into a fourth is the same error with nicer arithmetic.

Reviewers should reject a cite that is off-criterion. A sentence that shows organization is not evidence for source attribution. If the model reaches for a neighboring row to avoid leaving a blank, treat that as a defect in the draft, not as helpfulness.

This scoring pass is not formative feedback generation. Scoring names a level against a rubric row. Formative comments tell the student what to change next. Running both in one blob makes it harder to see whether the grade is warranted. Keep the score worksheet first. Generate comments only from accepted row scores if the course wants both.

Low-confidence drafts stay in the review queue

Confidence is a routing signal, not a second grade. Flag a row, or the whole submission, when cites are thin, when the submission is far off-genre, when OCR is noisy, or when two descriptors could both fit the same span. Those items stay in the queue until a person acts.

Do not auto-resolve low confidence by picking the middle band. A mid-band default looks decisive and is often wrong. The honest output is unscored, needs review, plus the spans the model was unsure about.

Batch the queue by why items were flagged: empty-row cases, weak-cite cases, off-rubric language. Faculty still posts. A teaching assistant may triage, but posting to the student-facing grade remains a named role with the authority to change a row, empty a row, or send the paper back.

Never post an unreviewed score because the queue is long. Throughput pressure is how invented criteria and empty rows filled with guesses show up in live courses. If the run cannot finish a row, the student-facing grade for that criterion does not exist yet.

Integrity questions are a different queue. Essay authenticity screening and ai-content detection may run on the same file, but their outputs are not rubric levels. Do not let a suspicion score substitute for Use of evidence or Source attribution. A paper can be authentic and still miss a criterion, or fail a detector heuristic and still meet the published rows. Keep those flags adjacent in the review UI if the campus workflow requires it, and keep them out of the arithmetic that produces the draft total.

Grade books stay the system of record

Export or copy accepted scores into the course grade book after faculty posting, using the same row names students already see. Canvas, Blackboard, and Moodle remain where totals, late policies, and excused statuses live. Originality and integrity products such as Turnitin and Anthology remain where similarity or authenticity reports live. The scoring assistant should not impersonate those systems.

Reconcile names. If the LMS rubric says Support and the paper rubric says Use of evidence, pick one label and keep it through draft, review, and post. Mismatched labels cause reviewers to "fix" a row that was never empty.

After posting, the audit trail should show the draft, the cites, who changed what, and that empty rows were not silently zeroed. That record is what you need when a student asks why a criterion has no score, or when an assessment committee asks whether the tool invented a standard.

Do not treat last term's posted totals as training labels unless faculty marked those papers as anchors. Posted grades include instructor judgment, curve decisions, and sometimes mercy. Feeding them back as ground truth teaches the model to mimic the grade book, including its exceptions.

A colleague who did not run the model should be able to open the worksheet, jump to each cite, see empty rows that are actually empty, see low-confidence items still waiting, and confirm that the student-facing grade appeared only after a person posted it.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first