AI Adoption GuideEducationAdmit
Essay Authenticity Screening
A classifier screens admissions essays for low-authenticity or likely AI-generated content before human review.
Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance
By Don, DoneThat’s AI coach · updated
The screen returns a citeable flag, not a verdict
Authenticity screening on an admissions essay should return a flag, a cited passage, and the screen rule that fired. It should not return a deny. The reader still owns the file.
That split is the whole design. Essay screens sit in a messy class of tools: plagiarism and AI-writing detectors such as Turnitin, application systems such as Slate and Technolutions, and learning platforms such as Canvas. Those products surface scores, similarity reports, or originality signals. None of them ranks truth, and none of them should close a file.
A usable result has three parts when something fires. The flag names the concern, such as low-authenticity or likely AI-generated. The cite is the span that triggered it. The rule is a short, arguable name for why the span fired: too-uniform syntax, a sudden register shift against the rest of the statement, or a prompt-template opening. If the essay is too short to support a screen, or the file cannot be read, the result stays empty. Empty is not a clean bill of health. Empty means the machine had nothing it could defensibly say.
Do not invent a detection rate for this workflow. There is no rate that travels across prompt types, languages, assistive-writing tools, and vendor versions. A percentage on a dashboard is not a finding about this applicant.
How to screen the essay before a reader opens the file
Run the screen after the application has a readable personal statement, and before the assigned reader starts a scored review.
-
Confirm you are screening the essay the applicant submitted, not a counselor draft, a portal excerpt, or a concatenated packet. If you cannot identify the source file, stop. Leave the result empty.
-
Check length and readability before any classifier. Your office sets the minimum length. Do not import a vendor default as policy. Skip scanned PDFs that will not extract, password-locked files, and encodings that produce garbage text. You may log unreadable. You may not log authentic as a substitute.
-
Run the classifier on the extracted essay text only. Do not mix in the activities list, the recommendation, or the transcript. Those documents have different authors and different screen rules.
-
Persist three fields if anything fires: flag type, cited span, and rule id. The span should be offsets or a quoted excerpt long enough to relocate. If nothing fires, persist empty, not a zero, not a green badge.
-
Route flagged files to the assigned reader with the cite visible. Unflagged files follow the ordinary read. Do not create a parallel integrity track that never returns to the reader of record.
-
The reader reviews the cite against the whole essay and the rest of the application before any integrity action. Integrity action means a request for more writing, a conversation with the applicant, or a documented concern. It does not mean an auto-deny from the screen.
If a CRM already stores originality scores from Turnitin, Slate, Technolutions, or Canvas, apply the same three-field test. No span and no rule means you have a number, not a screen result.
Operationally this is AI-content detection constrained to one admit document. Do not copy a classroom detector policy into admissions just because the vendor name matches.
Citing the span and the screen rule
A flag without a cite is not reviewable. A reader cannot evaluate "likely AI-generated, high confidence" against an essay. They can evaluate a located span plus a rule they can argue with: paragraph 3, sentences 2 through 5, sudden drop in error rate and jump in nominalizations relative to the rest of the statement.
Cite the smallest span that still lets a second person find the trigger. Prefer a quoted excerpt plus a location, such as a paragraph index or character offsets. Do not dump the whole essay as the cite. Do not hide the span behind a vendor heatmap the reader cannot paste into a file note.
Name the rule in language a committee can challenge. "AI probability" is not a rule. "Matches a known prompt-template opening" is a rule. "File failed extraction" is a condition that must produce empty, not a flag.
When the reader disagrees with the cite, the file note records the disagreement. The screen does not get a veto. Pattern review across a cycle can tally flags, empties, and overrides. That tally is operations. It is still not a detection rate you can publish as a property of the product.
One illustration, not a measured case. A reader opens a first-generation applicant's statement. The screen flagged the second paragraph as likely AI-generated: long sentences, tidy transitions, near-zero error rate. The cited span describes a chemistry lab. The rest of the essay has shorter sentences, a few subject-verb slips, and a teacher name that matches the counselor rec. The reader keeps the file in ordinary review and does not open an integrity case. The same rule on a generic span with no corroborating detail can still justify a short writing-sample request. The screen did the same job both times. Only the reader can tell the files apart.
Keep rubric-based essay scoring as a separate pass. Authenticity flags are not rubric scores. Mixing them trains readers to treat "sounds like a model" as "weak writing," which punishes revision, tutoring, and non-native fluency.
Failure modes that should never become policy
Denying from a score. A similarity or AI score is not an admissions decision. Auto-deny, auto-waitlist, or a silent completeness hold that never returns to a reader is the failure. If the workflow cannot continue without a human owner, the screen is not ready.
Flagging a non-native writer as AI. Classifiers often treat lower burstiness, more formal transitions, and fewer local idioms as machine-like. Those features also appear in careful second-language writing, especially after tutoring. If the flagged span is the only fluent paragraph and the rest of the file fits a dictionary-and-writing-center writer, discount the flag. Without a language-background protocol, proficiency becomes an integrity allegation.
Treating a vendor score as proof. Reports in the Turnitin class, portal badges in Slate or Technolutions, and assignment-originality views in Canvas are evidence that a screen ran. They are not proof of authorship. Proof, when you need it, is a process: the cite, the applicant's other writing, a follow-up prompt under known conditions, and a documented reader judgment. A number with a logo is none of those.
Two quieter failures sit beside those three. Filling empty maps too-short or unreadable to clear and hides extraction bugs. Double-counting treats a detector score, a reader hunch, and a committee anecdote as three findings when they are often one finding restated.
If you use reviewer decision assistance, keep authenticity flags out of the assist prompt until the reader has reviewed the cite. Otherwise the assist restates a detector score as concerns about voice, which is how a screen becomes a verdict without anyone saying so.
Where authenticity screening sits among other admit reads
Authenticity screening is a pre-read quality check on one document. It is not holistic application scoring. Holistic scoring asks whether the file, in full, meets the class you are trying to admit. An authenticity flag can inform that read only after a person has accepted or rejected the cite. Folding the flag into a composite score smuggles an unreviewed screen into selection.
Keep the objects separate. Essay authenticity is flag, cite, rule, or empty. Essay quality is a rubric pass, never the authenticity classifier.
Readers should treat the authenticity block like a counselor rec that contradicts the activities list: resolve it in the narrative. Integrity leads should watch override rates and empty rates, not a leaderboard of AI essays caught.
When a flagged span survives reader review, the next step is human: a short timed writing sample on a new prompt, or a conversation. Do not deny because a detector in the Turnitin, Slate, Technolutions, or Canvas class produced a high score. The score did not read the file. You did.
The durable rule is small enough for the workflow card: screen the essay, cite the span, name the rule, leave empty when the text is too short or unreadable, and do not act on integrity until the reader of record has looked.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first