AI Adoption GuideEducationAdmit
Holistic Application Scoring
An LLM scores applications across rubric dimensions such as academics, essays, and activities, then flags outliers for human review.
Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance
By Don, DoneThat’s AI coach · updated
Treat the output as a draft, not an admit
A language model can score an application only as a draft against the committee's rubric. Each dimension gets a score, a cite to a passage in the file, and a cite to the rubric line that justified it. Empty stays empty when the file has no evidence for that dimension. The packet is not a ranking, not a predicted first-year GPA, and not an admit. The reader still owns the recommendation.
Quality here means a reader can open the packet, jump to the cited page, and confirm or reject the cell in minutes. Completeness that invents an activity, averages dimensions the rubric keeps separate, or treats a total as a decision is a process failure, even if the numbers look tidy.
Use the draft as input to reviewer decision assistance. Do not skip the human write-up because a total already exists.
Score each dimension from the file
Work the rubric in the order the committee published it. For each dimension, search only the parts of the file that dimension is allowed to use.
Academics come from the transcript, the school profile, and any official test or course list the office already treats as academic evidence. Essays come from the essay text and any prompt-specific supplements. Activities come from the activity list, resume, or other activity section the application actually includes. Recommendations come from the rec letters. Do not borrow a sentence from the essay to populate activities, and do not use a glowing rec to fill a missing grade trend.
For every non-empty cell, record three things: the draft score or band the rubric uses, a short quote or locator in the file (page, section, or field name), and the rubric line (criterion ID or exact wording). If the rubric uses bands such as "limited / typical / distinctive," use those labels. Do not invent a 1-10 scale the committee never adopted.
If the search returns nothing on-point, leave the cell blank and write "no evidence in file" in the cite field. Do not infer a club, a job, or a sport from tone. Filling a blank activity is the most common silent error: the model treats a first-person essay about teamwork as proof of a captaincy that never appears in the activity list.
Keep dimensions separate until a human says otherwise. A high academic cell and a thin activity cell are two facts. Averaging them into a single "holistic" number erases the rubric's structure and hides which evidence was missing. If the committee later weights dimensions, that weighting is a policy choice in the reader workflow, not a step the model should take on its own.
Essay scores answer the rubric's essay criteria (argument, reflection, fit to prompt). They are not a substitute for essay authenticity screening. Authenticity is a different question. Do not dock the essay-quality cell because the prose "sounds generated" unless the rubric explicitly includes that criterion. Likewise, this scoring pass is not academic success prediction. A rubric score describes the file against published criteria. A success model guesses later outcomes. Mixing the two breaks the cite trail the reader needs.
One file, scored without filling gaps
Consider a transfer applicant whose packet includes an official transcript, a personal statement, one counselor rec, and an activities section that lists only a part-time retail job. The rubric has four dimensions: academics, essays, activities, and recommendations.
Academics: the transcript shows a completed quantitative sequence with a clear upward trend after a weak first term. Cite the term rows and the rubric line for "academic trajectory," then assign the band that line describes. Do not upgrade the band because the essay sounds ambitious.
Essays: the statement answers the prompt with a specific account of why the student is changing institutions. Cite the paragraphs that match the rubric's essay criteria. Do not pull the retail job into the essay score, and do not treat the essay as an activities source.
Activities: the list contains the retail job and nothing else. Score that job against the activities line if the rubric has a place for work. Leave every other activity sub-cell empty. Do not add "leadership" because the essay used the word "team," and do not create a campus club that is not in the file.
Recommendations: the counselor letter speaks to classroom habits. Cite those sentences against the rec rubric line. If the letter is silent on character or extracurriculars, those rec sub-cells stay empty rather than inheriting praise from the academic paragraph.
The draft that comes out of this pass is uneven on purpose. That unevenness is the point. A filled grid that invented a second activity would look more complete and would be wrong.
Flag outliers, then hand the file to a reader
After every dimension has a score or an explicit empty, run a second pass that does not rescore. Flag files that are unusual relative to the rubric and to the rest of the pool the office is reading that round.
Useful flags are specific: an academic band far above or below the rest of the file; an essay that does not address the prompt; a rec that contradicts the transcript; a dimension left empty where most complete files in this round have evidence; a cite the model cannot locate on the stated page. Each flag should name the dimension, the cite, and why it is unusual. "Review this one" without a reason is not a flag.
Outlier status is not a vote. A flag means a human opens the file. It does not mean auto-waitlist, auto-merit, or auto-deny. Do not convert a low total, or a missing activity cell, into a decision rule. The total, if you display one at all, is a convenience sum of scored cells. It is not an admit. Treating it as one launders a policy the committee never voted on.
When totals start deciding who gets a second read, you already have a de facto policy. Keep those human decisions in the audit trail, and use bias audit on admit decisions on the admits and denies the committee actually issued, not on the model's convenience sum.
Where the packet already lives
Most offices do not need a new system of record. Application PDFs, counselor recs, transcripts, and activity lists already sit in the campus stack the office runs: Slate, Technolutions, Salesforce, or Ellucian. Export or assemble the same packet a human reader would see. Score that packet. Write the draft scores and cites back as a reader aid, not as a new source of truth.
Those platforms differ in workflow and data model. Treat them as a class: they hold the file. Do not assume a shared scoring API, a shared rubric object, or a shared "AI review" feature. If a field is missing in the export, it is missing in the file. Do not fill it from a neighboring CRM object that the reader would not have seen.
The reader still writes the recommendation, still applies institutional priorities, and still signs the decision. The model's job ends when each rubric line has a cite, an empty, or a flag.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first