AI Adoption GuideEducationAdmit
Academic Success Prediction
ML predicts first-year GPA and retention probability for each applicant from academic and contextual signals to support admit decisions.
Education processRecruitAdmitEnrollTeachAssessCredentialGraduateAdvance
By Don, DoneThat’s AI coach · updated
A cited prior, not an admit recommendation
Academic success prediction at admit is a cited prior for first-year GPA band and first-year retention, trained on your own enrolled students. It is not an admit. It is not a deny. It does not invent a GPA when the cell is too thin to cite.
The committee still owns the decision. Place the prior on the file the way institutional research would place a sourced range: labeled, dated, and optional. If the cell is thin, both the GPA band and the retention prior stay empty. A blank that tells the truth is a quality signal. A filled number that was never observed at your institution is not.
Do not treat the predicted band as merit. First-year GPA is an outcome you recorded after people enrolled, in your courses, under your grading, with your supports. It is not belonging, grit, or "fit." Using it as a proxy for who deserves a seat turns a quality prior into a ranking of people.
Keep it next to holistic application scoring, not above it. Transcript pattern, recommendations, and context still carry the file. The prior is one cited statement about likely first-year grades and return, or it is absent.
Train on enrolled first-year outcomes only
Train only on first-party enrolled history. The unit of observation is a student who matriculated, their first-year GPA mapped into the bands your registrar already uses, and whether they met your published first-year retention definition (return for year two, or completion of the first year in good standing, if that is what IR already reports). Do not train on denied applicants. Do not train on deposits who never enrolled. Those groups never produced the outcomes you are citing.
Cites belong on the training set the same way they belong on the file. Record cohort years, census date, GPA banding rules, retention definition, and exclusions (pre-census withdrawals, visiting students, dual-enrollment-only records). When a reader sees a band and a retention prior, they should be able to open the cite and see which enrolled classes generated it.
Use academic and contextual signals you already collect for review: transcript pattern, curriculum rigor as you code it, prior college credit, sending-school or sending-college identifiers, and context flags readers already see (first-generation, gap year, English-language pathway). Do not add a belonging score. Do not add a "likely to thrive" label. Those are not first-year GPA and retention.
Score the prior against later enrolled cohorts, not against last year's admit list. Error belongs on first-year GPA band and retention among students who actually enrolled, sliced by sending school, curriculum cluster, and the groups your bias audit on admit decisions already requires. If you skip that audit, you will not know whether a low prior is a student signal or a historical enrollment pattern you should not repeat.
Refresh on an IR calendar, not mid-cycle. Changing bands while files are in committee rewrites the cite under people already under review.
Put the band and probability next to the file
Show the prior beside the application. The same view should hold the transcript, recommendations, and the rest of holistic review, plus a short cited block: first-year GPA band or blank, retention prior or blank, training cite, and a one-line thin-cell note when you cannot cite.
Reviewer decision assistance is the right pattern: the model proposes a cited range; the reader may use it, ignore it, or flag it. It must not auto-sort the queue toward "predicted to struggle" or move a file to deny.
Here is the one picture to keep in mind. A first-year applicant from a small sending school that has enrolled only a handful of your students in recent cohorts. The transcript is complete. The counselor letter is specific. The curriculum does not match the large feeders that dominate enrolled history. The honest output is an empty GPA band and an empty retention prior, with a cite that the sending-school cell is too thin. The dishonest output is a made-up GPA because the dashboard hates blanks. The committee still reads the file. The empty cell is the correct quality signal.
Do not stamp "below our bar" on emptiness. Do not encode "insufficient data equals risk." That rule is a deny dressed up as data quality.
Leave thin cells empty
Thin is a published IR rule, not a reader's hunch. Define it before go-live: minimum enrolled outcomes in the cell you will cite (sending school, curriculum cluster, or another slice you actually use), how recent those outcomes must be, and what happens when the applicant sits in more than one sparse slice. When any gate fails, both the GPA band and the retention prior stay empty.
Do not interpolate from the next-largest feeder. Do not show a campus-wide average as if it were this applicant's prior. A campus-wide average is not cited history for this file.
Empty is usable. Readers already work with incomplete files. What they cannot defend is a precise-looking GPA that no enrolled student like this one ever produced at your institution.
If later cohorts fill the cell, the prior can appear on future files. That is an IR refresh. It is not a reason to rewrite people already in committee.
SIS and CRM systems hold history, not the decision
Application files, inquiry records, and term-by-term outcomes usually already live in enrollment CRM and SIS products. That class includes Slate, Technolutions, EAB, Ellucian, and Salesforce. Treat them as systems of record for the file and for first-party enrolled history. Do not assume any of them ships a first-year GPA and retention prior you can cite. If you attach a model, you still owe the cite, the thin-cell rule, and the committee vote.
Keep the prior out of automated student messages. A retention probability at admit is not an early engagement alert. Engagement alerts belong after someone enrolls and misses a signal you defined. Mixing the two trains staff to treat a pre-enroll forecast as a student who is already failing.
What goes wrong when the prior becomes the vote
Three failure modes show up as soon as the prior is convenient.
Denying on a thin prior. The cell was empty or barely populated. Someone treated emptiness, or a wide band, as evidence the student would not succeed. You cannot defend that deny with first-year outcomes you never had.
Treating GPA prediction as merit. The band is a forecast of grades in your first-year courses, not a ranking of worth. Sorting admits by predicted GPA recreates historical enrollment as if it were excellence. Applicants from sending schools you rarely enroll will look worse because you lack their outcomes, not because they lack ability.
Skipping a bias audit on admit decisions. If you never check whether the prior, or the way readers use it, shifts admit rates across groups you are obligated to watch, you will not see the proxy-for-belonging problem until someone else asks. Audit decisions, not only model error. A well-labeled GPA band can still become a biased vote if reviewers use it as a shortcut.
The quality bar stays narrow: cite first-year GPA band and retention from enrolled history, leave thin cells empty, and let the committee decide.
Is this worth automating for you?
Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other. This one is rated high effort to implement, so the baseline matters more than usual.
DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.
Measure the baseline first