Skip to main content
DoneThat

AI Adoption GuideConsultingRecommend

Initiative Prioritization Matrix Generator

LLM scores initiatives on impact, effort, and urgency using client-agreed criteria and outputs a ranked priority stack.

Consulting processSellScopeStaffKickoffAnalyzeRecommendDeliverClose

By Don, DoneThat’s AI coach · updated

A model that invents weights is ranking its own taste

The useful output is a short ranked stack scored on criteria and weights the sponsor signed, with readiness sitting in its own column. A board of forty sticky notes is not a decision. Neither is a composite score a model invented because nobody locked the axes.

Quality fails when scoring starts before the client has agreed what impact means. Filling cells fast is a side effect, and a cheap one. A rank that encodes the model's taste, or the partner's, will not survive a CFO who asks who chose the weights.

The list usually already lives somewhere in the Excel, Miro, Mural, and Smartsheet class: a workshop board, a tracker, a two-by-two for steering. Those tools store the rows. They do not decide the criteria. Do not ask a model to prioritize this list until the sponsor has named the axes and put a name on the weights.

Lock the axes with the sponsor before any score exists

Sit with the sponsor, not the full steering committee and not the analyst, and write three things down before a single cell is filled.

  1. Criteria. What impact is in this engagement. On-time departure, cash inside twelve months, and regulatory exposure are different questions. Pick the ones this recommendation is for. Define effort in a unit the delivery team can own (named-person weeks, not "medium"). Define urgency as time-to-harm if you wait, not as "the board is impatient."
  2. Weights. Numeric, in a dated note, with the sponsor's initials. If the model proposes a 50/30/20 split because that is a common teaching matrix, you are ranking a textbook, not this client.
  3. What is not a criterion. Readiness, politics, "the CEO will hate this," "we already promised the board." Those are real. They are not impact.

A model that invents weights is ranking its own taste. Treat an unsigned weight set as a defect, the same way you would treat an unsigned exhibit.

Impact, effort, and urgency can sit on one page. Readiness cannot be blended into impact. A high-impact, not-ready item is still high-impact. Sequencing and staffing change. The score does not. Pass political feasibility and change capacity to the client organizational readiness classifier as a separate column the steering committee can see.

This matrix is not a substitute for comparing strategic options. If you have not yet chosen a path, use the strategic option comparison synthesizer on client-agreed option criteria. Rank initiatives inside a chosen path. Do not use a forty-row stack to pretend the option decision was analytical.

Score from evidence, and keep readiness off impact

Once the axes are locked, score each initiative against evidence you already have: interview quotes, cost extracts, service files, inspection calendars. "High impact" with no source is an adjective. Force a one-line evidence note per cell, or leave the cell empty.

Effort is not the consulting team's guess. The client's delivery leads set it, in the unit the sponsor defined. If the delivery lead is not in the room, you do not have an effort score. You have a placeholder the client will disprove in week one.

Show how close adjacent items are. A thin gap is a near-tie, not a rank. Presenting one through thirty-eight as a strict order turns noise into a decision.

Ranking is not sequencing. Dependencies and capacity decide what can start. A ranked list that implies row one starts Monday will be disproved as soon as two items share a scarce team. Keep a sequence view next to the stack. Once work is live, delivery risk is a different job: the workstream delivery risk predictor watches velocity and slippage. It does not belong in the recommend-stage composite.

Cap the "do now" stack at what the organization can actually start. If they can run five concurrent workstreams, the output is five, plus a wait list, plus a not-now list. Ranking forty so every owner still has a "priority" is how nothing is a priority.

Illustrative example: Harborline, thirty-eight rows

This is a worked example with made-up names, written to show the cuts, not a case study with results.

Elena is the partner on an operating-model recommendation at Harborline, a regional ports operator. A Mural workshop produced thirty-eight initiatives. The analyst copied them into Excel, parked a tracker in Smartsheet, and built a Miro two-by-two for the steering dry-run.

The analyst asked the model to score impact, effort, and urgency. The model invented weights (impact heavy, then effort, then urgency) and wrote a ranked stack. Shared-services finance sat near the bottom. Plant night-count discipline sat near the top on operational impact, then fell when the model blended "the CFO is not ready" and "Westbrook will resist" into the impact cell.

Elena does not want night-count discipline in the recommendation. It looks like a process fix, not a transformation, and it would put a plant manager on stage instead of the shared-services story she has been socializing. She asks the analyst to re-run with more weight on structural change. That is a hidden weight. The matrix is being used to bury an idea the partner does not like.

The COO, as sponsor, had never signed the weights. Steering would have seen a thirty-eight-row stack in which every owner could still claim a priority.

Elena stops the re-run. She sits with the COO and locks three impact criteria: contribution to on-time vessel departure, cash released inside twelve months, and exposure on the next inspection cycle. Weights go in a dated note with the COO's initials. Effort is scored by Harborline's PMO lead in named-person weeks. Urgency is time-to-harm if delayed. Readiness is a separate column: CFO capacity, Westbrook change load, union notice periods. None of that is allowed to discount impact.

Night-count discipline stays high on impact against on-time departure, with the interview evidence that skipped counts delay outbound. It is marked not-ready on the readiness column. Shared-services finance is a strategic option question, not a row that should win by reweighting. The "do now" stack is five items, the number of concurrent workstreams the PMO says they can staff. Near-ties are shown as ties. Dependencies (night-count before a new yard system) sit on the sequence view.

If Elena still wants night-count off the page, she has to say so as a partner call, not as a score.

Hidden weights, a stack of forty, and a buried idea

Three failure modes show up in the same week if you do not design for them.

Hidden weights. Re-running with more weight on "strategic," or letting the model pick a textbook split, is the same move: the rank encodes taste. If weights change after first scores, show both runs. A silent second pass is a cooked exhibit.

Ranking forty so nothing is a priority. A full sort feels complete. Owners keep their rows. Steering nods. Capacity does not appear. Force a hard cap on "do now" from the client's delivery lead, not from the length of the workshop list.

Using the matrix to bury an idea the partner does not like. Down-weighting a criterion until an unloved row falls is not analysis. Kill it in the open, or leave it on the stack with a readiness or sequence note. Steering can smell a cooked rank.

Before the readout, run the stack through the recommendation adversarial stress-tester. Ask what a hostile CFO would say about the weights, the cap, and any row that mysteriously died. Do not let the stress-tester rewrite the scores.

The artifact that should leave the room is a short do-now stack, evidence notes on the cells, a readiness column that did not leak into impact, a sequence view for dependencies, and the signed weight note. Top items that need money get a range from the implementation business case generator after the rank exists, not before. Do not invent a payback figure to justify a row you already wanted.

If the first output is a thirty-eight-row sort with model-chosen weights and a two-by-two that hid a near-tie, you ran the wrong job.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first