Skip to main content
DoneThat

AI Adoption GuideLogisticsPlan

CO2 Emission Plan Scorer

ML scores each route plan variant by estimated carbon footprint and flags plans that exceed regulatory or corporate targets.

Logistics processBookPlanPickLoadMoveDeliverConfirmClose

By Don, DoneThat’s AI coach · updated

Why route plans need a carbon score before dispatch

A route plan that looks cheap and on-time can still blow a carbon budget. Logistics teams now face EU ETS for maritime, FuelEU Maritime, Corporate Sustainability Reporting Directive (CSRD) disclosure pressure, and shipper contracts that write absolute CO2 caps into tender awards. When planners compare five multimodal variants for the same order set, distance and cost alone no longer decide which plan is acceptable.

A CO2 emission plan scorer sits after candidate generation and before commit. It takes each plan variant, estimates well-to-wheel or tank-to-wheel CO2e for every lane segment, rolls those estimates up to a plan-level footprint, and compares the result to regulatory thresholds and corporate targets. The score is a quality signal, not an autopilot. The planner still picks the route. The scorer’s job is to make the carbon consequence of that pick explicit, auditable, and comparable across variants.

Without that step, carbon targets live in a spreadsheet that nobody opens at 06:00 when a reefer lane fails and three alternatives appear. With it, every variant carries a footprint, a pass/fail against policy, and a citation trail that explains how the number was built.

What the scorer consumes and what it refuses to invent

The scorer needs structured plan variants: ordered stops, modes per leg (road, rail, short-sea, deep-sea, air), equipment or vessel class where known, and distance or duration for each lane segment. It also needs an emission-factor catalog keyed by mode, fuel or energy type, geography, and methodology version. Those factors typically come from libraries aligned with GLEC Framework or ISO 14083 practice, or from specialist calculators such as EcoTransIT World for multimodal freight and Searoutes for ocean and inland-waterway legs.

Distance is non-negotiable. If a lane segment has no distance (and no reliable proxy the factor set accepts), the scorer returns empty for that plan rather than guessing. An empty score is intentional. Fabricated kilometers produce fabricated tonnes, and fabricated tonnes destroy trust in both the dashboard and the audit file. Upstream systems (Trimble for road network distance, project44 for live shipment context that may refine mode or empty-return assumptions) should fill gaps before rescoring, not after a false green light.

When data is complete, each plan score cites:

  • The lane segments included in the rollup (origin–destination pairs, mode, distance basis)
  • The emission factor IDs applied to each segment (catalog key, version, scope such as TTW vs WTW)
  • Any uplift or empty-return rules applied at plan level
  • The corporate or regulatory target used for the flag (absolute tonnes, intensity such as gCO2e/tkm, or both)

That citation package is what turns a single number into something a sustainability team, a customer, or an auditor can replay.

How scoring and flagging work in the planning loop

Candidate generation (including AI-assisted route optimization) produces N feasible plans. The scorer runs in batch against those N. For each plan it:

  1. Decomposes the itinerary into lane segments with mode and distance.
  2. Looks up emission factors by segment attributes and applies them consistently (same methodology version across all variants in a comparison set).
  3. Aggregates to plan CO2e, optionally normalized by tonne-kilometers when intensity targets apply.
  4. Compares against configured limits: hard regulatory caps where they bind the move, softer corporate science-based targets, and customer-specific SLAs when present.
  5. Emits a quality outcome: numeric score, status (within target / exceeds target / unscorable), and the citation payload above.

Exceeding a target does not auto-reject the plan. It flags it. Planners may still choose a higher-carbon option when capacity, service failure risk, or ML transit time prediction shows the greener variant will miss a hard delivery window. The flag exists so that choice is conscious and logged, not accidental.

Rescoring matters as much as first-pass scoring. When a dynamic replanning agent swaps a missed rail slot for trucking, the carbon profile changes immediately. The scorer should re-run on the new variant set with the same factor catalog version used in the original comparison, so deltas are apples-to-apples. Pairing carbon flags with a KPI trend anomaly monitor then surfaces whether a corridor is drifting toward chronic exceedances rather than one-off exceptions.

Emission factors, vendors, and methodology discipline

Carbon numbers only travel if the methodology is stable and named. Teams typically mix:

  • EcoTransIT for standardized multimodal freight emission calculations across road, rail, inland waterway, and sea, useful when plans hop modes inside one itinerary.
  • Searoutes for ocean and related routing geometry plus emissions estimates when deep-sea legs dominate the footprint.
  • Trimble for accurate road distances and network attributes that feed factor application on truck and last-mile segments.
  • project44 for visibility and shipment-state context that keeps scored plans aligned with what is actually moving, including mode shifts and dwell that affect empty or repositioning assumptions.

The scorer should not silently blend incompatible factor sets inside one comparison. If EcoTransIT factors score three variants and a fourth pulls a one-off carrier spreadsheet factor, the comparison is invalid. Pin factor IDs and catalog versions on every score. When the catalog upgrades (new GLEC year, revised vessel class averages), reprocess historical plans only under an explicit version bump so trend lines do not jump without explanation.

Scope must be stated on every output: tank-to-wheel versus well-to-wheel, CO2 versus CO2e, inclusion or exclusion of refrigerants for cold chain. Corporate targets and regulatory reports often require different scopes. The scorer can compute both, but each flag must name which scope it used.

Operating the quality outcome without overriding the planner

Outcome type for this use case is quality. The artifact is a scored, cited plan assessment, not a booked itinerary. Empty scores when distance data is missing keep quality honest. Non-empty scores that exceed targets keep quality visible.

Recommended operating rules:

  • Block on empty only in audit modes. In live planning, empty means “do not claim a carbon advantage,” not “freeze the desk.” Require distance backfill before any customer-facing green claim.
  • Surface deltas, not only absolutes. Show tonnes and intensity versus the cheapest and versus the lowest-carbon variant so trade-offs are readable in one glance.
  • Log planner overrides. When a flagged plan is selected, capture reason codes (capacity, SLA, customer request). Those codes feed later network design and carrier RFP work.
  • Keep humans on commit. Automation can rank by carbon among feasible plans; it should not force the lowest-carbon variant when service constraints conflict unless policy explicitly says so for that lane.

Success looks like fewer silent exceedances, faster answers to “what does this plan emit,” and cleaner evidence when customers or regulators ask how a reported figure was derived. Failure modes to watch: scoring only the winning plan (comparison disappears), caching stale distances after a network update, and treating intensity targets as absolute caps (or the reverse) without labeling.

Implementation checklist for a trustworthy scorer

Start with a closed set of high-volume corridors where distance data is reliable and modes are well labeled. Wire factor IDs from a single catalog version. Integrate EcoTransIT and/or Searoutes for non-road legs, Trimble (or equivalent) for road kilometers, and project44 (or equivalent) for execution feedback that triggers rescoring. Define corporate and regulatory thresholds per lane family. Emit scores with segment and factor citations. Return empty when distance is missing. Leave route selection with the planner, and feed flags into replanning and KPI anomaly workflows so carbon quality stays in the same loop as cost and time.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first