โ†
AI for Trucking, Fleet & Freight
Visionary ยท M10 ยท lesson 10 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Fleet AI Pilots to Evidence
๐Ÿ“–
now learning

Running Fleet AI Pilots to Evidence

15 min

The fleet manager at a 500-truck regional carrier signed a 12-month contract with an AI dispatch optimization vendor in January after seeing a conference demo that showed a 22 percent deadhead reduction on a Southeastern fleet. By June, the fleet's own deadhead number had improved by 3 percent on the pilot lane but the safety director had flagged two instances where the AI-generated route plans conflicted with current hours-of-service (HOS) rules, and the owner was asking why the ROI projections from the vendor's slide deck bore no resemblance to the carrier's actual results. The problem was not the AI. The problem was the pilot design: no pre-defined success criteria, no HOS compliance gate built into the run conditions, no control lane to establish a baseline, and a vendor benchmark derived from a fleet that ran different lanes, different freight, and a different driver population. By the time the evidence was in, it was too late to renegotiate the contract or adjust the rollout plan. This lesson teaches the pilot discipline that prevents this failure: the structured, safety-gated, evidence-producing fleet AI pilot that satisfies an owner's demand for ROI proof and a safety director's demand for compliance assurance before a single additional lane goes live.

Why Most Fleet AI Pilots Fail to Produce Evidence

A freight AI pilot that does not produce usable evidence is not a pilot. It is an expensive experiment with no conclusion. The distinction matters because the purpose of a pilot is not to try the technology. The purpose is to generate the specific evidence that answers the decision question: does this AI application, running on this carrier's actual freight, with this carrier's actual drivers and actual HOS constraints, produce the improvement it was designed to produce at a safety and compliance level that is sustainable? That question has a yes or no answer, and the pilot is the mechanism for reaching it before the carrier commits to network-wide deployment.

Most fleet AI pilots fail to produce that evidence because they are designed to confirm a hypothesis rather than test one. A vendor demo creates a positive prior: the dispatcher has seen the technology work in a best-case environment, the owner has approved a budget, and the operations team wants to validate the investment. When the pilot is designed by people who want it to succeed, the bias shows up in structural choices that prevent negative evidence from emerging. Success criteria are set vaguely enough to be satisfied by almost any result. The pilot lane is the carrier's best lane rather than a representative one. The control condition is a period from a different season or a different market environment. The HOS compliance check is informal rather than systematic. The result is a pilot that produces a positive-looking summary with no statistical validity and no honest comparison to the alternative, and a carrier that now has to either deploy a system it cannot actually defend or admit the pilot was inconclusive.

The second failure mode is the pilot that runs too long without clear decision rules. A carrier that runs a pilot for 12 months without pre-specified criteria for what result would trigger a go decision and what result would trigger a stop decision is not running a pilot. It is running the technology on real freight with real drivers and real safety exposure while calling it a test. A pilot without pre-specified decision rules will always find a way to continue because the sunk cost of the trial creates organizational inertia toward a go decision regardless of the evidence.

The third failure mode is a pilot that measures the wrong things. Vendors naturally measure the metrics their technology improves most visibly: dispatch cycle time, load matches per dispatcher hour, or average route score. These metrics may be real improvements that do not translate into margin or safety outcomes. The carrier's owner and safety director need to see deadhead percentage, revenue per truck per day, HOS violation rate, and CSA (Compliance, Safety, Accountability) score change, because those are the metrics that determine whether the technology is working for the carrier's actual business problem. A pilot designed around vendor metrics rather than carrier business outcomes produces vendor evidence, not carrier evidence.

A pilot that cannot tell you what success looks like before it runs cannot tell you whether it succeeded when it ends. Define the evidence before the first load is touched, not after the results are in.

The Pilot Design Framework: Four Elements That Matter

A fleet AI pilot that produces actionable evidence has four structural elements that must be defined before the first truck runs: a pre-specified primary metric with a minimum success threshold, a valid control condition for comparison, a safety compliance gate that is non-negotiable, and a pre-specified decision rule that will be honored when the pilot ends. Each element is simple in concept and difficult in practice because each one requires the carrier's operations leadership to make commitments in advance that limit their flexibility to interpret the results in their favor.

Element One: The Primary Metric and Success Threshold

The primary metric for a freight AI pilot should be the carrier's actual operational problem, not the vendor's feature set. For a dispatch optimization pilot, the primary metric is deadhead percentage on the pilot lanes or revenue per truck per day, not load matches per dispatcher hour. For a predictive maintenance pilot, the primary metric is avoided roadside breakdown events or total maintenance cost per truck per month, not alert volume or alert acknowledgment rate. For an HOS compliance prediction pilot, the primary metric is HOS violation rate on pilot lanes, not HOS recommendation volume or recommendation acceptance rate by dispatchers.

The success threshold is the minimum improvement that justifies network-wide deployment, stated as a specific number before the pilot begins. A carrier whose dispatch optimization baseline is 18 percent deadhead on the pilot lanes might set a success threshold of 15 percent deadhead, a 3 percentage point improvement, sustained over the pilot period. If the pilot produces 16.5 percent deadhead, the technology improved the metric but did not reach the threshold, and the decision rule for that outcome (renegotiate terms, require a second pilot period, or decline deployment) must already be in place. The threshold removes the negotiation that happens at the end of a pilot when operations wants to deploy and the vendor presents a positive-looking number that does not quite meet the carrier's original expectation.

The success threshold should be set by the carrier's operations leadership based on the economics of the carrier's specific freight, not by the vendor. A vendor that proposes the success threshold for its own technology has a conflict of interest that will consistently produce thresholds the technology can satisfy rather than thresholds that reflect the carrier's actual business need. The carrier's owner or fleet manager sets the threshold, the vendor agrees to it before the pilot begins, and the contract includes language specifying what happens to the commercial relationship if the threshold is not reached.

Element Two: The Control Condition

The most common error in fleet AI pilot design is the absence of a valid control condition. A carrier that runs a dispatch optimization AI on its Atlanta-to-Miami lane in Q1 and compares the results to the same lane's performance in Q4 of the prior year has not established that the AI drove the improvement. Seasonal freight volume changes, fuel price fluctuations, driver turnover, and shipper mix changes all affect deadhead and revenue metrics independently of any AI application. A pilot without a valid control condition is not measuring AI performance. It is measuring the difference between two periods that differ in many ways besides the presence of AI.

The cleanest control condition for a freight AI pilot is a simultaneously running control lane: a lane of similar freight type, similar distance, similar driver count, and similar shipper mix that runs on existing dispatch processes during the same period as the pilot lane. The difference in performance between the pilot lane and the control lane during the same period, holding other variables as constant as the carrier's network allows, is the closest thing a freight operation can produce to a controlled experiment. It eliminates the seasonality bias that plagues before-and-after comparisons and gives the carrier a defensible answer to the owner's question: did this AI improve our results, or did everything improve because it was a better freight market?

When a simultaneous control lane is not operationally feasible (because the carrier's network does not have a comparable parallel lane, or because the AI system cannot be selectively applied to some lanes and not others without operational disruption), the alternative is a holdout dispatcher design: one or two dispatchers running the same lanes on existing processes while a matched group of dispatchers uses the AI system. This design is more susceptible to the Hawthorne effect (the phenomenon where people perform differently because they know they are being observed) but produces a concurrent comparison that is more valid than a before-and-after comparison across periods.

Element Three: The Safety Compliance Gate

The safety compliance gate is the non-negotiable element that distinguishes a responsible fleet AI pilot from a technology experiment on real freight and real drivers. The gate operates as a mandatory check on every AI-generated recommendation before it is acted upon during the pilot period, and it is the mechanism that prevents the pilot from generating ROI evidence while simultaneously generating CSA score damage or driver safety risk.

For a dispatch optimization pilot, the safety compliance gate has three components. The first is the HOS verification: every AI-proposed route plan is checked against the driver's current ELD (electronic logging device) status and remaining HOS hours before the plan is dispatched. An AI route plan that would require the driver to exceed HOS limits is not a valid plan and must not be dispatched, regardless of its optimization score. This check must be performed by a qualified person with access to the real-time ELD data, not by the AI itself, because the AI's HOS calculation may be based on a regulatory version that predates recent FMCSA (Federal Motor Carrier Safety Administration) updates or may use simplifying assumptions that do not account for the driver's specific exemptions or exceptions.

The second component is the driver vehicle inspection report (DVIR) status check: no AI-optimized load assignment may be confirmed for a truck with an open DVIR defect that has not been cleared by a certified technician. An AI that schedules the most efficient load for a truck with a pending brake inspection is not optimizing safely; it is creating a maintenance compliance failure.

The third component is the safety director sign-off threshold: if the pilot generates more than a pre-specified number of safety compliance gate trips (situations where the AI recommendation was blocked because it failed the HOS or DVIR check) in any pilot period, the safety director reviews the trip pattern before the pilot continues. A high gate trip rate indicates that the AI's recommendations are systematically miscalibrated relative to the carrier's compliance reality, and that miscalibration should halt the pilot while the vendor investigates rather than produce a series of gate trips that never get evaluated collectively.

For a predictive maintenance pilot, the safety compliance gate works differently but serves the same function. Every AI-generated maintenance recommendation that involves a safety-critical component (brakes, tires, steering, lights) must be reviewed by a certified diesel technician before the truck is cleared to continue operating. The AI's prediction is a decision-support input, not a pass/fail determination. A technician who reviews the AI's brake-wear prediction and finds it inconsistent with a physical inspection has the authority to override the AI recommendation and must document the override. That documentation is part of the pilot evidence: a pattern of technician overrides indicates the AI model needs recalibration, while a pattern of AI predictions confirmed by technician inspection indicates the model is performing correctly.

Element Four: The Pre-Specified Decision Rule

The pre-specified decision rule converts the pilot from an open-ended trial into an actual experiment. It states, in writing before the pilot begins, exactly what the carrier will do at the end of the pilot period given each possible outcome. A carrier that pre-specifies its decision rules cannot be maneuvered by organizational inertia or vendor pressure into deploying a technology that did not reach the success threshold.

The decision rule framework has three outcome buckets. The first bucket is the success case: the primary metric reached the success threshold, the safety compliance gate ran without systematic issues, and the control comparison supports the conclusion that AI drove the improvement. In this case, the decision rule authorizes the carrier to proceed with network-wide deployment under the terms already negotiated, with the safety compliance gate requirements carried forward into production operations.

The second bucket is the marginal case: the primary metric improved but did not reach the success threshold, or the metric reached the threshold but the safety compliance gate tripped at a rate above the pre-specified limit. In this case, the decision rule specifies a structured path: renegotiate the vendor contract to reflect actual performance, require a second pilot period with an adjusted configuration, or decline to proceed with network-wide deployment. The marginal case is the one most likely to produce ambiguous pressure on the carrier's decision-makers, which is why it requires the most explicit pre-specification.

The third bucket is the failure case: the primary metric did not improve meaningfully, or the safety compliance gate generated a pattern of HOS miscalculations or systematic errors that the vendor could not explain. In this case, the decision rule terminates the pilot and requires the vendor to provide a root cause analysis before any further commercial discussion. Pilot failure is not a catastrophe. It is the point of the pilot: to find out at pilot scale whether the technology works on this carrier's freight before it is deployed at network scale.

What Evidence an Owner Actually Needs

The carrier owner who authorized the pilot budget wants to answer one question: is this AI worth the money? That question has a specific answer structure that is different from the evidence a vendor wants to show and different from the evidence a technology enthusiast finds compelling. The owner needs three pieces of evidence that together constitute a defensible investment thesis.

The first piece is the margin impact: a specific dollar figure representing the improvement in operating contribution per truck per day attributable to the AI application on the pilot lanes, net of the AI's cost. For a dispatch optimization pilot, this is the revenue recovered from deadhead reduction (dead-mile cost avoided plus revenue from loads that replaced empty legs) minus the total cost of the AI system over the pilot period, divided by the number of trucks on the pilot lanes. If this number is positive and above the carrier's minimum acceptable return threshold, the first piece of evidence supports deployment.

For a predictive maintenance pilot, the avoided cost calculation has two components: the cost of roadside breakdowns that were predicted and prevented (including the breakdown event cost, the towing cost, the repair cost premium for roadside versus in-shop work, the driver downtime cost, and the shipper relationship cost), and the change in total maintenance cost per truck per month from the shift toward scheduled in-shop intervention rather than reactive roadside repair. The documented benchmark for this calculation is approximately 34 percent maintenance cost reduction on a roughly 44-day payback, which the carrier's own pilot evidence should either confirm, exceed, or fall short of. If the carrier's pilot produces 28 percent maintenance cost reduction, the technology is working but at a different level than the industry benchmark, and the network deployment economics should reflect the carrier's own evidence rather than the benchmark.

The second piece is the compliance trajectory: a comparison of the CSA score trajectory on pilot lanes and control lanes during the pilot period. If the AI application is improving dispatch efficiency and reducing HOS violations, the carrier's HOS compliance BASIC score should improve on pilot lanes. If the AI is producing maintenance interventions that prevent vehicles with safety defects from running, the vehicle maintenance BASIC score should improve. If neither metric moves, or if they move in the wrong direction, the AI application is not producing the compliance benefit that makes it safe to deploy at network scale. The compliance trajectory evidence is the safety director's primary input to the deployment decision, and it carries a veto right: an owner who wants to deploy an application that the safety director cannot endorse on compliance grounds is assuming personal liability for the safety risk the application creates.

The third piece is the scalability assessment: a realistic analysis of whether the pilot conditions can be reproduced at network scale. A pilot that ran on the carrier's most favorable lanes with the carrier's best dispatchers and the carrier's most experienced drivers is not evidence that the technology will produce the same results when it is deployed across all lanes, all dispatchers, and all drivers. The scalability assessment asks three questions: does the AI perform comparably on the carrier's harder lanes (longer distances, more complex freight, tighter HOS margins) as on the pilot lanes? Does the AI's performance depend on dispatcher engagement levels that cannot be maintained across the entire dispatch team? And does the AI require data quality (TMS data completeness, ELD data accuracy, telematics feed reliability) that is uniform across the network or only present on the pilot lanes?

The Safety Director Sign-Off Structure

The safety director's role in a fleet AI pilot is not to approve or reject the technology. The safety director is not a technology evaluator. The safety director's role is to certify that the AI application, as it ran during the pilot period and as it is proposed to run in network-wide deployment, does not create safety or compliance risk that the carrier is unable to manage within its existing safety management system.

The safety director sign-off at pilot conclusion has four components. The first is the safety compliance gate review: a summary of every instance during the pilot where the gate blocked an AI recommendation, including the reason the recommendation was blocked, the action that was taken instead, and whether a similar recommendation from the same AI system was blocked more than once (indicating a systematic issue rather than an isolated error).

The second component is the HOS audit: a review of the ELD records for all drivers on pilot lanes during the pilot period, comparing the AI-recommended route plans to the actual hours logged, to verify that the AI recommendations did not contribute to any HOS violations, even in cases where the violation was attributed to driver behavior rather than dispatch decision. A dispatcher who follows an AI recommendation that turns out to push the driver's HOS to the limit is exposed even if the driver technically chose to accept the assignment, because the AI recommendation established the operational expectation.

The third component is the DVIR and maintenance incident review: a comparison of DVIR defect discovery rates on pilot lanes versus control lanes, and a review of any maintenance incidents (roadside breakdowns, repair events, inspection failures) on pilot trucks during the pilot period. For a predictive maintenance pilot, this comparison is the core of the safety director's evidence. For a dispatch optimization pilot, the maintenance incident review is a secondary check that confirms the dispatch optimization pressure did not encourage drivers or dispatchers to accept assignments with trucks that had pending maintenance concerns.

The fourth component is the network deployment safety plan: a written plan, reviewed and signed by the safety director, specifying how the safety compliance gate will operate in network-wide deployment, what the safety director's intervention criteria are if the network-wide deployment produces an increase in gate trips or HOS violations, and what the safety director's role is in the ongoing monitoring of the deployed system. A safety director who signs off on network deployment without a written intervention plan has certified that the system is safe to deploy but has not built in the mechanism to act if the deployment does not maintain pilot-level safety performance.

Pilot Structure for Autonomous Lane Integration

The pilot discipline for autonomous freight capacity integration through Aurora's McLeod TMS (transportation management system) integration has additional structural requirements beyond the standard fleet AI pilot framework, because autonomous lane pilots involve the public road operation of a SAE Level 4 system within its operational design domain (ODD) and the coordination of autonomous capacity with human-driven loads in a mixed dispatch environment.

The primary metric for an autonomous lane integration pilot is lane contribution margin: revenue per loaded mile on the autonomous lane minus the fully loaded cost of the autonomous capacity (per-mile rate, terminal handling cost, and any first-mile or last-mile driver cost for legs outside the ODD) compared to the same calculation for human-driven capacity on comparable lanes during the same period. Aurora's current commercial pricing is structured on a per-mile basis through the TMS booking interface, which allows the carrier to calculate a direct lane-level contribution margin comparison without the ambiguity of overhead allocation.

The safety compliance gate for an autonomous lane pilot has a different structure than a human-driven fleet pilot. Because Aurora's SAE Level 4 system manages its own in-ODD operation, the carrier's HOS-based safety gate does not apply to the autonomous driving portion. The carrier's gate responsibilities in an autonomous lane pilot are concentrated at the ODD boundary: verifying that every load booked on an autonomous lane falls within the Aurora system's current published ODD limits (specific highway segments, weather conditions, daylight operating windows, and load weight limits), and that the first-mile and last-mile segments handled by human drivers are planned with adequate HOS margin. An autonomous lane booking that sends a driver on a 90-minute first-mile segment immediately after they have consumed 10 hours of their HOS allowance has not produced a safe plan; it has moved the HOS risk from the highway to the drayage segment.

The pre-specified decision rule for an autonomous lane integration pilot must address two specific outcomes that do not arise in standard AI pilots. The first is an ODD violation: a situation where an Aurora vehicle operates outside its declared ODD due to a dispatch booking that did not correctly verify ODD constraints. The decision rule for an ODD violation should specify an immediate review by the carrier's safety director and a temporary hold on new autonomous lane bookings until the root cause (whether a TMS integration error, a dispatcher procedure failure, or an ODD boundary ambiguity) is identified and corrected. The second specific outcome is an autonomous vehicle intervention event: a situation where the Aurora system encounters a condition it cannot handle within its ODD and takes a safe action (such as stopping the vehicle in a safe location and requesting human assistance). The decision rule for an intervention event should specify how the carrier will coordinate with Aurora's operations team, what the driver notification procedure is, and how the incident will be documented for FMCSA records.

Key Takeaways

  • A fleet AI pilot that does not produce usable evidence is an expensive experiment with no conclusion. The pilot's purpose is to generate specific evidence answering whether the AI produces the designed improvement on this carrier's freight at a compliance level that is sustainable.
  • The four structural elements of a valid pilot are: a pre-specified primary metric with a minimum success threshold set by the carrier's operations leadership, a valid simultaneous control condition, a non-negotiable safety compliance gate covering HOS verification and DVIR status, and a pre-specified decision rule that will be honored regardless of sunk cost or organizational pressure.
  • The safety compliance gate for a dispatch optimization pilot requires that every AI-proposed route plan be verified against real-time ELD status by a qualified person before dispatch, because the AI's HOS calculation may be based on regulatory text that predates recent FMCSA updates.
  • The evidence an owner needs is three-part: the margin impact in dollars per truck per day net of AI cost, the CSA compliance trajectory on pilot versus control lanes, and a scalability assessment confirming the pilot conditions can be reproduced across harder lanes and less experienced dispatchers.
  • The safety director's sign-off at pilot conclusion certifies that the AI application does not create safety or compliance risk the carrier cannot manage within its existing safety management system, and must include a written network deployment safety plan with intervention criteria, not just an approval of the pilot-period results.
  • For autonomous lane integration pilots through Aurora's McLeod TMS integration, the primary metric is lane contribution margin, the safety gate is concentrated at ODD boundary verification rather than HOS compliance on the highway segment, and the pre-specified decision rule must address both ODD violations and autonomous vehicle intervention events.
  • Accountability stays human throughout the pilot: the fleet manager who signs the pilot design brief, the safety director who approves the compliance gate, and the owner who authorizes network deployment based on the pilot evidence are each on record as having made a decision that they own, regardless of what the AI recommended.
  • The predictive maintenance pilot benchmark to verify against is approximately 34 percent cost savings on approximately a 44-day payback. If the carrier's pilot evidence diverges significantly from this benchmark, the network deployment economics should be based on the carrier's actual evidence, not the industry figure.