Back
The 4 Levels of Freight Execution: How to Tell Whether a Freight AI Platform Actually Does the Work

Alex Amorello

Let me guess how your week went. You sat through three freight AI platform demos, and they all sounded the same. Every vendor said “agentic,” and every demo featured a carrier who answered on the first ring with a truck 20 minutes out. You’ve been in freight long enough to know that carrier doesn’t exist.
Vendors love the word “agentic” the way gas stations love the word “fresh.” Gartner calls the practice agent washing, which means relabeling ordinary software as agentic AI. In June 2025, Gartner estimated only about 130 of the thousands of vendors making that claim were legitimate. It warned supply chain buyers again in May 2026.
Full disclosure: I sell Ellie for Envoy, so I’m one of those vendors in a quarter zip. Read this the way you’d read any rep on commission, with one eyebrow raised. In the end, I’ll give you a test that works even if you don’t trust a word I say.
What are the 4 Levels of Freight Execution on One Screen?
First and foremost, the four levels of freight execution grade any freight AI platform by how much work it finishes without a person touching it.
Save this table for your next demo, and watch the sales guy sweat.
Level | What the software does | What’s still your rep’s job | How to spot it in a demo |
1. Reports | Flags the problem and turns the load red | Everything: find a truck, negotiate, book, call the customer | Dashboards, and “visibility” on every slide |
2. Recommends | Picks carriers, suggests a rate and drafts the emails | Click approve, then chase callbacks and book | The demo ends on an Approve button |
3. Executes | Tenders, books, updates the TMS and follows up, once someone tells it to | Notice the problem and kick it off | Nothing happens until someone clicks Run |
4. Owns the exception | Spots the problem, books a vetted truck inside your rules and tells the customer | Only the calls your policy saves for a human | It fixes a problem nobody told it about |
What Do These Freight Execution Levels Look Like in Practice?
In practice, each level hands the same rep a different Monday morning. So I’m going to ruin a Sunday night in Chicagoland.
You’ve got a reefer of frozen deep-dish pizza picking up in Bolingbrook at 6 a.m. Monday, with a $1,400 max buy. It’s headed to a grocery DC in St. Louis. Our CEO, Robby Nathan, is a die-hard Cubs fan, so he considers that enemy territory.
At 9:40 Sunday night, the driver texts your rep that his truck died on I-55 outside Joliet. She’s asleep with her phone face down on the nightstand, which is the healthiest thing happening at your brokerage all weekend.
What she finds at 6:30 Monday morning depends on your software.
What Is Level 1 Freight Execution?
Level 1 freight execution is software that can tell you something went wrong, but can’t do anything about it.
Your visibility tool sees the truck stop outside Joliet at 9:41 Sunday night. It flags the load, turns the screen red, and maybe sends an alert. Useful, sure. But the actual recovery still waits for your rep.
She opens the laptop at 6:30 Monday morning and now owns the whole mess: find a replacement reefer, stay under the $1,400 buy, protect the St. Louis appointment, and somehow keep the frozen pizza from becoming somebody else’s problem.
That’s Level 1. The software observes. The human executes.
A lot of brokerages have stacked several tools around this model. The TMS records what happened. Visibility shows where it happened. BI tells you later how often it happened. Robby calls it the software paradox: plenty of software, while the rep still does the work.
What Is Level 2 Freight Execution?
Level 2 freight execution goes one step beyond Level 1: the software can recommend the recovery, but it still needs a human to approve it.
So our reefer dies at 9:41 Sunday night, and this time the system gets busy. It finds five carriers that run Chicagoland to St. Louis, suggests a $1,250 buy, and drafts the outreach.
Then it waits.
Nothing goes to the carriers until your rep wakes up and clicks approve at 6:30 Monday morning. By then, two trucks are gone, and the other three know exactly what a Monday-morning rescue is worth. Suddenly that $1,400 ceiling looks decorative.
This is why Level 2 gets confused with automation. The software did some thinking and prepared the next move. Your rep still had to be there at the exact moment the move needed to happen.
Robby has a pretty simple test for this stuff: What happens if nobody clicks approve? If the answer is, “It sits in a queue,” you’ve got an email drafter wearing an agent costume.
What Is Level 3 Freight Execution?
Level 3 freight execution is where the software finally stops suggesting the work and starts doing it. A human still has to kick things off, but once they do, the system can run the recovery itself.
For our stranded reefer, that means your rep doesn’t wake up to five drafted emails and a blinking approval button. She starts the recovery, and the software takes over from there: tendering the load, working carriers, booking one, updating the TMS, rescheduling the appointment, and chasing the rate con.
Sounds simple until you remember one load can bounce across a TMS, load boards, inboxes, phone calls, and a shipper portal designed by someone who apparently hates freight brokers.
Ellie, though, already works this way. On one customer’s first posted load, she handled 76 calls in five minutes, pulled 20 offers, and booked below the posted rate. What’s more, across live accounts, about 75% of freight books through her.
The limitation is timing. Level 3 can run the play, but somebody still has to call it.
What Is Level 4 Freight Execution?
Level 4 freight execution is where the software owns the exception from start to finish. It detects the problem, decides what to do within rules you’ve already set, takes action across your systems, and pulls in a human only when the situation falls outside those rules.
So when our driver texts at 9:41 Sunday night, the recovery starts right then. The system reads the message, recognizes the dead truck outside Joliet, works the St. Louis lane, vets the available carriers, and books one at $1,325. Nobody had to script this exact scenario, and nobody had to wake up to launch it.
The guardrails still matter. Carrier selection has legal consequences, so the system follows your vetting standards, documents the decision, and escalates anything outside policy. Ellie handles exceptions as far as you let her, while every one adds context to your Carrier Context Graph.
So, How Can You Tell Which Freight Execution Level a Platform Really Reaches?
Once you understand the four levels, the final job is figuring out which one a vendor actually delivers. The cleanest way is to stop watching polished demos and give the platform one ugly, live exception inside your own operation.
Use a real lane, your real SOPs, and your real max buy. Then watch what happens when nobody helps.
See If It Notices First: Recreate a recent exception, and don’t tell the system where to look. If a rep has to flag the problem before anything happens, you already know how far up the ladder you are.
Push on the Guardrails: Give it a carrier that looks tempting but fails your standards. A three-week-old MC and a truck 10 minutes away is exactly the kind of shortcut that exposes whether the system follows policy when pressure hits.
Make It Finish the Work: Don’t stop at a recommendation or draft. See whether it can actually tender, vet, book, update the TMS, and close the loop without handing the load back to a person halfway through.
Make It Show Its Homework: “Done” means very little if nobody can see what was checked. Berkeley’s Agents’ Last Exam found agents often claimed success without completing the underlying checks. In freight, that can mean a load marked covered on a carrier whose insurance lapsed Tuesday.
Loads Per Rep Is Where the Levels Show Their Work
Eventually, all four levels land in the same place: how much freight one person can actually run.
Level 1 tells your rep there’s a problem. Level 2 prepares a response. Level 3 does the work after somebody starts it. Level 4 takes the exception off the rep’s plate unless your rules say otherwise. Every step removes a little more human effort, but the jump gets interesting when your best people stop spending half the afternoon rescuing loads.
Robby calls it execution per human decision. I like the term because it forces a harder conversation than: “How much AI do you have?”
C.H. Robinson reported 15%+ year-over-year productivity growth while headcount fell 10.8%. Most brokerages aren’t going to build Robinson’s engineering organization. They also shouldn’t need one.
Before you add another rep to the 2027 headcount plan, book 30 minutes with me. Bring the ugliest exception your team dealt with last week, give us your rules, and let Ellie run it live inside your TMS.
If she can’t handle it, don’t buy her.
But, of course, if you’d rather hear it from somebody who already did, contact me and I’ll set up a reference call with a brokerage using Ellie in production. I’m more than happy to put my money where my mouth is.


