Back
The Wrong AI Race: Why Enterprise AI Will Be Won by Execution, Not Models

Robert Nathan

The best carrier rep I ever worked with knew which reefer outfits picked up before 7 a.m. and which ones went quiet the week that produce season started. Ten years of pattern recognition walking the floor with a headset on.
And what did I have her doing most of the day? Retyping rate confirmations into the TMS and sending the same capacity email to the same 40 carriers.
I can’t believe I paid humans to send emails. I’ve said it enough that it’s become furniture, but sit with what it actually means. I hired intelligence and spent it on clicking. The sharpest minds in my building did the dumbest work in it, and so did yours, because for 20 years no other way to run a floor existed.
The AI industry is now repeating my mistake at scale, worshiping the intelligence and ignoring the work, which is why so much of enterprise AI still can’t cover a load. Wrong race. I built Envoy for the other one: Ellie, the execution layer for logistics, now running whole floors through Ellie Workforce and booking about 75% of freight in our live accounts.
We Already Ran This Race Once
Brokerages ran their own version of the model race for years. We competed on who could hire the sharpest reps, train them hardest, feed them the most lane knowledge. Then we sat those people in front of five portals and a TMS and converted all of it into data entry.
The loads we lost, we almost never lost on judgment. We lost them on the clock. My rep knew the right carrier before she finished reading the posting; she got to it fourth, behind two confirmations and a check call, and by then the cheap trucks were gone. Intelligence was never the scarce input on a brokerage floor. Execution capacity was.
So when a vendor opens a pitch with benchmark scores, I hear a guy bragging about his hiring class while the floor drowns. If it doesn’t change execution, it isn’t real.
The Model Race Just Ended in a Tie
You don’t have to take my word on the model part, and given what I sell, you probably shouldn’t.
Take Stanford’s instead. The 2026 AI Index puts the top frontier labs within about 25 Elo points of each other, the leader up 2.7%. Six labs, one cluster, functionally a tie, and every shop in your revenue band can rent any of them by this afternoon.
The same research cycle found agents failing roughly 1-in-3 production attempts. Genius on the benchmark, a no-show on one of every three tries on the job.
I stopped having model opinions sometime in 2025. The model matters the way an engine matters, and no shipper has ever asked what’s under the hood. What decides enterprise AI outcomes now is everything wrapped around the model, the tools, the memory, the policies, and who welded all of it to freight.
The Buyers Figured It Out Before the Vendors
Two years ago I said there’s no future in selling technology to freight brokers and got told I was being dramatic.
Then the pricing data started agreeing with me. Which is an annoying way to win an argument.
Futurum’s survey has fewer than 1-in-5 enterprise software buyers still wanting per-seat pricing; 43% want consumption, and 27% want to pay for outcomes outright. Bessemer politely calls it the AI pricing pivot. I call it services and SaaS converging, right on schedule.
Buyers quit paying for access to intelligence and started paying for finished work, and the supply side can’t keep up: McKinsey finds no more than 10% of organizations scaling AI agents in any single function. Real budgets, scarce production.
Outcome buying has a catch, though. You have to watch the outcomes happen. That’s what TOAS exists for.
Where the Pilots Meet Freight
The autopsy numbers are public. MIT’s researchers found 95% of enterprise AI pilots produced no measurable P&L impact, Gartner expects over 40% of agentic projects canceled by the end of 2027, and BCG went looking through logistics in January and found exactly 1% of shippers with AI embedded in core operations: 1-in-100.
None of those companies were stupid. Their pilots met freight. The shipper SOP that lives in one planner’s head, the portal that’s never heard of an API, the commodity exception the model had never once seen. Clean loads are easy. The ugly ones, which are most of them, take domain depth and hands.
I’ve sat across from leadership teams living inside those statistics, eight months after a full-automation pitch went nowhere, and their skepticism is earned. The vendors now complaining about AI fatigue built the fatigue themselves.
The Part You Can’t Rent
Back to my rep for a second. All that unique knowledge, the carriers who answer, the ones who ghost, the receiver that backs up on Mondays. Where did any of it live?
Her head. Nowhere else.
The day she left, 10 years of pattern recognition walked out through the parking lot, and my P&L never even registered the loss because no line item existed for it.
A model can’t fix that. A frontier model wakes up every session with amnesia and will never know your carriers better tomorrow than it did today. An execution layer writes everything down. Every negotiation, booked load, and compliance check Ellie runs becomes a shared structure in a Carrier Context Graph the brokerage owns outright, the same operating-layer advantage MIT Technology Review says the whole category is converging on.
Rent the intelligence. Own the memory.
8 Weeks
All of this would matter eventually. Freight decided it matters now. June’s Logistics Managers’ Index scored transportation capacity at 28.4, miles under the neutral 50, with enforcement pushing carriers out for good. ACT Research also reads 2026 as structural tightening rather than a demand pop.
The old answer to a tightening board was a hiring class. Those seats cost more this cycle and cover less.
Montgomery v. Caribe stacked courtroom risk on top, 9-0, and now every carrier you select is a liability decision, with a manual vetting process sitting there as a plaintiff’s exhibit.
Peak season lands in about eight weeks. A pilot scoped in Q4 delivers in Q2, maybe. The window inside this squeeze belongs to the brokerages already executing at machine speed when the boards fill, and it doesn’t hold for a proof of concept.
What I Built Instead
I got the chance to fix my own mistake. Most operators don’t, so I’ll keep this part short and sweet.
Ellie runs in the browser on top of the TMS, portals, and load boards your floor already uses. She sources, reaches carriers over email, SMS, and voice at once, negotiates inside your guardrails, verifies MC and DOT before anything books, and pauses for your rep’s approval. SOC 2 Type II, no integration project.
In live accounts, reps made her their default and now book about three of every four loads through her. They didn’t get automated out of a job. They got promoted into running a workforce.
So book a demo and bring a real lane and the portal your team actually lives in. We’ll source it while you watch, no sandbox.
Better yet, skip the form and email me, robby@tryenvoy.ai, with one number: the percentage of your freight booking through a tool today. If it’s zero, good. That’s an honest starting line. Come talk to us before the boards fill.
My best rep never got an Ellie. Yours can have one by peak season.


