Planners have gone from asking AI questions to letting it write to the plan. The next step is writing down how the team plans so agents can follow it.
Q4 2026 Edition
In-Stock
The unhedged truth about retail planning.
Planners have gone from asking AI questions to letting it write to the plan. The next step is writing down how the team plans so agents can follow it.
A few weeks ago, a planner at a children's apparel brand asked Tooli, our AI planner, to build their department's FY26 open-to-buy deck: plan versus budget versus last year, by channel and subclass, in the team's own format. Tooli built it and checked every number against the plan.
Then the planner asked a different kind of question: "What's the prompt so that we can ask you to pull this again?" By the end of the day, other planners on the team were using that one line to build decks for their own departments.
That planner went beyond building a deck. They wrote down how the work gets done. That step separates teams that use AI from teams that hand work to it, and it's what this is about.
86% of our customers now have planners working with AI, either in Tooli or through our MCP server. The growth is coming from more people using it, not from a handful of power users. One in five outputs is a finished deliverable: a spreadsheet, deck or saved view that's ready to share.
What planners hand off tells you a lot. Weekly reporting, buy checks before the PO goes out, risk scans for the next quarter and store setup are the jobs that used to eat Monday morning, and they're the first ones going to AI.
Across those conversations I see three stages. Most teams are in the first, a growing share are in the second, and nobody is fully in the third yet.

89% of AI sessions in Toolio are read-only. Planners analyze, reconcile and report, and nothing gets written back to the plan. Some people read that as caution. I think it's the right order of operations.
The most common question isn't "what will sell." It's "why don't these two numbers match." Planners use an agent to understand the plan before they trust it with the plan, because nobody commits a season of receipts on a number they can't tie out.
A buyer at a multi-brand streetwear retailer asked for a size-level buy sheet for a key brand's spring buy. Tooli built it, then flagged that the fixed size packs were 38% small and medium while those sizes made up 62% of demand. The mismatch surfaced before the PO was confirmed.
11% of sessions now write changes straight back into the plan. That number is small, and it's the one I watch most closely.
A planner at a $1B+ apparel brand opening a new store asked Tooli to copy an existing store's min and max by size across every presentation profile. Tooli checked each profile, spelled out what would change, and wrote the rows back into Toolio, so the store setup happened in one conversation.
The planner still directs and approves, and the trust comes from checking the work.
At a national footwear brand, a planner asked why one outlet got so many units. Before recommending anything, Tooli reproduced the actual transfers to within a few units at 83% of stores, traced the cause to the retrend window, and modeled the options.
Half an hour later the planner signed off on a new safety-stock policy: "I'm convinced, thanks Tooli."
The part I find most interesting is what planners do after they get an answer: they save the procedure. A returns-by-channel analysis at an outdoor brand became a saved skill, and the planner re-ran it the next week with one line.
An over- and understock scan at an intimates brand became a standing tracker the whole team watches. Planners are starting to write down how they work. That's where stage three begins.
In stage three, agents do most of the work and planners decide how it gets done. They set the guardrails, define when work starts and who signs off, and spend their time on exceptions and judgment calls.
The planner becomes the reviewer, and the designer of the system. Peter covers how agent orchestration works in his piece in this issue.
For planning leaders the lesson is simpler: an agent can't follow a process your team has never written down. Getting to stage three takes guardrails, but first it takes something harder.
Organizations have to get explicit about how they plan, and that comes down to five things.
Every decision from target setting to in-season review, with its inputs, rules, dependencies and downstream impacts. In most teams this lives in a few people's heads.
When the financial plan locks, when the assortment gets built, when buys commit and when receipts flow, all on your own lead times. The calendar becomes the clock agents work to.
Alongside it you define the business events that should start off-cycle work, like a demand shift, a late receipt or an upstream plan getting approved.
Who owns which decision, and what an agent may do without asking: margin floors, open-to-buy limits, cover targets, markdown budgets. Inside those limits the agent acts on its own, and outside them it asks first. An event can start an agent working, but it doesn't give the agent permission to write.
What passes to the next step, what happens when a check fails, and who hears about it.
What was decided and how it turned out, so the system gets better each season.
This isn’t new for good planning teams. The best teams always had a calendar, a process, and people who knew when to escalate. The difference now is that it has to be written down clearly enough for an agent to follow.
In Q2 I called this the early innings. In Q3 I said the planner's job was being rewritten. Both still hold. What's changed is the pace.
No brand or retailer we work with has handed the controls to agents, and the models still have to earn that trust, one checked answer and one reversible edit at a time.
Every team will move at its own pace, and some decisions will move faster than others. Reporting will get there long before the buy does.
A year ago most planners were curious about AI. Today they're reconciling numbers, building decks, writing to live plans and saving their own routines. The step from there to stage three is shorter than it looks.
Pick one planning process and write it down end to end. Put your calendar and lead times somewhere an agent can read them. Then set the tolerances for one decision: what an agent can change on its own, and what needs your sign-off.
Those are the inputs that let agents do the work and let your team move up into review. It's still early, and it's moving faster.
Outerwear is tracking 12% under plan, and three AI agents have already acted on it. The allocation agent pulled cover from slow stores. The replenishment agent reordered for the stores still selling through. The markdown agent pulled promotions forward, discounting the same styles replenishment had just bought more of.
Each decision was right on its own terms. None was reconciled against open-to-buy, margin targets or each other. Together, they pushed the category past its budget and gave away margin no one approved.
This is the fragmentation planners have lived with for years, just automated.
Faster execution of uncoordinated decisions is simply a quicker way to drift off plan. The next leap is bigger than a smarter individual agent. It's an orchestrator that runs the whole planning chain, calls on specialist agents for each part of it, and holds every one of them to a single plan.
Agentic AI has arrived in retail planning, one task at a time. A replenishment agent that reorders against live demand. A markdown agent that tunes price cadence. An allocation agent that rebalances cover across stores. Each is genuinely useful.
But a merchandise planning process is more than a collection of tasks. It is a chain of decisions, from the financial plan to the assortment to the buy to the units in each store.
The next leap is bigger than a smarter individual agent. It’s an orchestrator that runs the whole chain, calls on specialist agents for each part of it, and holds every one of them to a single plan.
The better model looks like a well-run planning team: an orchestration agent at the center, surrounded by specialists.
An exception agent watches the business continuously against the tolerance bands the merchandise plan defines. When outerwear breaks its band, it hands the orchestrator a scoped problem: which category, how far off, which guardrails are at risk.
What counts as an exception is not a generic threshold. It is a deviation from your plan, at the tolerances your planners set.

The orchestrator then decomposes the response. The MFP agent reforecasts the category and quantifies the gap. The assortment agent evaluates which choices to cut, deepen, or chase. The buying agent checks open-to-buy and vendor lead times. The allocation and replenishment agents model how inventory should flow to the stores that can sell it.
The orchestrator assembles their work into one coherent recommendation instead of five competing ones.
What makes this model safe is that the merchandise plan stops being a scorecard and becomes the contract every agent operates within.
Margin floors, open-to-buy limits, cover targets, and markdown budgets define what agents can do on their own and what requires a human decision. A replenishment adjustment inside approved OTB executes automatically. A chase buy that would push a department over budget routes to the planner with the trade-offs already modeled.
The guardrails live in one place, and every agent reads from it.
Most agentic approaches assume a textbook planning process, and no retailer runs that way. A department store may lock the financial plan before assortment begins and require GMM sign-off on large buys. A DTC brand may iterate between plan and range for weeks and let planners chase freely within OTB.
The specialist agents can be standardized. The orchestration cannot.
That makes adaptability the orchestrator's defining feature. Retailers should be able to configure their own stages, approval gates, and escalation paths, and change them in days when the business adds a channel or tightens approvals for a tough season.
The orchestrator follows your process rather than asking you to adopt its own.
Orchestration only works if every agent sees the same truth. When MFP, assortment, and allocation and replenishment live in separate systems, five agents means five versions of the plan, and the guardrails drift out of sync the moment one system updates before another.
On an integrated platform with one data model, a reforecast is instantly visible to the assortment and buying agents, a chase carries its OTB and allocation implications at once, and an approved recommendation becomes plan data directly.
Retailers evaluating agentic AI should ask less about any single agent and more about whether the foundation beneath them is connected.
None of this removes the planner, but it does change the job. Instead of manually carrying decisions from one system to the next, the planner designs the workflow, sets the guardrails, and judges the recommendations that reach their desk. They supervise a team of agents the way a planning director supervises a team of analysts, focusing their time on the calls that genuinely need judgment.
The retailers pulling ahead will be the ones whose agents work together, inside a plan they trust, following a process built around how their business runs.
For a couple of years now, tariffs showed up in planning conversations the way weather shows up in a forecast. Something to react to, plan around, wait out. That's not how planners are talking about them anymore. This quarter, tariffs weren't just a topic that came up, they were the condition underneath every other topic.
One home goods leader summed up the mood without meaning to, opening a call with: "just living the life over here in the tariff world we live in."
That's a shift as opposed to a storm to weather; a world to operate in.
Tariffs have moved from finance's problem to everyone's daily operations. A jewelry supplier told us they had to cut a call short to "run to our weekly tariff meeting;” a recurring line on the calendar that didn't exist a year ago. An apparel brand opened a session with an apology:
"Everyone, my apologies — today's all hands on deck with tariffs."
The changes are landing mid-order instead of a hypothetical or safely in the future. One brand described product sitting at the port under a hold. Another said they can't hold pricing even ten weeks out, because a new tariff could push costs up 50% before the goods arrive.
India made that concrete for several customers. With rates jumping to 50 percent, a bedding brand described trying to reflect the change and hitting a wall: "even if I go into every SKU and press update prices and then refresh and then do it again, it doesn't necessarily capture it all." An outdoor brand asked whether they could plan against the cost at all:
"We're expecting a 50% tariff in India to come in this future season. Would we be able to model different costs on a forward looking basis?"
That question “Can I plan against a cost I know is coming but don't have yet” came up again and again.
Here's the structural problem: Tariffs corrupted the historical record planners rely on to forecast.
When costs change this often, order cost and weighted cost drift apart, and last year is no longer a clean comparison. A lifestyle brand put it plainly: "over time, with price changes and tariffs, that is going to give me misleading data." Others described cost data that simply hadn't caught up: POs where the duty rate, freight, and landed cost no longer matched what was in the system.
Margin stopped being a number you check after the fact and became the number you plan around. Teams told us they're now planning to landed margin as a primary metric; that it's "the one that we know is a moving target." A luggage brand described watching healthy margins evaporate on arrival: "when it lands in the US, those margins are being defeated by a good four or five quads because of the way we have to bring it in, and then tariffs." A footwear brand has started breaking unit FOB, duty, tariff, and freight into separate line items in their product feed, costs that never used to be tracked apart.

The other thing customers aren't doing is waiting for next season to react. They're changing where and how they make things right now, inside active buys.
An outdoor apparel brand said it flatly: "we've changed factories in some cases because of tariffs." An intimates brand described diversifying into new countries to soften the impact, and the ripple that creates: "we've had to create a new SKU." A wearables company is actively moving development out of China to save money. An apparel brand shifted production to Mexico when the duty-free math changed.
Each of these moves is sensible on its own. Together they create SKU proliferation, broken cost histories, and plans built on assumptions that are no longer true. Instead of resourcing a product,you’re forced to re-plan it with new cost, lead time, and country of origin often before the old version has fully cycled through.
The pattern turns into something bigger than a cost line. Tariff work is crowding out everything else.
Across a striking number of conversations, customers named tariffs as the reason projects stalled, training slipped, or decisions waited. "There's a lot of higher priority things right now, immediate fires to fight," one apparel leader said, pointing to tariffs and de minimis changes. A sporting goods brand was candid: "we had to worry about these tariffs and we changed a lot of our pricing, and I'll be honest, I just haven't had a lot of time to focus on [much else]”
Underneath that is real caution about spending. We heard about cash being conserved, budgets under fresh scrutiny, and lean teams asking hard questions about what's worth the maintenance. A consulting partner described the market behavior in three words: "companies hoarding cash."
For all the noise on the cost side, the demand side is changing as well. Traffic is softening, and customers are tying it directly to the broader climate. A beachwear retailer told us they'd been "a little bit overly confident" on sales, and that with everything going on, "we've definitely seen traffic coming down and not being as good as it used to be."
At the same time, a lot of brands are sitting on too much inventory. Call it a hangover from years of over-ordering against uncertainty. One outdoor brand said half their units are overstock, "tying up cash flow and hindering new product development." A tactical brand described $80 million in inventory against a $65 million target. When cash is already tight and costs are climbing, that overhang hurts twice.
Customers are naming the big picture out loud in ways they didn't used to. One planning leader, describing what an accurate forecast now has to account for, listed "tariffs and wars in Iran and inflation and all that other good stuff." That's the frame planners are working inside.
What ties all of this together is that the ground keeps moving under the plan: costs, sources, timelines, and now the reliability of the history itself. The teams who feel like they're managing tell us something close to what they told us earlier this year: the goal isn't to predict the next tariff. It's to re-plan fast when it lands.
But there's a new layer on top of that. It used to be enough to move quickly. This quarter, customers are asking a harder question first: “Can I even trust what I'm looking at?” For a lot of teams, right now, answering that is the whole job.
A year ago, boards were mandating AI. Now they want returns. We asked our advisory and SI partners what their clients can actually show for the last twelve months, and what's shifting underneath. Here's what came back.
When partners separate the AI built into planning workflows from the AI rolled out across the enterprise, two very different receipts come back.
Here's what clients can put on a slide, and nearly all of it came from inside the planning workflow:
Accuracy improved in specific categories where the data was clean enough for the model to work.
The same team covers more SKUs, or fewer people produce the same output. One partner called this the most honest ROI story most retailers can tell right now.
They built internal credibility, and boards are starting to accept that as a legitimate answer for year one.
Here's what's being written off; most of it sits outside planning, in enterprise-wide AI programs:
Governance frameworks, data catalogs and center-of-excellence buildouts used up budget and headcount without connecting to a decision or a dollar.
Adoption numbers were strong and business impact was weak. Usage isn't value, and boards have figured that out.
One write-off does land squarely in planning: software provider POCs that ran beautifully on the provider's data and stalled the moment the retailer's real data arrived. The lesson for evaluation is simple. If a proof of concept hasn't run on your own data, it hasn't proven anything yet.
"The delta between demo performance and production performance is the embarrassing number no one is publishing."
Partner estimates of how many 2026 planning pilots will be in production heading into 2027 ranged from about half to almost none. The ones that make it follow a similar path: they run in silent mode next to the existing process for months before anyone switches over.
The pattern across the year: the ROI clients can defend came from AI embedded in a specific workflow and measured on a specific outcome. The losses came from AI deployed in general and measured on activity.
Retailers don't demand explainability everywhere to the same degree. The pressure is highest in allocation and replenishment. When a store gets shorted or over-shipped, someone's performance review is attached to it.
Planners won't consistently act on a recommendation they can't explain to their merchant or defend on the Monday morning call. When the reasoning is hidden, override rates climb above 60%, which defeats the purpose.
Demand forecasting at the SKU-store level gets more slack. Planners have always accepted that the math is complex. If the number is directionally right and the override is easy, most don't ask why. That's starting to change, and more clients are asking how the model reached a future forecast.
For now, though, retailers still accept a black box when the output looks familiar and overriding it is cheap. They reject it when a person takes the blame for the result.
Twelve months ago, some $1B+ brands and retailers wanted to build planning AI themselves, on a foundation model or an in-house one. That conversation has retreated to buy.
Partners describe the same arc every time: the first six months go great, then performance and architecture problems show up, and that's where it fails.
What's taking its place is more modest. Retailers implement the core planning capabilities and build around them with process optimization, bolt-ons and extensions.
"The starry-eyed view of build everything and replace what I'm paying license for has been debunked."
Many retailers leaving their current planning tool are doing it because they went fast the first time and ported over what they already had. Partners are blunt about hard cutovers: one called them a great way to crater the business.
Aggressive go-live promises are still in the market, and partners are still warning clients off them. Data readiness remains the long pole.
The pitch lines that break most often haven't changed much:
On the peak freeze, the most useful advice this quarter was a staged go-live. Go live on a category or region that isn't exposed to peak, run parallel on the high-stakes business, and use those six weeks for the hardest hypercare work. You keep the timeline without taking the full risk.

AI councils are a new approval gate that sits outside the business sponsor's authority. Unlike procurement or legal, they don't have a standard playbook yet. Each council has its own risk thresholds, questions and timelines. The result is often 6–12 weeks of unplanned delay in the middle of a program the business thought was already decided.
Retailers are also choosing the SI and the software vendor on separate tracks. That's deliberate: prior ERP cycles taught them that a single-vendor selection creates leverage problems later. But it leaves an architecture gap that nobody formally owns, and that's where scope disputes and change orders come from.
Partners had one fix in common: evaluate the system, not the tool. Three planning decisions made in three separate RFPs will produce three technically correct answers, plus an integration and sequencing problem no one budgeted for.
"The vendors are selected; the architecture between them isn't."
Most forecasting models were trained on history that assumed a fairly stable customer mix. When demand splits between a resilient high end and a softening value shopper, the aggregate forecast still looks reasonable while being wrong in both directions.
High-end categories are underforecast and value categories are overforecast. Inventory lands in the wrong place, and markdown pressure builds before any dashboard flags a problem.
Allocation logic is the next thing to fail. Store clusters and size curves were built on a single demand profile. A cluster that made sense twelve months ago may now hold stores with very different demand. The algorithm is optimizing correctly against the wrong segmentation, and you usually don't see it until units are sitting in the wrong doors.
Asked what most retail software providers aren't ready for, partners pointed to continuous planning. Leading retailers will move from weekly and monthly cycles to rolling, event-triggered replanning: the system updates when a signal crosses a threshold, not because it's Tuesday.
Most providers are building better interfaces for the existing cadence instead of rebuilding for a world where the cadence goes away.
The harder gap is organizational. Continuous planning raises questions about who owns the exception and who has authority to act without a meeting. Almost no one is helping retailers design that. Underneath all of it is product data.
Partners expect it to be the main constraint on AI planning value over the next two years, because an agent can only act on what it can read.
Past Editions

See Toolio in Action
The best way to understand what Toolio could do for your team is to start a conversation.