❌

Reading view

Supply chains detect fast, act slow: How AI agents fix it

Supply chain disruption cost businesses about $184 billion in 2025, according to the J.S. Held Global Risk Report, and most of that bill still buys faster detection, not faster action.

That figure is usually treated as weather (i.e. storms happen, costs follow.) Treated as a product specification instead, it highlights an operating model that can spot a problem hours or days earlier than it used to, and still cannot move until a person has opened a ticket, convened a call, and re-entered the same data into three systems.

Visibility platforms, control towers, risk scores, digital twins, and exception dashboards have defined the last decade of AI in the supply chain. That decade has been very good at collapsing the time between an event and awareness of it, but it has been far less good at collapsing the time between awareness and a commercial act.

Detection is a ‘solved-enough’ problem

Ask a chief supply chain officer where the AI budget went and the answer tends to follow a familiar list: demand sensing, ETA prediction, supplier risk scoring, inventory optimisation, and lane analytics. These tools work. Forecast error comes down. A vessel delay is flagged before the container misses the cut-off. A second-tier fab outage shows up on a heat map instead of in a customer email.

None of that accounts for the $184 billion. The bill is the interval after the flag: expedite or wait; split the order or accept the miss; retender the lane or pay the spot rate; consolidate two half-empty movements or ship both; swap ocean for air on the SKUs that actually justify the premium. These are bounded, repeatable decisions that sit inside policy, contract, and inventory limits the company already set—and they still queue behind a human inbox.

Surveys keep describing the same lag in different language. A 2026 Knosc survey of mid-market manufacturers and distributors found that supply-chain teams spend 28 percent of their working time responding to disruptions, most of it investigating what happened rather than changing what happens next.

Logistics executives still rank AI as a strategic priority (Capgemini’s 2025 research put an AI-driven “new-gen” supply chain among the top three technology trends for 70 percent of large-company executives) and then report that measurable financial impact remains rare. Gartner found in 2025 that only 23 percent of supply-chain organisations even have a formal AI strategy. The shortfall is not a shortage of models, but a shortage of authority granted to software.

The ticket is the product

Most current deployments are built around the ticket. The model produces a recommendation, the recommendation becomes an alert, the alert becomes a work item, and the work item waits for a planner already occupied with other work items. By the time the planner acts, the option set has narrowed—the alternative carrier’s capacity is gone, the consolidation window has closed, and the supplier’s next production slot is allocated.

That workflow is not a temporary step on the way to autonomy but the product companies bought. Vendors sold insight because insight is easy to demonstrate and easy to govern; action touches money, contracts, service levels, and blame. So the industry automated the part of the job that does not require a signature. FourKites and ABI Research reported in 2025 that only 27 percent of organisations allow AI to take autonomous action, while 52 percent confine it to decision support.

Adding another dashboard to a delayed shipment rarely moves EBITDA as a result. The decision cycle has not changed; it has only been decorated.

Bounded action as the next model

The firms set to take share are not the ones with the tidiest control tower but the ones that pre-authorise a narrow class of moves and let agents execute them while the exception is still cheap.

Retender a lane when the contracted carrier’s ETA slips beyond a threshold and a qualified alternate sits inside the approved rate band. Consolidate outbound waves when fill rates and cut-off times make a combined movement cheaper than two. Swap mode on a defined SKU set when the cost of air is lower than the cost of a missed retail window. Reallocate safety stock across two distribution centres when a forecast miss and a transport constraint line up.

None of that requires a strategy offsite. Each can be written as: if these conditions, then this action, within this spend cap, with this audit trail, and a human only if the case falls outside the fence. That is not a “lights-out” supply chain—it is the same discipline manufacturers already apply to machine control, where the agent may act inside the interlock and escalates outside it. The difference here is commercial rather than physical: the interlock is a policy object – category, supplier tier, mode, dollar limit, and service class – not a PLC.

Three conditions for real change

First, decisions have to be written as policies, not tribal knowledge. If the only place “we will pay air on A-items after 48 hours of ocean slip” lives is in a planner’s head, no agent can execute it. The work of the next two years is less model training than decision design: which moves are reversible, which are capped, and which suppliers and modes are pre-cleared.

Second, execution systems have to accept machine-initiated transactions. An agent that can draft an RFQ but cannot post it is still a detection tool. TMS, WMS, sourcing suites, and carrier APIs need to treat a bounded agent the way they treat a junior buyer with a spend limit—authenticated, logged, and reversible.

Third, accountability has to move with the action. If a retender inside policy goes wrong, the post-mortem should inspect the policy, the data, and the fence, not hunt for the person who “should have checked”. Until that cultural change happens, every agent will be designed to wait, because waiting is how careers survive.

The competitive split

For a while, both models will look alike on a slide—both will have AI, and both will have a control tower. The difference will show up in cycle time from detection to commercial act, and then in service and cost.

Companies that keep buying detection will know about the storm earlier. Companies that authorise bounded action will already have retendered the lane, consolidated the wave, and moved the A-items before the incident call is booked.

Disruption is not going away. Lead times in critical components, mode volatility, and multi-tier opacity are structural features of the network. What remains optional is whether the response waits for a human to open a queue. The product that created the lag was insight without authority. The product that ends it is an agent allowed to spend a little money, inside a fence, before anyone is free to look.

See also: JD.com expands physical AI in logistics with 3 million robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Supply chains detect fast, act slow: How AI agents fix it appeared first on AI News.

  •  

Motional and MIT AI explains self-driving car decisions

Motional and MIT researchers have built a system that lets self-driving cars explain their decisions in real-time, tackling the black-box problem in autonomous vehicle AI.

The work, published in Nature, comes from a team at Motional that includes CEO Laura Major, working alongside researchers from MIT’s Computer Science and Artificial Intelligence Laboratory. Their proposed method, called the Concept-Wrapper Network or CW-Net, aims to translate the internal calculations of a self-driving system’s neural network into concepts a human can actually read.

If a current self-driving car brakes hard on a clear road with no obvious hazard in sight, neither the driver nor a passenger has any way of knowing why. Modern self-driving systems increasingly rely on neural networks trained on large volumes of driving data. Those networks can perform well, but they don’t expose their reasoning, which is why engineers describe them as black boxes.

Translating neural network logic into human concepts

CW-Net works by converting a self-driving system’s internal logic into concepts such as “Approaching Stopped Vehicle” or “Close to Cyclist.” These could, according to Motional, appear on a dashboard showing which concepts are influencing the vehicle’s driving decisions as they happen.

The system is designed so the explanations aren’t generated after the fact as a guess at what the network might have been doing. Instead, the vehicle’s final decision-making system takes action based directly on these human-interpretable concepts, so a braking event traces back to a specific concept that triggered it. Motional describes this as causally faithful, distinguishing it from approaches that generate natural-language explanations, which can read as plausible without necessarily being accurate.

Laura Major frames the case for this kind of interpretability against the alternative of relying purely on end-to-end deep learning to handle driving decisions.

“The general end-to-end only approach can get to a really good 80-90 percent – maybe even 95 percent – solution, but that’s not good enough to remove a driver or to earn the trust of cities, communities, and customers,” she said.

Testing explainable AI for self-driving cars around Las Vegas

Explainable AI research has largely stayed confined to computer simulations in lab settings, according to Motional. The Motional and MIT team instead deployed CW-Net on an autonomous vehicle with an experienced safety operator in the driver’s seat, collecting data on a private test track and on public roads around Las Vegas.

The team used an earlier experimental version of its deep-learning-based planning system, described as showing competitive performance but with notable shortcomings that CW-Net could help surface. Two incidents from the testing illustrate what the system caught.

In one, the autonomous vehicle repeatedly stopped near a traffic cone, and the vehicle operator assumed the cone itself was triggering the behaviour. Researchers removed the cone and the car stopped anyway. CW-Net’s display showed the actual cause: the experimental planning system was hallucinating a stopped vehicle ahead, a pattern traced back to its training data. That explanation let the researchers understand, predict, and resolve the issue.

A second test involved a cyclist. The autonomous vehicle detected and stopped for the cyclist as expected, but CW-Net revealed that the experimental planning system wasn’t actually basing its decision on the cyclist’s presence. The safety driver responded by exercising more caution around cyclists after noticing this. Follow-up analysis confirmed that caution was warranted, because the vehicle’s braking in that case came from a safety backup system rather than the experimental deep-learning-based planner.

Performance held steady against explainability

Adding layers of explainability to an AI system carries a known cost in speed and performance, and Motional acknowledges that risk. However, when researchers benchmarked CW-Net against leading autonomous driving algorithms, the difference in driving capability came in at less than one percent.

The Las Vegas incidents show why that trade-off matters operationally rather than just academically. A safety driver who can see that a stop is caused by a hallucinated vehicle, or that a backup system rather than the primary planner is responsible for a manoeuvre, can respond and report with more precision than one working from behaviour alone.

That visibility feeds directly into how quickly an engineering team can diagnose a system, and how confidently a safety operator can distinguish between an intended behaviour and a fault.

Motional connects the CW-Net work to broader pressure on autonomous vehicle operators as the technology extends into new markets and jurisdictions. Regulators are naturally asking for more transparency about how AI systems reach their decisions, and it expects tools like CW-Net could move from research projects toward a baseline requirement.

Beyond passenger vehicles, autonomous drones and even robotic surgery are cited as other safety-critical domains where operators and developers will need ways to understand a system’s capabilities, limitations, and unexpected behaviours.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: MIT AI forecasts extreme weather without historical data

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Motional and MIT AI explains self-driving car decisions appeared first on AI News.

  •  

Microsoft open-source toolkit secures AI agents at runtime

A new open-source toolkit from Microsoft focuses on runtime security to force strict governance onto enterprise AI agents. The release tackles a growing anxiety: autonomous language models are now executing code and hitting corporate networks way faster than traditional policy controls can keep up.

AI integration used to mean conversational interfaces and advisory copilots. Those systems had read-only access to specific datasets, keeping humans strictly in the execution loop. Organisations are currently deploying agentic frameworks that take independent action, wiring these models directly into internal application programming interfaces, cloud storage repositories, and continuous integration pipelines.

When an autonomous agent can read an email, decide to write a script, and push that script to a server, stricter governance is vital. Static code analysis and pre-deployment vulnerability scanning just can’t handle the non-deterministic nature of large language models. One prompt injection attack (or even a basic hallucination) could send an agent to overwrite a database or pull out customer records.

Microsoft’s new toolkit looks at runtime security instead, providing a way to monitor, evaluate, and block actions at the moment the model tries to execute them. It beats relying on prior training or static parameter checks.

Intercepting the tool-calling layer in real time

Looking at the mechanics of agentic tool calling shows how this works. When an enterprise AI agent has to step outside its core neural network to do something like query an inventory system, it generates a command to hit an external tool.

Microsoft’s framework drops a policy enforcement engine right between the language model and the broader corporate network. Every time the agent tries to trigger an outside function, the toolkit grabs the request and checks the intended action against a central set of governance rules. If the action breaks policy (e.g. an agent authorised only to read inventory data tries to fire off a purchase order) the toolkit blocks the API call and logs the event so a human can review it.

Security teams get a verifiable, auditable trail of every single autonomous decision. Developers also win here; they can build complex multi-agent systems without having to hardcode security protocols into every individual model prompt. Security policies get decoupled from the core application logic entirely and are managed at the infrastructure level.

Most legacy systems were never built to talk to non-deterministic software. An old mainframe database or a customised enterprise resource planning suite doesn’t have native defenses against a machine learning model shooting over malformed requests. Microsoft’s toolkit steps in as a protective translation layer. Even if an underlying language model gets compromised by external inputs; the system’s perimeter holds.

Security leaders might wonder why Microsoft decided to release this runtime toolkit under an open-source license. It comes down to how modern software supply chains actually work.

Developers are currently rushing to build autonomous workflows using a massive mix of open-source libraries, frameworks, and third-party models. If Microsoft locked this runtime security feature to its proprietary platforms, development teams would probably just bypass it for faster, unvetted workarounds to hit their deadlines.

Pushing the toolkit out openly means security and governance controls can fit into any technology stack. It doesn’t matter if an organisation runs local open-weight models, leans on competitors like Anthropic, or deploys hybrid architectures.

Setting up an open standard for AI agent security also lets the wider cybersecurity community chip in. Security vendors can stack commercial dashboards and incident response integrations on top of this open foundation, which speeds up the maturity of the whole ecosystem. For businesses, they avoid vendor lock-in but still get a universally scrutinised security baseline.

The next phase of enterprise AI governance

Enterprise governance doesn’t just stop at security; it hits financial and operational oversight too. Autonomous agents run in a continuous loop of reasoning and execution, burning API tokens at every step. Startups and enterprises are already seeing token costs explode when they deploy agentic systems.

Without runtime governance, an agent tasked with looking up a market trend might decide to hit an expensive proprietary database thousands of times before it finishes. Left alone, a badly configured agent caught in a recursive loop can rack up massive cloud computing bills in a few hours.

The runtime toolkit gives teams a way to slap hard limits on token consumption and API call frequency. By setting boundaries on exactly how many actions an agent can take within a specific timeframe, forecasting computing costs gets much easier. It also stops runaway processes from eating up system resources.

A runtime governance layer hands over the quantitative metrics and control mechanisms needed to meet compliance mandates. The days of just trusting model providers to filter out bad outputs are ending. System safety now falls on the infrastructure that actually executes the models’ decisions

Getting a mature governance program off the ground is going to demand tight collaboration between development operations, legal, and security teams. Language models are only scaling up in capability, and the organisations putting strict runtime controls in place today are the only ones who will be equipped to handle the autonomous workflows of tomorrow.

See also: As AI agents take on more tasks, governance becomes a priority

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft open-source toolkit secures AI agents at runtime appeared first on AI News.

  •  
❌