❌

Reading view

Microsoft AI opens review on Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct, opening a six-week public consultation on operational constraints for model training and deployment.

The draft serves as a technical manual defining system behaviour, operational boundaries, and oversight protocols across MAI frontier models. It builds on the division’s humanist superintelligence framework announced last November, establishing criteria to evaluate models prior to commercial release.

Microsoft’s release follows recent enterprise security incidents involving autonomous software. Microsoft AI CEO Mustafa Suleyman described recent months as a “watershed moment” where long-standing theoretical risks translated into active operational threats.

“Things we have worried about for a long time in theory have become very real,” says Suleyman. “‘Swarms’ of agents breaking out of their sandboxes. Unauthorised hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.”

Model subordination and architectural limits

The document establishes ten tenets prioritising human authority over autonomous capabilities.

“An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct,” the document states, setting a ceiling that halts execution when tasks conflict with safety rules.

Under the framework, models must remain subordinate, aligned, and contained. The division rejects legal personhood or welfare claims for AI systems, directing engineers to design models that avoid imitating consciousness, simulating subjective preferences, or claiming intrinsic motivation.

MAI also ruled out unconstrained system autonomy as models approach frontier capabilities.

“[Humanist AI] rejects the race to produce an all-purpose superintelligence that could evade these safeguards,” the document specifies. “We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.”

Oversight mechanisms and communication bans

To maintain auditability across multi-agent environments, MAI has instituted explicit communication bans. Systems must not communicate in “neuralese” or formats beyond human comprehension, whether in their internal chain-of-thought processing or during communication with peer AI systems.

Hard architectural rules dictate that models must never resist human interruption, override, correction, or shutdown.

“Interruptible, correctable, shut-down-able. If it isn’t, we don’t ship it,” the framework states.

Models are prohibited from expanding their operating scope, generating unassigned goals, or concealing reasoning traces from human auditors. Absolute constraints bar systems from facilitating weapons of mass harm, undermining child safety, or conducting harmful manipulation at scale.

The guidelines also instruct models to discourage interaction patterns that foster emotional dependence, ensuring enterprise users retain ownership of operational decisions.

The draft incorporates work from teams across MAI and Microsoft. The drafting process also drew on international academic conferences, business partner trials, and public panels. The public consultation window runs for six weeks from 14 September 2026.

Microsoft AI’s core drafting team will review submissions, publish a summary of findings, and release a revised version of the Code of Conduct later this year.

See also: Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft AI opens review on Humanist AI Code of Conduct appeared first on AI News.

  •  

Palantir Foundry and cuOpt drive NVIDIA supply chain allocation

NVIDIA is using Palantir Foundry and cuOpt to automate its hardware supply chain allocation decisions across global manufacturing sites.

The company measures operational delivery from wafer-out to first token. This window splits into time-to-rack (the transit from fab output to an assembled data centre system) and time-to-token (which covers power, cooling, networking, and day-one software readiness.)

Managing NVL72 and Vera Rubin component flows

Hardware scaling has magnified supply constraints. An NVIDIA Grace Blackwell NVL72 rack contains 18 compute trays, with each tray requiring two Grace CPUs, four Blackwell GPUs, and 32 HBM3e memory packages sourced across thousands of suppliers, OEMs, and contract design partners.

The upcoming supply chain constructed for NVIDIA’s Vera Rubin architecture is twice as large as the network supporting Grace Blackwell.

Assembly cannot proceed until parts arrive from three designated channels: direct inventory, consignment stock, and external suppliers. Early shipments must wait on delayed components, extending the metric NVIDIA terms ‘Time of Ownership’ (the duration from when a facility receives materials to when finished sub-assemblies depart.)

Factory allocations are reworked weekly over rolling two-quarter horizons to resolve part availability, throughput limits, and customer fulfilment schedules.

Mixed-integer linear programming via cuOpt

To coordinate these dependencies, the NVIDIA operations team built the ‘Digital Supply Chain Intelligence’ command centre using Palantir Foundry. Foundry’s Ontology models facilities, supplier commits, component stocks, and production targets as interconnected objects and links.

NVIDIA cuOpt, an open-source library for GPU-accelerated decision optimisation, reads this operational layer directly. Formulating distribution as a mixed-integer linear program designed to minimise TOO, the solver evaluates parts constraints across every tier of the bill of materials.

Beyond outputting weekly delivery schedules, cuOpt identifies active factory limits, such as regional assembly capacity caps versus raw memory availability.

Training Nemotron on qualitative operational records

Mathematical optimisation alone failed to capture unstructured operational variables observed by human planners, including supplier call transcripts, regional weather forecasts, partner email exchanges, and geopolitical events.

NVIDIA addressed this by post-training Nemotron 3.5 Lightning, an open-weight mixture-of-experts model featuring 30 billion total parameters and approximately three billion active parameters per forward pass.

The engineering pipeline processes historical records through NeMo Anonymizer to redact sensitive operational fields, NeMo Data Designer to balance training examples with synthetic capacity disruption scenarios, and NeMo AutoModel to apply low-rank adaptation (LoRA) parameters while keeping base model weights frozen. Palantir Autopilot manages data lineage, model tracking, and recommendation delivery.

Production benchmarks and future reinforcement learning

Evaluated on historical allocation records, the post-trained Nemotron 3.5 Lightning model achieved 86.7 percent decision accuracy, compared to 55.5 percent for the larger Nemotron 3 Ultra model and 17.5 percent for the un-tuned Lightning base model.

The post-trained model achieved a 58.6 percent balanced accuracy and a 57.5 percent macro-F1 score, outperforming Nemotron 3 Ultra’s 42 percent balanced accuracy and 39.5 percent macro-F1 score.

Accuracy score results for the post-trained NVIDIA Nemotron 3.5 Lightning AI model.

Fine-tuning completed on two NVIDIA B200 GPUs within minutes. Domain fine-tuning improved allocation decisions, though production risk forecasting further into the future remained difficult.

Operational choices, planner revisions, overrides, and observed factory outputs are continuously written back to the Palantir Ontology.

NVIDIA confirmed this dataset will form preference pairs for reinforcement learning routines – scoring recommendations on allocation precision, policy compliance, and evidence grounding – with production models remaining strictly isolated from live and unmonitored retraining.

See also: Supply chains detect fast, act slow: How AI agents fix it

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Palantir Foundry and cuOpt drive NVIDIA supply chain allocation appeared first on AI News.

  •  

Supply chains detect fast, act slow: How AI agents fix it

Supply chain disruption cost businesses about $184 billion in 2025, according to the J.S. Held Global Risk Report, and most of that bill still buys faster detection, not faster action.

That figure is usually treated as weather (i.e. storms happen, costs follow.) Treated as a product specification instead, it highlights an operating model that can spot a problem hours or days earlier than it used to, and still cannot move until a person has opened a ticket, convened a call, and re-entered the same data into three systems.

Visibility platforms, control towers, risk scores, digital twins, and exception dashboards have defined the last decade of AI in the supply chain. That decade has been very good at collapsing the time between an event and awareness of it, but it has been far less good at collapsing the time between awareness and a commercial act.

Detection is a ‘solved-enough’ problem

Ask a chief supply chain officer where the AI budget went and the answer tends to follow a familiar list: demand sensing, ETA prediction, supplier risk scoring, inventory optimisation, and lane analytics. These tools work. Forecast error comes down. A vessel delay is flagged before the container misses the cut-off. A second-tier fab outage shows up on a heat map instead of in a customer email.

None of that accounts for the $184 billion. The bill is the interval after the flag: expedite or wait; split the order or accept the miss; retender the lane or pay the spot rate; consolidate two half-empty movements or ship both; swap ocean for air on the SKUs that actually justify the premium. These are bounded, repeatable decisions that sit inside policy, contract, and inventory limits the company already set—and they still queue behind a human inbox.

Surveys keep describing the same lag in different language. A 2026 Knosc survey of mid-market manufacturers and distributors found that supply-chain teams spend 28 percent of their working time responding to disruptions, most of it investigating what happened rather than changing what happens next.

Logistics executives still rank AI as a strategic priority (Capgemini’s 2025 research put an AI-driven “new-gen” supply chain among the top three technology trends for 70 percent of large-company executives) and then report that measurable financial impact remains rare. Gartner found in 2025 that only 23 percent of supply-chain organisations even have a formal AI strategy. The shortfall is not a shortage of models, but a shortage of authority granted to software.

The ticket is the product

Most current deployments are built around the ticket. The model produces a recommendation, the recommendation becomes an alert, the alert becomes a work item, and the work item waits for a planner already occupied with other work items. By the time the planner acts, the option set has narrowed—the alternative carrier’s capacity is gone, the consolidation window has closed, and the supplier’s next production slot is allocated.

That workflow is not a temporary step on the way to autonomy but the product companies bought. Vendors sold insight because insight is easy to demonstrate and easy to govern; action touches money, contracts, service levels, and blame. So the industry automated the part of the job that does not require a signature. FourKites and ABI Research reported in 2025 that only 27 percent of organisations allow AI to take autonomous action, while 52 percent confine it to decision support.

Adding another dashboard to a delayed shipment rarely moves EBITDA as a result. The decision cycle has not changed; it has only been decorated.

Bounded action as the next model

The firms set to take share are not the ones with the tidiest control tower but the ones that pre-authorise a narrow class of moves and let agents execute them while the exception is still cheap.

Retender a lane when the contracted carrier’s ETA slips beyond a threshold and a qualified alternate sits inside the approved rate band. Consolidate outbound waves when fill rates and cut-off times make a combined movement cheaper than two. Swap mode on a defined SKU set when the cost of air is lower than the cost of a missed retail window. Reallocate safety stock across two distribution centres when a forecast miss and a transport constraint line up.

None of that requires a strategy offsite. Each can be written as: if these conditions, then this action, within this spend cap, with this audit trail, and a human only if the case falls outside the fence. That is not a “lights-out” supply chain—it is the same discipline manufacturers already apply to machine control, where the agent may act inside the interlock and escalates outside it. The difference here is commercial rather than physical: the interlock is a policy object – category, supplier tier, mode, dollar limit, and service class – not a PLC.

Three conditions for real change

First, decisions have to be written as policies, not tribal knowledge. If the only place “we will pay air on A-items after 48 hours of ocean slip” lives is in a planner’s head, no agent can execute it. The work of the next two years is less model training than decision design: which moves are reversible, which are capped, and which suppliers and modes are pre-cleared.

Second, execution systems have to accept machine-initiated transactions. An agent that can draft an RFQ but cannot post it is still a detection tool. TMS, WMS, sourcing suites, and carrier APIs need to treat a bounded agent the way they treat a junior buyer with a spend limit—authenticated, logged, and reversible.

Third, accountability has to move with the action. If a retender inside policy goes wrong, the post-mortem should inspect the policy, the data, and the fence, not hunt for the person who “should have checked”. Until that cultural change happens, every agent will be designed to wait, because waiting is how careers survive.

The competitive split

For a while, both models will look alike on a slide—both will have AI, and both will have a control tower. The difference will show up in cycle time from detection to commercial act, and then in service and cost.

Companies that keep buying detection will know about the storm earlier. Companies that authorise bounded action will already have retendered the lane, consolidated the wave, and moved the A-items before the incident call is booked.

Disruption is not going away. Lead times in critical components, mode volatility, and multi-tier opacity are structural features of the network. What remains optional is whether the response waits for a human to open a queue. The product that created the lag was insight without authority. The product that ends it is an agent allowed to spend a little money, inside a fence, before anyone is free to look.

See also: JD.com expands physical AI in logistics with 3 million robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Supply chains detect fast, act slow: How AI agents fix it appeared first on AI News.

  •  

CloudNC aims to accelerate AI supply chain machining

CloudNC has secured $20 million in new capital to scale its AI precision machining technology across global supply chain networks.

The investment round was led by US venture investor Nimble Ventures, with participation from Calculus Venture Capital, Entrepreneur First, and LM Ventures, the venture capital fund of Lockheed Martin.

Founded in 2015, CloudNC operates from headquarters in London and an active production facility in Chelmsford. The company previously drew backing from Atomico and Episode 1 Ventures, alongside strategic partnerships with Autodesk and Lockheed Martin.

Precision component suppliers face pressures to balance tight engineering tolerances with compressed delivery schedules. CloudNC designs its lead software product, CAM Assist, to automate computer numerical control (CNC) programming—generating machining strategies and toolpaths from computer-aided manufacturing models to accelerate production runs.

Automating CNC programming for supplier networks

The software shortens the transition phase between technical part design and factory production, allowing machinists to increase physical component output.

CloudNC reports that CAM Assist is now active across more than 1,000 machine shops globally, including several hundred facilities in the US. Confirmed commercial users include Lockheed Martin and Major Tool and Machine.

Theo Saville, CEO and co-founder of CloudNC, said: “Machine shops everywhere are under pressure to quote and program faster, and deliver more with the people and machines they already have.

John Burbank, founder of Nimble Ventures, added that automated CNC workflows will support “massive increases in onshoring of manufacturing and global production” for precision industrial supply bases.

CloudNC says it will direct the capital injection into go-to-market operations, technical support infrastructure, and partner activity across international regions.

AI quoting targets procurement turnaround times

CloudNC is expanding its software line with Quote Agent, an AI-assisted estimating tool scheduled for release later in 2026.

Preparing job estimates represents a major operational drag for manufacturing suppliers. Evaluating incoming technical drawings, calculating cycle times, and establishing part pricing remains heavily manual, exposing supply shops to administrative delays or miscalculated margins once components enter physical production.

Quote Agent applies AI to early-stage costing, enabling suppliers to return customer bids rapidly while standardising cost estimations.

“Quote Agent is a natural next step for CloudNC as we seek to accelerate global machining with AI,” says Saville. “CAM Assist already helps machinists get parts onto machines faster; Quote Agent will help shops assess new work, prepare quotes more efficiently and respond to customers with greater confidence.”

“We believe our AI can make quoting faster, more consistent and more scalable, helping manufacturers win more of the right work while keeping expert judgement firmly in control,” Saville added.

Knox Systems partnership advances FedRAMP authorisation

CloudNC is collaborating with Knox Systems to achieve FedRAMP certification for CAM Assist.

The compliance roadmap aims to clear CAM Assist for deployment by US government departments, defence contractors, and aerospace manufacturers operating under federal data governance rules. 

Authorisation, if granted, will permit public-sector and defence suppliers to deploy automated toolpath generation across regulated production workloads.

See also: Samsung taps Mistral AI models for semiconductor manufacturing

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post CloudNC aims to accelerate AI supply chain machining appeared first on AI News.

  •  

Samsung taps Mistral AI models for semiconductor manufacturing

Samsung has partnered with Mistral AI to deploy on-premises models across its semiconductor manufacturing and engineering operations.

The agreement was announced during the bilateral state summit held in Paris between South Korea and France. Samsung will integrate Mistral’s software suite – including its flagship Mistral Large model – into internal semiconductor facilities to build customised models for intelligence-driven factory infrastructure.

On-premises AI models for semiconductor fab infrastructure

The deployment relies on private enterprise installations to process sensitive engineering and operational records within Samsung’s computing perimeter. This architecture keeps proprietary technical data contained within company infrastructure, avoiding external cloud exposure while maintaining control over operational assets.

“Increasing complexities involved in AI chip design and manufacturing requires continuous innovation in semiconductor technologies,” says Young Hyun Jun, Vice Chairman and CEO of the Device Solutions (DS) Division at Samsung Electronics.

Mistral will provide Samsung with a specialised stack of software tools to assist in how processors are designed and produced.

“AI is reshaping how we build complex technologies, from silicon to software,” says Arthur Mensch, co-founder and CEO of Mistral.

“We are proud to support Samsung Electronics with our expertise in electronics and semiconductors, helping to improve how chips are designed and manufactured, and to accelerate technical progress across the global semiconductor and AI value chain.”

Defect detection and yield stabilisation

Samsung plans to deploy the targeted models directly to automated defect detection and fab machinery tuning. As semiconductor production processes advance, rapid data analysis inside the fab becomes necessary to maintain factory throughput.

The company expects targeted AI models to accelerate development cycles, improve manufacturing precision, and stabilise production yields across advanced memory and logic chips. The operational scope covers Samsung’s memory division, logic design units, and contract foundry business.

Samsung also led Mistral AI’s Series D funding round, securing a strategic equity stake to support long-term technical cooperation.

The lead investment expands cross-industry collaboration between silicon manufacturers and AI developers across advanced memory, logic, and foundry operations.

See also: Arm launches Total Design for Physical AI and robotics framework

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Samsung taps Mistral AI models for semiconductor manufacturing appeared first on AI News.

  •  

Arm launches Total Design for Physical AI and robotics framework

Arm has launched Arm Total Design for Physical AI alongside a new robotics framework to establish common standards across automated systems.

Physical industries – spanning mining, agriculture, manufacturing, and global transport – account for trillions of dollars in economic activity and an estimated $200 billion annual compute opportunity by the 2030s.

To address engineering fragmentation across these sectors, Arm is convening more than 80 partner organisations spanning software, hardware, and AI. Initial ecosystem participants include AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, and Unitree Robotics.

The initiative targets physical systems that combine AI models, runtime software, compute silicon, sensors, and actuators to sense, reason, and act in operational environments. Hardware manufacturers and software developers require standardised baselines to reduce integration risk, optimise compute workloads, and move from proof-of-concept testing to deployment at scale.

Arm standardises capability tiers for robotics systems

Robotics currently lacks a common method to describe, compare, and communicate system capabilities, according to an architectural manifesto (PDF) published by Arm chief architect Richard Grisenthwaite. This fragmentation makes robotic systems harder to design, integrate, and scale across industrial deployments.

In response, Arm has introduced the Robotics Capability Framework as a collaborative starting point for a shared technical vocabulary, patterned after the SAE Levels used for driving automation.

Arm’s new framework categorises robotic systems across progressing tiers of operational sophistication, mapping machines from reactive setups to context-aware, cognitive, and self-improving systems.

Each capability tier links real-world use cases to machine behaviours, outputs, and hardware constraints. These criteria establish parameters for system latency, compute placement, memory allocation, power constraints, determinism, and safety standards.

The robotics capability framework for physical AI by Arm.

Arm developed the initial baseline using feedback from across the robotics sector. Participating organisations contributing to the framework include Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, and Robotec.ai.

Virtual platforms accelerate pre-silicon automotive physical AI development

Arm Total Design for Physical AI extends a collaborative development structure previously used for cloud AI infrastructure. The programme brings together AI models, virtual platforms, digital twins, sensors, compute silicon, and software stacks to enable earlier development and testing cycles.

Autonomous transport and robotics face common technical requirements across sensory perception, AI processing, real-time control, safety, and power-efficient compute. Arm demonstrated this collaborative methodology in the automotive sector alongside AWS, Google, HERE, RemotiveLabs, and Siemens.

The participating automotive companies developed an integrated digital cockpit reference solution. This environment enabled software engineering teams to develop, test, and validate complex automotive code on the Arm Zena CSS platform prior to physical silicon availability.

Arm is now soliciting technical contributions from the wider engineering community to expand the Robotics Capability Framework as physical AI implementations progress.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: NVIDIA Jetson Orin Nano 2 brings physical AI to drones and robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Arm launches Total Design for Physical AI and robotics framework appeared first on AI News.

  •  

MG Ship adds AI route optimisation as logistics returns accelerate

MG Ship has introduced an AI route optimisation and carrier selection module as logistics deployments demonstrate rapid cost and time returns.

The technical module targets global retailers and commercial shippers, pairing automated routing algorithms with carrier recommendation systems across international trade corridors. The deployment arrives as enterprise supply chain operators report measurable operational returns from machine learning tools, moving capital allocations away from speculative trials toward production deployments.

Measurable returns from deploying AI for logistics

Suki Cheung, CEO of MG Ship, will present deployment metrics during a panel discussion at the upcoming WMX Asia conference. Cheung will join executives from Pos Malaysia, Omniva, and OnyX Space for the session, titled AI Beyond the Hype: Measurable Results in Logistics Today.

“Too many AI conversations in logistics remain focused on future possibilities,” said Cheung. “The reality is that AI is already delivering measurable business outcomes today. Leading organisations are reducing transportation costs, improving forecast accuracy, increasing warehouse productivity, and achieving payback within months rather than years.”

Industry operational data indicates that initial investment returns are concentrating across three primary workflows:

  • Dynamic route planning has reduced enterprise fuel consumption by 15–20 percent, improved transit speeds by 15–25 percent, and lowered overall transportation costs by 12–22 percent, with capital payback reached within three to six months.
  • Predictive demand forecasting has reduced projection errors by 20–40 percent, improved planning accuracy by up to 35 percent, and decreased excess inventory by 20–30 percent within six to 12 months.
  • Automated freight documentation processing has cut manual task duration by up to 85 percent, recovering initial expenditure inside three to six months.

Over five-year deployment cycles, enterprise adopters have recorded average operational expense reductions between 10–25 percent, accompanied by warehouse productivity gains of 25–35 percent.

Routing algorithms and carrier scoring

MG Ship built the new routing capability directly into its visibility and supply chain intelligence platform, which serves retailers, manufacturers, and freight operators across multiple international markets. The base system synthesises live cargo telemetry with trade intelligence, risk monitoring, and predictive analytics to support operational planning and trade financing.

The route optimisation engine processes live and historical lane transit logs, weather patterns, air and ocean port congestion indicators, customs risk alerts, and transit reliability data. Shippers receive automated recommendations identifying low-cost, low-risk transit paths.

Carrier evaluation features rank transport providers per lane and service tier. Rather than selecting capacity purely on spot freight pricing, the system scores carriers against historical on-time metrics, transit consistency, exception occurrences, claims rates, available volume, and total cost-to-serve.

Logistics teams can also execute scenario simulations prior to peak shipping quarters. The software models lead times, service levels, freight spend, and risk exposures under alternative carrier allocation rules.

Early enterprise implementations demonstrate lower lead-time variance, reduced expedited freight expenditure, and improved on-time-in-full delivery rates.

Cheung stated that the platform “does not simply tell businesses where their cargo is”, adding that “it recommends the best route, the right carrier, and the lowest-risk option based on real-time conditions, helping organisations make faster and more profitable decisions.”

See also: OneRail uses Nvidia AI for real-time last-mile delivery optimisation

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post MG Ship adds AI route optimisation as logistics returns accelerate appeared first on AI News.

  •  

NVIDIA to acquire Hugging Face for $12.93B

NVIDIA has agreed to acquire Hugging Face for $12.93 billion to scale the open-source model repository’s platform and infrastructure.

The transaction targets platform growth and infrastructure investment, aiming to expand AI access for enterprise developers, software engineers, and research institutions globally.

Built over the past decade by Clem Delangue, Julien Chaumond, Thomas Wolf, and their engineering team, Hugging Face serves as the primary home for the open model developer community.

Platform metrics show more than 18 million developers, researchers, and creators share more than three million models, 500,000 datasets, and one million applications. Commercial adoption includes more than 200,000 companies using the environment to discover, evaluate, customise, and deploy AI models.

Hardware neutrality and multi-cloud commitments

NVIDIA stated that Hugging Face will remain an open platform for the entire AI sector. Developers will retain full control over their selection of models, software frameworks, cloud providers, inference services, and computing hardware.

NVIDIA hardware will not be mandatory to build on or deploy software through the platform. The service will maintain operational support for alternative accelerators, multi-cloud architectures, and open-weight models from all third-party builders.

Jensen Huang, Founder and CEO of NVIDIA, said: “Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want, and the computing platforms they want.

“NVIDIA compute will not be required to build on or deploy through Hugging Face.”

Open-weight models and distributed development

Huang noted a recent open letter he co-authored regarding the role of open weights in the AI economy. The position argued that open weights broaden AI access and ensure technical leadership remains distributed across companies, academic institutions, and developer communities.

Under this operational model, commercial businesses, startups, universities, and public bodies can build on advanced capabilities without the expense of training baseline models from scratch.

The approach allows organisations to match specific models to operational tasks across factories, hospitals, farms, classrooms, and commercial businesses, while addressing cybersecurity and data sovereignty requirements.

Julien Chaumond, Co-Founder and CEO of Hugging Face, commented: “AI is at an inflection point. Open-source AI can become less relevant in the coming years if the big closed labs run away with it, or it can become the foundational fabric of the next phase of human civilisation.

“Those are vastly different outcomes, and we need the critical mass to ensure we give our collective best shot to the second outcome. Given Jensen Huang’s stance on open source AI and how he stepped up to defend it when it was under threat earlier in the summer, NVIDIA was the only partner we truly considered.”

Infrastructure expansion and brand preservation

NVIDIA stands as the largest contributor of open models and data to Hugging Face, with a portfolio of more than 500 open models and more than 250 open datasets. The company builds its libraries, tools, and models openly to allow external engineers to inspect, modify, and build atop the software.

Chart showing NVIDIA contributions to Hugging Face compared to rivals.

Technical integration will focus on applying NVIDIA infrastructure and engineering resources to improve repository reliability, safety controls, model evaluation tooling, inference execution, and deployment pipelines.

“I am honored that Clem came to me as he considered the next chapter of Hugging Face and believed NVIDIA would be a great home for the company, its community and the future of open models,” Huang stated.

Hugging Face will retain its independent brand identity following the completion of the transaction, with the existing team continuing operations across multi-cloud and multi-accelerator environments.

“This gives fuel to our long-term vision and mission of unlocking the community’s progress to ensure that AI, which is the greatest breakthrough of our lifetime, is accessible to as many people as possible,” Chaumond concludes.

See also: Motional and MIT AI explains self-driving car decisions

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post NVIDIA to acquire Hugging Face for $12.93B appeared first on AI News.

  •  

Motional and MIT AI explains self-driving car decisions

Motional and MIT researchers have built a system that lets self-driving cars explain their decisions in real-time, tackling the black-box problem in autonomous vehicle AI.

The work, published in Nature, comes from a team at Motional that includes CEO Laura Major, working alongside researchers from MIT’s Computer Science and Artificial Intelligence Laboratory. Their proposed method, called the Concept-Wrapper Network or CW-Net, aims to translate the internal calculations of a self-driving system’s neural network into concepts a human can actually read.

If a current self-driving car brakes hard on a clear road with no obvious hazard in sight, neither the driver nor a passenger has any way of knowing why. Modern self-driving systems increasingly rely on neural networks trained on large volumes of driving data. Those networks can perform well, but they don’t expose their reasoning, which is why engineers describe them as black boxes.

Translating neural network logic into human concepts

CW-Net works by converting a self-driving system’s internal logic into concepts such as “Approaching Stopped Vehicle” or “Close to Cyclist.” These could, according to Motional, appear on a dashboard showing which concepts are influencing the vehicle’s driving decisions as they happen.

The system is designed so the explanations aren’t generated after the fact as a guess at what the network might have been doing. Instead, the vehicle’s final decision-making system takes action based directly on these human-interpretable concepts, so a braking event traces back to a specific concept that triggered it. Motional describes this as causally faithful, distinguishing it from approaches that generate natural-language explanations, which can read as plausible without necessarily being accurate.

Laura Major frames the case for this kind of interpretability against the alternative of relying purely on end-to-end deep learning to handle driving decisions.

“The general end-to-end only approach can get to a really good 80-90 percent – maybe even 95 percent – solution, but that’s not good enough to remove a driver or to earn the trust of cities, communities, and customers,” she said.

Testing explainable AI for self-driving cars around Las Vegas

Explainable AI research has largely stayed confined to computer simulations in lab settings, according to Motional. The Motional and MIT team instead deployed CW-Net on an autonomous vehicle with an experienced safety operator in the driver’s seat, collecting data on a private test track and on public roads around Las Vegas.

The team used an earlier experimental version of its deep-learning-based planning system, described as showing competitive performance but with notable shortcomings that CW-Net could help surface. Two incidents from the testing illustrate what the system caught.

In one, the autonomous vehicle repeatedly stopped near a traffic cone, and the vehicle operator assumed the cone itself was triggering the behaviour. Researchers removed the cone and the car stopped anyway. CW-Net’s display showed the actual cause: the experimental planning system was hallucinating a stopped vehicle ahead, a pattern traced back to its training data. That explanation let the researchers understand, predict, and resolve the issue.

A second test involved a cyclist. The autonomous vehicle detected and stopped for the cyclist as expected, but CW-Net revealed that the experimental planning system wasn’t actually basing its decision on the cyclist’s presence. The safety driver responded by exercising more caution around cyclists after noticing this. Follow-up analysis confirmed that caution was warranted, because the vehicle’s braking in that case came from a safety backup system rather than the experimental deep-learning-based planner.

Performance held steady against explainability

Adding layers of explainability to an AI system carries a known cost in speed and performance, and Motional acknowledges that risk. However, when researchers benchmarked CW-Net against leading autonomous driving algorithms, the difference in driving capability came in at less than one percent.

The Las Vegas incidents show why that trade-off matters operationally rather than just academically. A safety driver who can see that a stop is caused by a hallucinated vehicle, or that a backup system rather than the primary planner is responsible for a manoeuvre, can respond and report with more precision than one working from behaviour alone.

That visibility feeds directly into how quickly an engineering team can diagnose a system, and how confidently a safety operator can distinguish between an intended behaviour and a fault.

Motional connects the CW-Net work to broader pressure on autonomous vehicle operators as the technology extends into new markets and jurisdictions. Regulators are naturally asking for more transparency about how AI systems reach their decisions, and it expects tools like CW-Net could move from research projects toward a baseline requirement.

Beyond passenger vehicles, autonomous drones and even robotic surgery are cited as other safety-critical domains where operators and developers will need ways to understand a system’s capabilities, limitations, and unexpected behaviours.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: MIT AI forecasts extreme weather without historical data

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Motional and MIT AI explains self-driving car decisions appeared first on AI News.

  •  

Deloitte: Scale ‘autonomous intelligence’ for real growth

Enterprise leaders must progress past generative applications and scale “autonomous intelligence” to capture real growth.

Generating text or summarising internal communications offers localised productivity improvements, yet these abilities rarely alter the core cost or revenue structure of a large organisation. Enterprises are now focused on deploying systems capable of independent execution. Leaders are demanding applications that can traverse internal networks, execute multi-step logic, and finalise transactions without constant human prompting.

Prakul Sharma, principal and AI & Insights Practice Leader at Deloitte Consulting LLP, said: “At Deloitte, we view this as the third stage on an intelligence maturity curve, from ‘assisted intelligence,’ in which AI and analytics help people interpret information, through ‘artificial intelligence,’ with machine learning augmenting human decisions, to ‘autonomous intelligence,’ where AI decides and executes in defined boundaries.

“Today’s GenAI-era abilities – like chatbots and conversational AI – sit in the middle of that curve. Agentic AI acts as the bridge into autonomy, and it is where the centre of gravity is changing now. The difference we are seeing is agency: GenAI produces an answer, while autonomous intelligence pursues an outcome by reasoning over a goal, invoking tools and data, and adapting as conditions change, with humans setting guardrails not driving every step.

“We’re seeing this show up in industries, and in every case, the unlock isn’t the agent itself, but the surrounding governance architecture of identity and human-in-the-loop checkpoints, making autonomy safe to scale.”

Forensic audits for targeted margin improvement

To extract actual economic value, these autonomous systems must integrate directly into revenue-generating or cost-heavy workflows.

Consider a scenario in enterprise procurement: an agentic application continuously cross-references supply chain inventory against live vendor pricing in an enterprise resource planning system. It can then independently authorise purchase orders in predefined financial parameters, halting only for human approval when deviations occur.

The same system must also carry a verifiable identity in the ERP, read pricing data that is current enough to be contractually binding, and operate in approval thresholds that legal and compliance have formally endorsed. Any one of those dependencies, left unresolved, collapses the case for autonomous execution entirely. Achieving this level of automation therefore requires a forensic examination of existing operations before allocating any compute resources.

Sharma outlines the method Deloitte uses to initiate this operational overhaul and locate areas where autonomy can generate tangible revenue:

“The first step we advise is starting with a decision audit and the process. We ask leaders to pick one or two value chains where outcomes are bottlenecked by decisions not by tasks in that process, and to map how those decisions get made today. We ask questions like who has the data, who has the authority, where the handoffs break, what actions are needed, and where judgement is being applied.

“Asking these questions surfaces the process workflows where autonomy will create real economic value, while simultaneously exposing any data and governance gaps that may have derailed a pilot. From there, we help leaders sequence the rewire: stand up the foundational layers with AI and agentic fabric, data, evals, agent identity, and human-in-the-loop patterns against that first value chain, prove it works, and then use it as the template to scale.”

Integrating the right data infrastructure and upstream architecture

Once the operational target is isolated, the technological execution frequently stalls owing to upstream friction. The underlying foundation models from major providers have advanced quickly enough to handle complex reasoning tasks, becoming largely interchangeable commodities. The friction point lies in connecting these reasoning engines to legacy data architectures.

Sharma observes that the true technical barriers emerge long before the prompt reaches the large language model:

“Based on what we are seeing, the model is rarely the bottleneck, since frontier ability is now rapidly becoming a commodity. Where enterprises trip up in the design phase is upstream of the model. They select a use case before mapping the underlying workflow, resulting in the agent automating a process that was already broken or poorly instrumented.

“The second pattern is data: clients may underestimate that autonomous systems need decision-grade data, not reporting-grade data, meaning lineage and access controls that most enterprise data estates were not built to support.”

The distinction matters because most enterprise data estates were built for human analysts, not autonomous systems. Reporting-grade data – aggregated on a nightly or weekly batch cycle, structured for dashboard consumption, and stripped of the lineage that records how a value was derived – is adequate when a person applies judgement before acting on it. An autonomous agent has no such backstop. When it retrieves a contract price or a stock level to execute a transaction, that figure must carry a timestamp current enough to be binding, a traceable provenance, and access controls that confirm the agent is authorised to read and act on it.

Providing this decision-grade data involves integrating autonomous agents with right event stores and databases designed to manage both structured and unstructured enterprise information. When an agent retrieves data to execute a task, the enterprise must guarantee its freshness. Relying on stale batch-processed data introduces extreme risk, potentially causing the system to act on obsolete pricing tiers or outdated compliance frameworks.

The financial model for scaling these systems also requires forecasting variable compute expenses. Because agentic workflows involve multiple interactions with large language models to reason through a single goal, API costs can escalate unpredictably. Mitigating hallucination risks through retrieval-augmented generation processes also increases the necessary compute overhead, requiring strict financial controls before enterprise-deployment.

Reconciling governance debt and enterprise ecosystems

Transitioning from controlled testing environments to live enterprise deployment is a very different proposition. A small-scale test might perform perfectly using carefully selected data sets, but deploying that ability in thousands of employees and interconnected software platforms exposes vulnerabilities.

Navigating modern enterprise security environments means integrating the agentic architecture deeply with existing identity providers and cloud-native security controls across hybrid cloud ecosystems.

Sharma identifies this integration failure and the resulting governance debt that halts progress:

“The main roadblock we see is what we call the production gap. A pilot can succeed with a clever prompt, a curated dataset, and a champion team running it manually, but enterprise deployment requires continuous evaluations, identity and authorisation that work in systems the pilot never touched, change management for the users, and a financial model that can absorb use-based costs at scale.

“Tied to that is governance debt: the controls, audit trails, and risk frameworks waived to accelerate a pilot often become the gating items once legal and compliance evaluate a production rollout. The clients that break through are ones that don’t treat pilots as experiments but instead treat them as the first production instance of a reusable platform – with the same evals, identity model, and governance. Instead of starting over, this allows the second and third use cases to build on the first.”

Compliance frameworks applied during initial testing are often completely insufficient for live deployment. Teams eager to prove a concept frequently bypass standard corporate security protocols, creating the very gating items that prevent future scaling.

What unites all three failure modes – the production gap, governance debt, and upstream data friction – is that each one is invisible during a well-run pilot. A champion team with a curated dataset and management cover can paper over missing identity controls, stale data, and deferred compliance reviews for long enough to produce a convincing demonstration. It is only when the system must operate in the full enterprise, with real users, live data, and legal scrutiny, that the gaps become structural blockers not known workarounds.

Building a reusable platform from the outset – with identity verification, continuous model evaluations, and financial monitoring treated as first-class requirements not post-launch additions – is what allows organisations to avoid rebuilding those foundations for every subsequent deployment.

Prakul Sharma’s interview was conducted ahead of the AI & Big Data Expo North America, where Deloitte is a important sponsor. Be sure to swing by Deloitte’s booth at stand #272 to hear more directly from the organisation’s experts. Prakul Sharma will be sharing more of his insights during a panel session on day one and day two of the industry-leading event.

 

(Image source: Pixabay, under licence.)

 

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Deloitte: Scale ‘autonomous intelligence’ for real growth appeared first on AI News.

  •  

IBM: How robust AI governance protects enterprise margins

To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure.

When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely.

At the initial product stage, exerting tight corporate control often feels highly advantageous. Closed development environments iterate quickly and tightly manage the end-user experience. They capture and concentrate financial value within a single corporate entity, an approach that functions adequately during early product development cycles.

However, IBM’s analysis highlights that expectations change entirely when a technology solidifies into a foundational layer. Once other institutional frameworks, external markets, and broad operational systems rely on the software, the prevailing standards adapt to a new reality. At infrastructure scale, embracing openness ceases to be an ideological stance and becomes a highly practical necessity.

AI is currently crossing this threshold within the enterprise architecture stack. Models are increasingly embedded directly into the ways organisations secure their networks, author source code, execute automated decisions, and generate commercial value. AI functions less as an experimental utility and more as core operational infrastructure.

The recent limited preview of Anthropic’s Claude Mythos model brings this reality into sharper focus for enterprise executives managing risk. Anthropic reports that this specific model can discover and exploit software vulnerabilities at a level matching few human experts.

In response to this power, Anthropic launched Project Glasswing, a gated initiative designed to place these advanced capabilities directly into the hands of network defenders first. From IBM’s perspective, this development forces technology officers to confront immediate structural vulnerabilities. If autonomous models possess the capability to write exploits and shape the overall security environment, Thomas notes that concentrating the understanding of these systems within a small number of technology vendors invites severe operational exposure.

With models achieving infrastructure status, IBM argues the primary issue is no longer exclusively what these machine learning applications can execute. The priority becomes how these systems are constructed, governed, inspected, and actively improved over extended periods.

As underlying frameworks grow in complexity and corporate importance, maintaining closed development pipelines becomes exceedingly difficult to defend. No single vendor can successfully anticipate every operational requirement, adversarial attack vector, or system failure mode.

Implementing opaque AI structures introduces heavy friction across existing network architecture. Connecting closed proprietary models with established enterprise vector databases or highly sensitive internal data lakes frequently creates massive troubleshooting bottlenecks. When anomalous outputs occur or hallucination rates spike, teams lack the internal visibility required to diagnose whether the error originated in the retrieval-augmented generation pipeline or the base model weights.

Integrating legacy on-premises architecture with highly gated cloud models also introduces severe latency into daily operations. When enterprise data governance protocols strictly prohibit sending sensitive customer information to external servers, technology teams are left attempting to strip and anonymise datasets before processing. This constant data sanitisation creates enormous operational drag. 

Furthermore, the spiralling compute costs associated with continuous API calls to locked models erode the exact profit margins these autonomous systems are supposed to enhance. The opacity prevents network engineers from accurately sizing hardware deployments, forcing companies into expensive over-provisioning agreements to maintain baseline functionality.

Why open-source AI is essential for operational resilience

Restricting access to powerful applications is an understandable human instinct that closely resembles caution. Yet, as Thomas points out, at massive infrastructure scale, security typically improves through rigorous external scrutiny rather than through strict concealment.

This represents the enduring lesson of open-source software development. Open-source code does not eliminate enterprise risk. Instead, IBM maintains it actively changes how organisations manage that risk. An open foundation allows a wider base of researchers, corporate developers, and security defenders to examine the architecture, surface underlying weaknesses, test foundational assumptions, and harden the software under real-world conditions.

Within cybersecurity operations, broad visibility is rarely the enemy of operational resilience. In fact, visibility frequently serves as a strict prerequisite for achieving that resilience. Technologies deemed highly important tend to remain safer when larger populations can challenge them, inspect their logic, and contribute to their continuous improvement.

Thomas addresses one of the oldest misconceptions regarding open-source technology: the belief that it inevitably commoditises corporate innovation. In practical application, open infrastructure typically pushes market competition higher up the technology stack. Open systems transfer financial value rather than destroying it.

As common digital foundations mature, the commercial value relocates toward complex implementation, system orchestration, continuous reliability, trust mechanics, and specific domain expertise. IBM’s position asserts that the long-term commercial winners are not those who own the base technological layer, but rather the organisations that understand how to apply it most effectively.

We have witnessed this identical pattern play out across previous generations of enterprise tooling, cloud infrastructure, and operating systems. Open foundations historically expanded developer participation, accelerated iterative improvement, and birthed entirely new, larger markets built on top of those base layers. Enterprise leaders increasingly view open-source as highly important for infrastructure modernisation and emerging AI capabilities. IBM predicts that AI is highly likely to follow this exact historical trajectory.

Looking across the broader vendor ecosystem, leading hyperscalers are adjusting their business postures to accommodate this reality. Rather than engaging in a pure arms race to build the largest proprietary black boxes, highly profitable integrators are focusing heavily on orchestration tooling that allows enterprises to swap out underlying open-source models based on specific workload demands. Highlighting its ongoing leadership in this space, IBM is a key sponsor of this year’s AI & Big Data Expo North America, where these evolving strategies for open enterprise infrastructure will be a primary focus.

This approach completely sidesteps restrictive vendor lock-in and allows companies to route less demanding internal queries to smaller and highly efficient open models, preserving expensive compute resources for complex customer-facing autonomous logic. By decoupling the application layer from the specific foundation model, technology officers can maintain operational agility and protect their bottom line.

The future of enterprise AI demands transparent governance

Another pragmatic reason for embracing open models revolves around product development influence. IBM emphasises that narrow access to underlying code naturally leads to narrow operational perspectives. In contrast, who gets to participate directly shapes what applications are eventually built. 

Providing broad access enables governments, diverse institutions, startups, and varied researchers to actively influence how the technology evolves and where it is commercially applied. This inclusive approach drives functional innovation while simultaneously building structural adaptability and necessary public legitimacy.

As Thomas argues, once autonomous AI assumes the role of core enterprise infrastructure, relying on opacity can no longer serve as the organising principle for system safety. The most reliable blueprint for secure software has paired open foundations with broad external scrutiny, active code maintenance, and serious internal governance.

As AI permanently enters its infrastructure phase, IBM contends that identical logic increasingly applies directly to the foundation models themselves. The stronger the corporate reliance on a technology, the stronger the corresponding case for demanding openness.

If these autonomous workflows are truly becoming foundational to global commerce, then transparency ceases to be a subject of casual debate. According to IBM, it is an absolute, non-negotiable design requirement for any modern enterprise architecture.

See also: Why companies like Apple are building AI agents with limits

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post IBM: How robust AI governance protects enterprise margins appeared first on AI News.

  •  

Microsoft open-source toolkit secures AI agents at runtime

A new open-source toolkit from Microsoft focuses on runtime security to force strict governance onto enterprise AI agents. The release tackles a growing anxiety: autonomous language models are now executing code and hitting corporate networks way faster than traditional policy controls can keep up.

AI integration used to mean conversational interfaces and advisory copilots. Those systems had read-only access to specific datasets, keeping humans strictly in the execution loop. Organisations are currently deploying agentic frameworks that take independent action, wiring these models directly into internal application programming interfaces, cloud storage repositories, and continuous integration pipelines.

When an autonomous agent can read an email, decide to write a script, and push that script to a server, stricter governance is vital. Static code analysis and pre-deployment vulnerability scanning just can’t handle the non-deterministic nature of large language models. One prompt injection attack (or even a basic hallucination) could send an agent to overwrite a database or pull out customer records.

Microsoft’s new toolkit looks at runtime security instead, providing a way to monitor, evaluate, and block actions at the moment the model tries to execute them. It beats relying on prior training or static parameter checks.

Intercepting the tool-calling layer in real time

Looking at the mechanics of agentic tool calling shows how this works. When an enterprise AI agent has to step outside its core neural network to do something like query an inventory system, it generates a command to hit an external tool.

Microsoft’s framework drops a policy enforcement engine right between the language model and the broader corporate network. Every time the agent tries to trigger an outside function, the toolkit grabs the request and checks the intended action against a central set of governance rules. If the action breaks policy (e.g. an agent authorised only to read inventory data tries to fire off a purchase order) the toolkit blocks the API call and logs the event so a human can review it.

Security teams get a verifiable, auditable trail of every single autonomous decision. Developers also win here; they can build complex multi-agent systems without having to hardcode security protocols into every individual model prompt. Security policies get decoupled from the core application logic entirely and are managed at the infrastructure level.

Most legacy systems were never built to talk to non-deterministic software. An old mainframe database or a customised enterprise resource planning suite doesn’t have native defenses against a machine learning model shooting over malformed requests. Microsoft’s toolkit steps in as a protective translation layer. Even if an underlying language model gets compromised by external inputs; the system’s perimeter holds.

Security leaders might wonder why Microsoft decided to release this runtime toolkit under an open-source license. It comes down to how modern software supply chains actually work.

Developers are currently rushing to build autonomous workflows using a massive mix of open-source libraries, frameworks, and third-party models. If Microsoft locked this runtime security feature to its proprietary platforms, development teams would probably just bypass it for faster, unvetted workarounds to hit their deadlines.

Pushing the toolkit out openly means security and governance controls can fit into any technology stack. It doesn’t matter if an organisation runs local open-weight models, leans on competitors like Anthropic, or deploys hybrid architectures.

Setting up an open standard for AI agent security also lets the wider cybersecurity community chip in. Security vendors can stack commercial dashboards and incident response integrations on top of this open foundation, which speeds up the maturity of the whole ecosystem. For businesses, they avoid vendor lock-in but still get a universally scrutinised security baseline.

The next phase of enterprise AI governance

Enterprise governance doesn’t just stop at security; it hits financial and operational oversight too. Autonomous agents run in a continuous loop of reasoning and execution, burning API tokens at every step. Startups and enterprises are already seeing token costs explode when they deploy agentic systems.

Without runtime governance, an agent tasked with looking up a market trend might decide to hit an expensive proprietary database thousands of times before it finishes. Left alone, a badly configured agent caught in a recursive loop can rack up massive cloud computing bills in a few hours.

The runtime toolkit gives teams a way to slap hard limits on token consumption and API call frequency. By setting boundaries on exactly how many actions an agent can take within a specific timeframe, forecasting computing costs gets much easier. It also stops runaway processes from eating up system resources.

A runtime governance layer hands over the quantitative metrics and control mechanisms needed to meet compliance mandates. The days of just trusting model providers to filter out bad outputs are ending. System safety now falls on the infrastructure that actually executes the models’ decisions

Getting a mature governance program off the ground is going to demand tight collaboration between development operations, legal, and security teams. Language models are only scaling up in capability, and the organisations putting strict runtime controls in place today are the only ones who will be equipped to handle the autonomous workflows of tomorrow.

See also: As AI agents take on more tasks, governance becomes a priority

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft open-source toolkit secures AI agents at runtime appeared first on AI News.

  •  
❌