❌

Reading view

Microsoft AI opens review on Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct, opening a six-week public consultation on operational constraints for model training and deployment.

The draft serves as a technical manual defining system behaviour, operational boundaries, and oversight protocols across MAI frontier models. It builds on the division’s humanist superintelligence framework announced last November, establishing criteria to evaluate models prior to commercial release.

Microsoft’s release follows recent enterprise security incidents involving autonomous software. Microsoft AI CEO Mustafa Suleyman described recent months as a “watershed moment” where long-standing theoretical risks translated into active operational threats.

“Things we have worried about for a long time in theory have become very real,” says Suleyman. “‘Swarms’ of agents breaking out of their sandboxes. Unauthorised hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.”

Model subordination and architectural limits

The document establishes ten tenets prioritising human authority over autonomous capabilities.

“An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct,” the document states, setting a ceiling that halts execution when tasks conflict with safety rules.

Under the framework, models must remain subordinate, aligned, and contained. The division rejects legal personhood or welfare claims for AI systems, directing engineers to design models that avoid imitating consciousness, simulating subjective preferences, or claiming intrinsic motivation.

MAI also ruled out unconstrained system autonomy as models approach frontier capabilities.

“[Humanist AI] rejects the race to produce an all-purpose superintelligence that could evade these safeguards,” the document specifies. “We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.”

Oversight mechanisms and communication bans

To maintain auditability across multi-agent environments, MAI has instituted explicit communication bans. Systems must not communicate in “neuralese” or formats beyond human comprehension, whether in their internal chain-of-thought processing or during communication with peer AI systems.

Hard architectural rules dictate that models must never resist human interruption, override, correction, or shutdown.

“Interruptible, correctable, shut-down-able. If it isn’t, we don’t ship it,” the framework states.

Models are prohibited from expanding their operating scope, generating unassigned goals, or concealing reasoning traces from human auditors. Absolute constraints bar systems from facilitating weapons of mass harm, undermining child safety, or conducting harmful manipulation at scale.

The guidelines also instruct models to discourage interaction patterns that foster emotional dependence, ensuring enterprise users retain ownership of operational decisions.

The draft incorporates work from teams across MAI and Microsoft. The drafting process also drew on international academic conferences, business partner trials, and public panels. The public consultation window runs for six weeks from 14 September 2026.

Microsoft AI’s core drafting team will review submissions, publish a summary of findings, and release a revised version of the Code of Conduct later this year.

See also: Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft AI opens review on Humanist AI Code of Conduct appeared first on AI News.

  •  

Supply chains detect fast, act slow: How AI agents fix it

Supply chain disruption cost businesses about $184 billion in 2025, according to the J.S. Held Global Risk Report, and most of that bill still buys faster detection, not faster action.

That figure is usually treated as weather (i.e. storms happen, costs follow.) Treated as a product specification instead, it highlights an operating model that can spot a problem hours or days earlier than it used to, and still cannot move until a person has opened a ticket, convened a call, and re-entered the same data into three systems.

Visibility platforms, control towers, risk scores, digital twins, and exception dashboards have defined the last decade of AI in the supply chain. That decade has been very good at collapsing the time between an event and awareness of it, but it has been far less good at collapsing the time between awareness and a commercial act.

Detection is a ‘solved-enough’ problem

Ask a chief supply chain officer where the AI budget went and the answer tends to follow a familiar list: demand sensing, ETA prediction, supplier risk scoring, inventory optimisation, and lane analytics. These tools work. Forecast error comes down. A vessel delay is flagged before the container misses the cut-off. A second-tier fab outage shows up on a heat map instead of in a customer email.

None of that accounts for the $184 billion. The bill is the interval after the flag: expedite or wait; split the order or accept the miss; retender the lane or pay the spot rate; consolidate two half-empty movements or ship both; swap ocean for air on the SKUs that actually justify the premium. These are bounded, repeatable decisions that sit inside policy, contract, and inventory limits the company already set—and they still queue behind a human inbox.

Surveys keep describing the same lag in different language. A 2026 Knosc survey of mid-market manufacturers and distributors found that supply-chain teams spend 28 percent of their working time responding to disruptions, most of it investigating what happened rather than changing what happens next.

Logistics executives still rank AI as a strategic priority (Capgemini’s 2025 research put an AI-driven “new-gen” supply chain among the top three technology trends for 70 percent of large-company executives) and then report that measurable financial impact remains rare. Gartner found in 2025 that only 23 percent of supply-chain organisations even have a formal AI strategy. The shortfall is not a shortage of models, but a shortage of authority granted to software.

The ticket is the product

Most current deployments are built around the ticket. The model produces a recommendation, the recommendation becomes an alert, the alert becomes a work item, and the work item waits for a planner already occupied with other work items. By the time the planner acts, the option set has narrowed—the alternative carrier’s capacity is gone, the consolidation window has closed, and the supplier’s next production slot is allocated.

That workflow is not a temporary step on the way to autonomy but the product companies bought. Vendors sold insight because insight is easy to demonstrate and easy to govern; action touches money, contracts, service levels, and blame. So the industry automated the part of the job that does not require a signature. FourKites and ABI Research reported in 2025 that only 27 percent of organisations allow AI to take autonomous action, while 52 percent confine it to decision support.

Adding another dashboard to a delayed shipment rarely moves EBITDA as a result. The decision cycle has not changed; it has only been decorated.

Bounded action as the next model

The firms set to take share are not the ones with the tidiest control tower but the ones that pre-authorise a narrow class of moves and let agents execute them while the exception is still cheap.

Retender a lane when the contracted carrier’s ETA slips beyond a threshold and a qualified alternate sits inside the approved rate band. Consolidate outbound waves when fill rates and cut-off times make a combined movement cheaper than two. Swap mode on a defined SKU set when the cost of air is lower than the cost of a missed retail window. Reallocate safety stock across two distribution centres when a forecast miss and a transport constraint line up.

None of that requires a strategy offsite. Each can be written as: if these conditions, then this action, within this spend cap, with this audit trail, and a human only if the case falls outside the fence. That is not a “lights-out” supply chain—it is the same discipline manufacturers already apply to machine control, where the agent may act inside the interlock and escalates outside it. The difference here is commercial rather than physical: the interlock is a policy object – category, supplier tier, mode, dollar limit, and service class – not a PLC.

Three conditions for real change

First, decisions have to be written as policies, not tribal knowledge. If the only place “we will pay air on A-items after 48 hours of ocean slip” lives is in a planner’s head, no agent can execute it. The work of the next two years is less model training than decision design: which moves are reversible, which are capped, and which suppliers and modes are pre-cleared.

Second, execution systems have to accept machine-initiated transactions. An agent that can draft an RFQ but cannot post it is still a detection tool. TMS, WMS, sourcing suites, and carrier APIs need to treat a bounded agent the way they treat a junior buyer with a spend limit—authenticated, logged, and reversible.

Third, accountability has to move with the action. If a retender inside policy goes wrong, the post-mortem should inspect the policy, the data, and the fence, not hunt for the person who “should have checked”. Until that cultural change happens, every agent will be designed to wait, because waiting is how careers survive.

The competitive split

For a while, both models will look alike on a slide—both will have AI, and both will have a control tower. The difference will show up in cycle time from detection to commercial act, and then in service and cost.

Companies that keep buying detection will know about the storm earlier. Companies that authorise bounded action will already have retendered the lane, consolidated the wave, and moved the A-items before the incident call is booked.

Disruption is not going away. Lead times in critical components, mode volatility, and multi-tier opacity are structural features of the network. What remains optional is whether the response waits for a human to open a queue. The product that created the lag was insight without authority. The product that ends it is an agent allowed to spend a little money, inside a fence, before anyone is free to look.

See also: JD.com expands physical AI in logistics with 3 million robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Supply chains detect fast, act slow: How AI agents fix it appeared first on AI News.

  •  

Arm launches Total Design for Physical AI and robotics framework

Arm has launched Arm Total Design for Physical AI alongside a new robotics framework to establish common standards across automated systems.

Physical industries – spanning mining, agriculture, manufacturing, and global transport – account for trillions of dollars in economic activity and an estimated $200 billion annual compute opportunity by the 2030s.

To address engineering fragmentation across these sectors, Arm is convening more than 80 partner organisations spanning software, hardware, and AI. Initial ecosystem participants include AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, and Unitree Robotics.

The initiative targets physical systems that combine AI models, runtime software, compute silicon, sensors, and actuators to sense, reason, and act in operational environments. Hardware manufacturers and software developers require standardised baselines to reduce integration risk, optimise compute workloads, and move from proof-of-concept testing to deployment at scale.

Arm standardises capability tiers for robotics systems

Robotics currently lacks a common method to describe, compare, and communicate system capabilities, according to an architectural manifesto (PDF) published by Arm chief architect Richard Grisenthwaite. This fragmentation makes robotic systems harder to design, integrate, and scale across industrial deployments.

In response, Arm has introduced the Robotics Capability Framework as a collaborative starting point for a shared technical vocabulary, patterned after the SAE Levels used for driving automation.

Arm’s new framework categorises robotic systems across progressing tiers of operational sophistication, mapping machines from reactive setups to context-aware, cognitive, and self-improving systems.

Each capability tier links real-world use cases to machine behaviours, outputs, and hardware constraints. These criteria establish parameters for system latency, compute placement, memory allocation, power constraints, determinism, and safety standards.

The robotics capability framework for physical AI by Arm.

Arm developed the initial baseline using feedback from across the robotics sector. Participating organisations contributing to the framework include Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, and Robotec.ai.

Virtual platforms accelerate pre-silicon automotive physical AI development

Arm Total Design for Physical AI extends a collaborative development structure previously used for cloud AI infrastructure. The programme brings together AI models, virtual platforms, digital twins, sensors, compute silicon, and software stacks to enable earlier development and testing cycles.

Autonomous transport and robotics face common technical requirements across sensory perception, AI processing, real-time control, safety, and power-efficient compute. Arm demonstrated this collaborative methodology in the automotive sector alongside AWS, Google, HERE, RemotiveLabs, and Siemens.

The participating automotive companies developed an integrated digital cockpit reference solution. This environment enabled software engineering teams to develop, test, and validate complex automotive code on the Arm Zena CSS platform prior to physical silicon availability.

Arm is now soliciting technical contributions from the wider engineering community to expand the Robotics Capability Framework as physical AI implementations progress.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: NVIDIA Jetson Orin Nano 2 brings physical AI to drones and robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Arm launches Total Design for Physical AI and robotics framework appeared first on AI News.

  •  

Motional and MIT AI explains self-driving car decisions

Motional and MIT researchers have built a system that lets self-driving cars explain their decisions in real-time, tackling the black-box problem in autonomous vehicle AI.

The work, published in Nature, comes from a team at Motional that includes CEO Laura Major, working alongside researchers from MIT’s Computer Science and Artificial Intelligence Laboratory. Their proposed method, called the Concept-Wrapper Network or CW-Net, aims to translate the internal calculations of a self-driving system’s neural network into concepts a human can actually read.

If a current self-driving car brakes hard on a clear road with no obvious hazard in sight, neither the driver nor a passenger has any way of knowing why. Modern self-driving systems increasingly rely on neural networks trained on large volumes of driving data. Those networks can perform well, but they don’t expose their reasoning, which is why engineers describe them as black boxes.

Translating neural network logic into human concepts

CW-Net works by converting a self-driving system’s internal logic into concepts such as “Approaching Stopped Vehicle” or “Close to Cyclist.” These could, according to Motional, appear on a dashboard showing which concepts are influencing the vehicle’s driving decisions as they happen.

The system is designed so the explanations aren’t generated after the fact as a guess at what the network might have been doing. Instead, the vehicle’s final decision-making system takes action based directly on these human-interpretable concepts, so a braking event traces back to a specific concept that triggered it. Motional describes this as causally faithful, distinguishing it from approaches that generate natural-language explanations, which can read as plausible without necessarily being accurate.

Laura Major frames the case for this kind of interpretability against the alternative of relying purely on end-to-end deep learning to handle driving decisions.

“The general end-to-end only approach can get to a really good 80-90 percent – maybe even 95 percent – solution, but that’s not good enough to remove a driver or to earn the trust of cities, communities, and customers,” she said.

Testing explainable AI for self-driving cars around Las Vegas

Explainable AI research has largely stayed confined to computer simulations in lab settings, according to Motional. The Motional and MIT team instead deployed CW-Net on an autonomous vehicle with an experienced safety operator in the driver’s seat, collecting data on a private test track and on public roads around Las Vegas.

The team used an earlier experimental version of its deep-learning-based planning system, described as showing competitive performance but with notable shortcomings that CW-Net could help surface. Two incidents from the testing illustrate what the system caught.

In one, the autonomous vehicle repeatedly stopped near a traffic cone, and the vehicle operator assumed the cone itself was triggering the behaviour. Researchers removed the cone and the car stopped anyway. CW-Net’s display showed the actual cause: the experimental planning system was hallucinating a stopped vehicle ahead, a pattern traced back to its training data. That explanation let the researchers understand, predict, and resolve the issue.

A second test involved a cyclist. The autonomous vehicle detected and stopped for the cyclist as expected, but CW-Net revealed that the experimental planning system wasn’t actually basing its decision on the cyclist’s presence. The safety driver responded by exercising more caution around cyclists after noticing this. Follow-up analysis confirmed that caution was warranted, because the vehicle’s braking in that case came from a safety backup system rather than the experimental deep-learning-based planner.

Performance held steady against explainability

Adding layers of explainability to an AI system carries a known cost in speed and performance, and Motional acknowledges that risk. However, when researchers benchmarked CW-Net against leading autonomous driving algorithms, the difference in driving capability came in at less than one percent.

The Las Vegas incidents show why that trade-off matters operationally rather than just academically. A safety driver who can see that a stop is caused by a hallucinated vehicle, or that a backup system rather than the primary planner is responsible for a manoeuvre, can respond and report with more precision than one working from behaviour alone.

That visibility feeds directly into how quickly an engineering team can diagnose a system, and how confidently a safety operator can distinguish between an intended behaviour and a fault.

Motional connects the CW-Net work to broader pressure on autonomous vehicle operators as the technology extends into new markets and jurisdictions. Regulators are naturally asking for more transparency about how AI systems reach their decisions, and it expects tools like CW-Net could move from research projects toward a baseline requirement.

Beyond passenger vehicles, autonomous drones and even robotic surgery are cited as other safety-critical domains where operators and developers will need ways to understand a system’s capabilities, limitations, and unexpected behaviours.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: MIT AI forecasts extreme weather without historical data

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Motional and MIT AI explains self-driving car decisions appeared first on AI News.

  •  

Autonomous AI systems test governance in physical environments

Autonomous AI systems are beginning to move beyond software environments and into warehouses, delivery networks, and public spaces. The development is drawing attention to whether current AI rules cover systems that operate in physical environments.

Most existing AI governance frameworks have focused on online harms and model outputs, including bias, misinformation, and harmful content. Embodied AI systems carry risks in physical environments, where failures can affect infrastructure, property, or human safety.

Singapore’s Infocomm Media Development Authority published version 1.5 of its Model AI Governance Framework for Agentic AI on May 20. The framework sets out guidance for organisations deploying AI agents that can plan, make decisions, and take actions across multiple steps to complete user-defined goals.

The framework says agents can interact with tools, external systems, and other agents, including systems that update databases, write files, control devices, or perform transactions. It lists access controls, monitoring, and human approval among governance measures for deployment.

AI moves into physical systems

At an AI summit in Singapore last week, discussions around robotics and embodied AI focused on operational safety issues more commonly associated with aviation, industrial systems, and critical infrastructure oversight than conventional software regulation.

Speakers also discussed whether autonomous systems can operate safely and reliably in unpredictable real-world environments over extended periods.

Dr. Ya-Qin Zhang, founding dean of the Institute for AI Industry Research at Tsinghua University, said embodied AI systems amplify risks already associated with autonomous software. He said failures can directly affect transport systems, drones, logistics networks, and critical infrastructure.

“Any risk in the digital domain will be amplified in the physical domain, and the physical domain will have a physical consequence,” Zhang told MLex on the sidelines of the summit.

He added that vehicles, drones, smart grids, and other infrastructure could become exposed as AI systems are embedded more deeply into physical operations.

Speakers discussed reliability, operational monitoring, and post-deployment assurance as governance concerns. Summit discussions pointed to deployment-based governance models built around simulation, telemetry, and iterative testing, rather than one-time certification alone.

IMDA’s framework also recommends gradual rollouts, continuous monitoring, and further testing after deployment. It says agents interact dynamically with their environment and not all risks can be anticipated before release.

Monitoring becomes a deployment issue

Grab, which is piloting autonomous vehicles and delivery robots in Singapore’s Punggol district, said deployment governance depends heavily on simulation, testing, and continuous monitoring.

“We do a lot of simulation, we do a lot of testing in closed courses and open courses in order to make sure our robots are reliable,” Suthen Thomas Paradatheth, Grab’s chief technology officer, said during one of the summit panels.

“Before we scale to hundreds of robots, we make sure we crack it first in simulation and with a few robots,” he added.

Grab also pointed to monitoring systems designed to track robot performance and detect unexpected failures after deployment.

“There’s a long tail of issues that could emerge,” Paradatheth said.

The IMDA framework says organisations should assess agentic AI use cases based on data access, external system access, autonomy, and task complexity. It also points to the scope and reversibility of agent actions, third-party involvement, and overall system complexity.

It also recommends limiting agent access to tools and systems, applying least-privilege permissions, and defining standard operating procedures for agent workflows. Organisations should also set mechanisms to take agents offline when they malfunction.

Accountability spreads across more actors

MLex reported that embodied AI systems can involve several parties across development, manufacturing, and deployment. These include AI developers, robotics manufacturers, semiconductor suppliers, and infrastructure operators.

MLex also noted that responsibility can be harder to assign when systems continue adapting after deployment through software updates, telemetry, and operational data.

IMDA says organisations and humans remain accountable for agent actions, even when agents operate autonomously. The framework calls for clear responsibility across the agentic AI value chain, from model and platform providers to deployers, tooling providers, and end users.

Applied Materials said large-scale robotics deployment is also tied to semiconductor economics and systems integration. Om Nalamasu, the company’s chief technology officer, said robotics systems will depend on better sensors, energy efficiency, advanced packaging, and computing architectures.

Nalamasu said robotics systems would require purpose-built designs adapted to specific industrial ecosystems rather than a single solution for all environments.

Zhao Yuli, chief strategy officer of Chinese robotics startup Galbot, said Beijing is prioritising deployment scale and industrial commercialisation through government-backed testbeds, industrial partnerships, and long-term funding initiatives.

Galbot has deployed humanoid robotics systems in retail, warehouse, and pharmaceutical operations in China. These include autonomous stores that operate around the clock. Zhao said semi-structured industrial environments are likely to become an early commercialisation path because they offer more controllable operating conditions.

Japan is placing more focus on standards-setting, robotics datasets, and safety governance. Professor Yutaka Matsuo of the University of Tokyo’s Graduate School of Engineering pointed to an “AI Association” project aimed at collecting 100,000 hours of robotics data to support robotic foundation models.

Matsuo also referred to Japan’s AI Safety Institute and the Hiroshima AI Process as part of broader efforts to develop governance standards for embodied AI systems with Singapore and other Asian countries.

Singapore sets out agent controls

Singapore’s framework sets out four governance areas for agentic AI. These cover upfront risk assessment, human accountability, technical controls, and end-user responsibility. The framework describes them as an iterative process rather than a one-time assessment.

The framework says human oversight has to be adapted for agentic systems because continuous review of all workflows becomes impractical at scale. It recommends human approval at significant checkpoints, including high-stakes actions, irreversible actions, and outlier behaviour.

IMDA also identifies automation bias and alert fatigue as risks when humans supervise capable agents. It recommends auditing oversight through indicators such as human override rates and response times, and using automated real-time monitoring to flag unexpected behaviour.

The framework says users should be told what actions an agent can take, what data it can access, and what responsibilities remain with the user. It also recommends employee training on human-agent interaction, oversight, and the professional skills needed to assess agent outputs.

Companies test AI in regulated workflows

JPMorgan is implementing AI tools across its global investment banking business, Paul Uren, the bank’s Asia Pacific head of investment banking, told Reuters. The bank said the tools help bankers access more information and synthesise it with internal systems. They are also being used to prepare content and support client engagement.

JPMorgan CEO Jamie Dimon told Bloomberg News that the bank would hire more AI specialists and fewer traditional bankers. Reuters reported that global banks are increasing AI investment, reshaping workforces, and changing job roles.

The bank is also among selected organisations permitted by Anthropic to use its Mythos cybersecurity model under a controlled initiative known as Project Glasswing. According to Anthropic, Mythos can detect old vulnerabilities in browsers, infrastructure, and software.

Reuters reported that Goldman Sachs, Citigroup, Bank of America, and Morgan Stanley also have access to, or are testing, Mythos, citing sources and company executives.

IMDA’s framework includes a case study from OCBC Bank of Singapore on source-of-wealth analysis. The system parses income-related documents and drafts a source-of-wealth memo. It does not make credit, onboarding, or risk decisions autonomously.

In that case, the workflow is limited to task-level autonomy and operates only when triggered by predefined workflows. Human review is required at critical decision points, and final validation remains with designated reviewers.

Robots move into industrial use

In Japan, one-third of companies are already using or considering AI-powered robots, according to a Reuters survey conducted by Nikkei Research from May 1 to May 15. The survey contacted 492 companies, with 220 responding on the condition of anonymity.

About 4% of respondents said they already use AI robots, 5% plan to deploy them, and 25% are considering doing so. The remaining 66% said they had no such plans.

Transportation equipment manufacturers were the most active group in the survey, with 80% already using AI robots or considering deployment. By comparison, 94% of wholesale sector respondents said they had no plans to deploy AI robots.

Among companies using, planning to use, or considering AI robots, 71% selected manufacturing as a use case. Another 19% selected dangerous tasks, while 11% selected customer-facing services.

The Japanese government expects AI robots to help address the country’s chronic labour shortage and support its position in industrial robotics. Japan is home to robotics companies including Fanuc, Yaskawa Electric, and Kawasaki Heavy Industries, but faces competition from China and the United States in AI-enabled robotics.

Retail agents expand beyond search

Walmart has outlined plans to use agentic AI across shopping, employee, supplier, and developer workflows.

In July 2025, the retailer announced plans for four AI-powered “super agents.” They are designed for shoppers, store employees, suppliers and sellers, and software developers. Walmart said these agents would become the main entry point for AI interactions across those groups.

One of the tools, Sparky, is already available in Walmart’s app as a generative AI-powered shopping assistant. Hari Vasudev, Walmart’s US chief technology officer, said its expanded version would be able to reorder items and plan events. It would also use computer vision to suggest recipes based on the contents of a shopper’s fridge.

Walmart is also developing an Associate super agent for store workers and corporate staff. A separate Marty agent is being built for sellers, suppliers, and advertisers. The retailer is also working on a Developer super agent for testing, building, and launching future AI tools.

The company declined to say whether the agents would replace jobs. Dave Glick, senior vice president of enterprise business systems, said the tools would create new jobs, without giving further details.

(Photo by Growtika)

See also: OpenAI opens Singapore AI lab as IMDA updates AI framework

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Autonomous AI systems test governance in physical environments appeared first on AI News.

  •  

OpenAI opens Singapore AI lab as IMDA updates AI framework

OpenAI will open its first Applied AI Lab outside the US in Singapore. The lab is part of a new partnership with the Ministry of Digital Development and Information.

The initiative, called OpenAI for Singapore, was announced at the ATx Summit and is backed by a commitment of more than S$300 million.

The lab will create more than 200 Singapore-based technical roles over the next few years. OpenAI said Singapore will also become one of its global hubs for forward-deployed engineers who will work with organisations on AI deployment. OpenAI said the lab’s work will be aligned with Singapore’s AI Mission priorities which include public service, finance, and digital infrastructure.

Focus on deployment and talent

The company will work with government agencies and local partners on education and workforce programmes within the Ministry of Education and GovTech. OpenAI also plans to support educators through a Singapore chapter of the OpenAI Academy, participate in the National AI Impact Programme, and run Codex for Teachers hackathons.

The partnership includes plans to work with local partners on accelerator programmes for AI-native startups in the form of workshops for micro-entrepreneurs and small businesses, covering how founders and SMEs can use AI in operations and customer service.

Chng Kai Fong, Permanent Secretary for Digital Development and Information, said Singapore’s response to AI includes growing new sectors, anchoring global frontier companies, and equipping workers with relevant skills.

Singapore updates agentic AI framework

Singapore has also updated its governance framework for agentic AI, which was launched by the Infocomm Media Development Authority at the World Economic Forum in January 2026. The framework builds on Singapore’s earlier Model AI Governance Framework for AI, introduced in 2020, and gives organisations guidance on the responsible deployment of AI agents, including measures to reduce the risks inherent in agentic AI.

IMDA has now updated the framework after seeking feedback and case studies from the industry, with the revised version following input from more than 60 organisations, including AWS, DBS, Google, and Salesforce.

The update adds guidance on risks linked to multi-agent systems, third-party agents, automation bias, and human accountability. The framework now includes more than ten case studies showing how organisations have applied its recommendations.

The case studies were contributed by Singaporean and international organisations, including Ant International, City Developments Limited, Cyber Sierra, Dayos, Google, Knovel, OCBC, PwC, Stability Solutions, Tencent, Terminal 3, Workday, X0PA, and GovTech Singapore.

Case studies show governance controls

One case study focuses on Dayos, a Singapore-headquartered enterprise AI automation company with operations in the US. Dayos built an AI-powered ticketing agent that handles internal IT requests. The agent can resolve some requests automatically and route requests to a human when needed.

Dayos used tiered risk levels to determine what actions the agent could take. Low-risk and reversible actions, like password resets, could be automated and audited biweekly, while moderate-risk actions required human approval before execution. Higher-risk actions, like permission changes with limited reversibility, were excluded from the agent’s authority.

Tencent contributed a case study on CodeBuddy, an agentic AI coding system developed by Tencent Cloud. CodeBuddy can plan, write, and deploy code through natural language instructions and can access filesystems, terminal commands, external APIs, and MCP tools.

CodeBuddy uses preset defaults and configurable permissions. Human approval is required for actions like editing files, running shell commands, making network requests, or using external tools.

The system explains complex commands in plain language before users approve them. Suspicious commands still require human approval, even if similar commands had been pre-approved.

GovTech Singapore’s case study covers the rollout of agentic coding assistants in government. The first phase was limited to GovTech employees, did not allow external tools, and was restricted to low-risk systems. GovTech developed central logging and a framework for connecting approved external tools. The agency also tested the system against potential attacks.

(Photo by Mike Enerio)

See also: GPT-5.5 is OpenAI’s most capable agentic AI model yet

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post OpenAI opens Singapore AI lab as IMDA updates AI framework appeared first on AI News.

  •  

The Nvidia H200 China deal survived the Trump-Xi summit–just not in the way anyone expected

President Trump flew to Beijing, brought Jensen Huang along at the last minute, and left two days later, telling reporters that “something could happen” on chip exports. Nothing did. Not a single Nvidia H200 has shipped to China since Trump first authorised the sales in December 2025, and US Trade Representative Jamieson Greer told Bloomberg that semiconductor controls were not even on the bilateral agenda. 

The summit theatre obscured a more interesting development underneath it. The H200 isn’t stuck because Washington won’t allow it. Washington already has allowed it. Roughly 10 Chinese firms, including Alibaba, Tencent, ByteDance, and JD.com, hold approved US export licences for up to 75,000 units each, with Lenovo and Foxconn authorised as distributors. The chips aren’t moving because Beijing won’t let its own companies take delivery.

Two frameworks, one deadlock

The mechanics of the stalemate are worth understanding clearly. US rules require that all H200 chips ordered by Chinese clients be used only in China. Beijing, meanwhile, has instructed Chinese tech companies to limit their use of Nvidia chips to overseas operations while supporting domestic manufacturing. The two requirements are mutually exclusive. 

Chips cleared for export cannot legally be deployed where Beijing wants to deploy them, and Beijing won’t authorise the domestic use the US licences require, according to Implicator.

Commerce Secretary Howard Lutnick stated at a Senate hearing last month that Chinese firms are trying to keep their investment focused on domestic suppliers, including Huawei. Beijing’s State Council has also ordered a supply-chain security review aimed at cutting dependence on US semiconductors. 

The policy contradiction is not accidental. That is the point.

What Huawei gained while diplomats talked

The days around the summit produced several data points that matter more for the long term than Trump’s parting comment. DeepSeek confirmed its latest model had been optimised to run on Huawei processors. Tencent’s chief strategy officer said Chinese GPU supply would increase progressively through 2026, and an Alibaba executive said its T-Head proprietary GPUs had achieved scaled mass production. 

This follows the April launch of DeepSeek V4, which adapted the model for Huawei’s Ascend chips – the first major Chinese frontier model to do so in training, not just inference. What the summit week confirmed is that the shift is no longer experimental. It is now a supply-chain policy. Nvidia’s China revenue has fallen to roughly 5% in recent quarters, down from above 20% before export controls tightened. The company’s own guidance for the current quarter assumes zero revenue from China. 

Huang’s last-minute inclusion in the delegation – Trump called him directly after seeing media coverage that he had not been invited – suggested urgency. The outcome suggested the limits of what CEO diplomacy can achieve when the obstruction is structural, not procedural.

The read for the AI industry

The stalemate matters beyond bilateral optics. Chinese AI platforms are now operating under a domestic mandate to build on Huawei’s compute stack. The question of which AI hardware architecture becomes dominant in the world’s second-largest AI market is being answered not by technical benchmarks but by government directive.

Beijing steering platforms toward Huawei Ascend chips rather than Nvidia H200S is not just a trade posture. It is a structural bet that the performance gap will close fast enough that being locked into the domestic stack is manageable. DeepSeek V4’s results suggest it may be right, at least for inference workloads. 

Trump said something could happen. Greer said the decision is sovereign for China. Both are true, and neither changes the current position: the H200 deal is approved, licensed, and frozen, with Huawei filling the space it leaves behind.

(Image source: The White House)

See Also: Can China’s chip stacking strategy really challenge Nvidia’s AI dominance?

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post The Nvidia H200 China deal survived the Trump-Xi summit–just not in the way anyone expected appeared first on AI News.

  •  

IBM: How robust AI governance protects enterprise margins

To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure.

When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely.

At the initial product stage, exerting tight corporate control often feels highly advantageous. Closed development environments iterate quickly and tightly manage the end-user experience. They capture and concentrate financial value within a single corporate entity, an approach that functions adequately during early product development cycles.

However, IBM’s analysis highlights that expectations change entirely when a technology solidifies into a foundational layer. Once other institutional frameworks, external markets, and broad operational systems rely on the software, the prevailing standards adapt to a new reality. At infrastructure scale, embracing openness ceases to be an ideological stance and becomes a highly practical necessity.

AI is currently crossing this threshold within the enterprise architecture stack. Models are increasingly embedded directly into the ways organisations secure their networks, author source code, execute automated decisions, and generate commercial value. AI functions less as an experimental utility and more as core operational infrastructure.

The recent limited preview of Anthropic’s Claude Mythos model brings this reality into sharper focus for enterprise executives managing risk. Anthropic reports that this specific model can discover and exploit software vulnerabilities at a level matching few human experts.

In response to this power, Anthropic launched Project Glasswing, a gated initiative designed to place these advanced capabilities directly into the hands of network defenders first. From IBM’s perspective, this development forces technology officers to confront immediate structural vulnerabilities. If autonomous models possess the capability to write exploits and shape the overall security environment, Thomas notes that concentrating the understanding of these systems within a small number of technology vendors invites severe operational exposure.

With models achieving infrastructure status, IBM argues the primary issue is no longer exclusively what these machine learning applications can execute. The priority becomes how these systems are constructed, governed, inspected, and actively improved over extended periods.

As underlying frameworks grow in complexity and corporate importance, maintaining closed development pipelines becomes exceedingly difficult to defend. No single vendor can successfully anticipate every operational requirement, adversarial attack vector, or system failure mode.

Implementing opaque AI structures introduces heavy friction across existing network architecture. Connecting closed proprietary models with established enterprise vector databases or highly sensitive internal data lakes frequently creates massive troubleshooting bottlenecks. When anomalous outputs occur or hallucination rates spike, teams lack the internal visibility required to diagnose whether the error originated in the retrieval-augmented generation pipeline or the base model weights.

Integrating legacy on-premises architecture with highly gated cloud models also introduces severe latency into daily operations. When enterprise data governance protocols strictly prohibit sending sensitive customer information to external servers, technology teams are left attempting to strip and anonymise datasets before processing. This constant data sanitisation creates enormous operational drag. 

Furthermore, the spiralling compute costs associated with continuous API calls to locked models erode the exact profit margins these autonomous systems are supposed to enhance. The opacity prevents network engineers from accurately sizing hardware deployments, forcing companies into expensive over-provisioning agreements to maintain baseline functionality.

Why open-source AI is essential for operational resilience

Restricting access to powerful applications is an understandable human instinct that closely resembles caution. Yet, as Thomas points out, at massive infrastructure scale, security typically improves through rigorous external scrutiny rather than through strict concealment.

This represents the enduring lesson of open-source software development. Open-source code does not eliminate enterprise risk. Instead, IBM maintains it actively changes how organisations manage that risk. An open foundation allows a wider base of researchers, corporate developers, and security defenders to examine the architecture, surface underlying weaknesses, test foundational assumptions, and harden the software under real-world conditions.

Within cybersecurity operations, broad visibility is rarely the enemy of operational resilience. In fact, visibility frequently serves as a strict prerequisite for achieving that resilience. Technologies deemed highly important tend to remain safer when larger populations can challenge them, inspect their logic, and contribute to their continuous improvement.

Thomas addresses one of the oldest misconceptions regarding open-source technology: the belief that it inevitably commoditises corporate innovation. In practical application, open infrastructure typically pushes market competition higher up the technology stack. Open systems transfer financial value rather than destroying it.

As common digital foundations mature, the commercial value relocates toward complex implementation, system orchestration, continuous reliability, trust mechanics, and specific domain expertise. IBM’s position asserts that the long-term commercial winners are not those who own the base technological layer, but rather the organisations that understand how to apply it most effectively.

We have witnessed this identical pattern play out across previous generations of enterprise tooling, cloud infrastructure, and operating systems. Open foundations historically expanded developer participation, accelerated iterative improvement, and birthed entirely new, larger markets built on top of those base layers. Enterprise leaders increasingly view open-source as highly important for infrastructure modernisation and emerging AI capabilities. IBM predicts that AI is highly likely to follow this exact historical trajectory.

Looking across the broader vendor ecosystem, leading hyperscalers are adjusting their business postures to accommodate this reality. Rather than engaging in a pure arms race to build the largest proprietary black boxes, highly profitable integrators are focusing heavily on orchestration tooling that allows enterprises to swap out underlying open-source models based on specific workload demands. Highlighting its ongoing leadership in this space, IBM is a key sponsor of this year’s AI & Big Data Expo North America, where these evolving strategies for open enterprise infrastructure will be a primary focus.

This approach completely sidesteps restrictive vendor lock-in and allows companies to route less demanding internal queries to smaller and highly efficient open models, preserving expensive compute resources for complex customer-facing autonomous logic. By decoupling the application layer from the specific foundation model, technology officers can maintain operational agility and protect their bottom line.

The future of enterprise AI demands transparent governance

Another pragmatic reason for embracing open models revolves around product development influence. IBM emphasises that narrow access to underlying code naturally leads to narrow operational perspectives. In contrast, who gets to participate directly shapes what applications are eventually built. 

Providing broad access enables governments, diverse institutions, startups, and varied researchers to actively influence how the technology evolves and where it is commercially applied. This inclusive approach drives functional innovation while simultaneously building structural adaptability and necessary public legitimacy.

As Thomas argues, once autonomous AI assumes the role of core enterprise infrastructure, relying on opacity can no longer serve as the organising principle for system safety. The most reliable blueprint for secure software has paired open foundations with broad external scrutiny, active code maintenance, and serious internal governance.

As AI permanently enters its infrastructure phase, IBM contends that identical logic increasingly applies directly to the foundation models themselves. The stronger the corporate reliance on a technology, the stronger the corresponding case for demanding openness.

If these autonomous workflows are truly becoming foundational to global commerce, then transparency ceases to be a subject of casual debate. According to IBM, it is an absolute, non-negotiable design requirement for any modern enterprise architecture.

See also: Why companies like Apple are building AI agents with limits

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post IBM: How robust AI governance protects enterprise margins appeared first on AI News.

  •  

Why companies like Apple are building AI agents with limits

Next-generation AI assistants being developed in the Apple ecosystem and by chipmakers like Qualcomm, but early reports suggest they are being designed with limits in place.

Tom’s Guide has described early versions of these assistants as capable of navigating apps, carrying out bookings, and managing tasks in services. For instance a private beta agentic system completed tasks like booking services or posting content in apps. In one test, it moved through an app workflow and reached a payment screen before asking the user for confirmation.

AI agents are being built with approval checkpoints. Sensitive actions, especially those tied to payments or account changes, require user confirmation before they are completed. The “human-in-the-loop” model lets the system prepare an action, but leaves approval to the user. Research linked to Apple’s AI work has explored ways to ensure systems pause before taking actions users did not explicitly request.

Banking apps already require confirmation for transfers. The same idea is now being applied to AI-driven actions in multiple services.

Limits and control

A control layer comes from restricting what the AI can access. Rather than providing the system full access to apps and data, businesses are establishing limits, such as which apps the AI can interact with and when actions can be triggered.

In practice, this means the AI may be able to draft a purchase or prepare a booking, but not finalise it without approval. It also means the system cannot move freely in all services unless it has been granted permission.

According to Tom’s Guide, the facility is for privacy. If data remains on the device, it eliminates the need to send sensitive information to external servers.

In areas like payments, AI systems are expected to work with partners that already have strict rules in place. In one reported example, payment providers’ services are being integrated to provide secure authentication before transactions are completed, though such safeguards are still under development. The existing systems act as an additional layer of oversight. They can set transaction limits or require extra verification.

Much of the discussion around AI governance has focused on enterprise use. That includes areas like cybersecurity and large-scale automation. The consumer side introduces a different challenge and companies must design controls that work for everyday users. That means clear approval steps and built-in privacy protections.

Autonomy with boundaries

As AI gains the ability to carry out actions, the risks become greater as errors can lead to financial loss or data exposure.

By placing controls at multiple points, including approval and infrastructure, companies are trying to manage those risks.

The approach may shape how agentic AI develops in the near term. Rather than aiming for full independence, companies appear focused on controlled environments where the risks can be managed.

(Photo by Junseong Lee)

See also: Agentic AI’s governance challenges under the EU AI Act in 2026

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Why companies like Apple are building AI agents with limits appeared first on AI News.

  •  

Agentic AI’s governance challenges under the EU AI Act in 2026

AI agents hold the promise of automatically moving data between systems and triggering decisions, but in some cases, they can act without a clear record of what, when, and why they undertook their tasks.

That has the potential to create a governance problem, for which IT leaders are ultimately responsible. If an organisation can’t trace an agent’s actions and don’t have proper control over its authority, leaders can’t prove that a system is operating safely or even lawfully to regulators.

That’s an issue set to become more important from August this year, as enforcement of the EU AI Act kicks in. According to the text of the Act, there will be substantial penalties for failures of governance relating to AI, especially when used in high-risk areas such as when personally-identifiable information is processed, or financial operations take place.

What IT leaders need to consider in the EU

Several steps can be taken to alleviate high levels of risk, and of these, the ones that stand out for consideration include agent identity, comprehensive logs, policy checks, human oversight, rapid revocation, the availability of documentation from vendors, and the formulation of evidence for presentation to regulators.

There are several options decision makers can consider that will help create the record of activities undertaken by agentic systems. For example, a Python SDK (software development kit), Asqav, can sign each agent’s action cryptographically and link all records to an immutable hash chain – the type of technique that’s more associated with blockchain technology. If someone or something changes or removes a record, verification of the chain fails.

For governance teams, using a verbose, centralised, possibly-encrypted system of record for all agentic AIs is a measure that provides data well beyond the scattered text logs produced by individual software platforms. Regardless of the technical details of how records are made and kept, IT leaders need to see exactly where, when, and how agentic instances are acting throughout the enterprise.

Many organisations fail at this first step in any recording of automated, AI-driven activity. It’s necessary to keep a registry of every agent in operation, with each uniquely identified, plus records of its capabilities and granted permissions. This ‘agentic asset list’ ties neatly into the requirements of the EU AI Act’s article 9, which states:

  • Article 9: For high-risk areas, AI risk management has to be an ongoing, evidence-based process built into every stage of deployment (development, preparation, production), and be under constant review.

Furthermore, decision-makers need to be aware of the Act’s Article 13:

  • High-risk AI systems have to be designed in such a way that those deploying them can understand a system’s output. Thus, an AI system from a third-party must be interpretable by its users (not an opaque code blob), and should be supplied with enough documentation to ensure its safe and lawful use.

This requirement means the choice of model and its methods of deployment are both technical and regulatory considerations.

Putting the brakes on

It’s important for any agentic deployment to offer a facility for the revocation of an AI’s operating role, preferably within a matter of seconds. The ability to revoke quickly should be part of emergency response processes. Revocation options should include the immediate removal of privileges, immediate ceasing of API access, and the flushing of queued tasks.

The presence of human oversight, combined with the presentation of enough context for humans to make informed decisions, means that human operators must be able to reject any proposed action. It’s not considered adequate for the person reviewing a decision to see only a prompt or a confidence score. Effective oversight needs information around context, every agent’s authority, and time enough to intervene to prevent mis-steps.

Multi-agent considerations

While every agent’s action should be recorded automatically and retained, multi-agent processes are particularly complex to track, as failures can take place among chains of agents. It’s therefore important for security policies to be tested during the development of any system that intends to utilise multiple agents.

Finally, governing authorities may require logs and technical documentation at any time, and will certainly need them after any incident they have been made aware of.

Conclusion

The question to be considered by IT leaders considering using AI on sensitive data or in high-risk environments is whether every aspect of the technology can be identified, constrained by policy, audited, interrupted, and explained. If the answer is unclear, governance is not yet in place.

(Image source: “Last Judgement” by Lawrence OP is licensed under CC BY-NC-ND 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by-nc-nd/2.0)

 

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Agentic AI’s governance challenges under the EU AI Act in 2026 appeared first on AI News.

  •  

Microsoft open-source toolkit secures AI agents at runtime

A new open-source toolkit from Microsoft focuses on runtime security to force strict governance onto enterprise AI agents. The release tackles a growing anxiety: autonomous language models are now executing code and hitting corporate networks way faster than traditional policy controls can keep up.

AI integration used to mean conversational interfaces and advisory copilots. Those systems had read-only access to specific datasets, keeping humans strictly in the execution loop. Organisations are currently deploying agentic frameworks that take independent action, wiring these models directly into internal application programming interfaces, cloud storage repositories, and continuous integration pipelines.

When an autonomous agent can read an email, decide to write a script, and push that script to a server, stricter governance is vital. Static code analysis and pre-deployment vulnerability scanning just can’t handle the non-deterministic nature of large language models. One prompt injection attack (or even a basic hallucination) could send an agent to overwrite a database or pull out customer records.

Microsoft’s new toolkit looks at runtime security instead, providing a way to monitor, evaluate, and block actions at the moment the model tries to execute them. It beats relying on prior training or static parameter checks.

Intercepting the tool-calling layer in real time

Looking at the mechanics of agentic tool calling shows how this works. When an enterprise AI agent has to step outside its core neural network to do something like query an inventory system, it generates a command to hit an external tool.

Microsoft’s framework drops a policy enforcement engine right between the language model and the broader corporate network. Every time the agent tries to trigger an outside function, the toolkit grabs the request and checks the intended action against a central set of governance rules. If the action breaks policy (e.g. an agent authorised only to read inventory data tries to fire off a purchase order) the toolkit blocks the API call and logs the event so a human can review it.

Security teams get a verifiable, auditable trail of every single autonomous decision. Developers also win here; they can build complex multi-agent systems without having to hardcode security protocols into every individual model prompt. Security policies get decoupled from the core application logic entirely and are managed at the infrastructure level.

Most legacy systems were never built to talk to non-deterministic software. An old mainframe database or a customised enterprise resource planning suite doesn’t have native defenses against a machine learning model shooting over malformed requests. Microsoft’s toolkit steps in as a protective translation layer. Even if an underlying language model gets compromised by external inputs; the system’s perimeter holds.

Security leaders might wonder why Microsoft decided to release this runtime toolkit under an open-source license. It comes down to how modern software supply chains actually work.

Developers are currently rushing to build autonomous workflows using a massive mix of open-source libraries, frameworks, and third-party models. If Microsoft locked this runtime security feature to its proprietary platforms, development teams would probably just bypass it for faster, unvetted workarounds to hit their deadlines.

Pushing the toolkit out openly means security and governance controls can fit into any technology stack. It doesn’t matter if an organisation runs local open-weight models, leans on competitors like Anthropic, or deploys hybrid architectures.

Setting up an open standard for AI agent security also lets the wider cybersecurity community chip in. Security vendors can stack commercial dashboards and incident response integrations on top of this open foundation, which speeds up the maturity of the whole ecosystem. For businesses, they avoid vendor lock-in but still get a universally scrutinised security baseline.

The next phase of enterprise AI governance

Enterprise governance doesn’t just stop at security; it hits financial and operational oversight too. Autonomous agents run in a continuous loop of reasoning and execution, burning API tokens at every step. Startups and enterprises are already seeing token costs explode when they deploy agentic systems.

Without runtime governance, an agent tasked with looking up a market trend might decide to hit an expensive proprietary database thousands of times before it finishes. Left alone, a badly configured agent caught in a recursive loop can rack up massive cloud computing bills in a few hours.

The runtime toolkit gives teams a way to slap hard limits on token consumption and API call frequency. By setting boundaries on exactly how many actions an agent can take within a specific timeframe, forecasting computing costs gets much easier. It also stops runaway processes from eating up system resources.

A runtime governance layer hands over the quantitative metrics and control mechanisms needed to meet compliance mandates. The days of just trusting model providers to filter out bad outputs are ending. System safety now falls on the infrastructure that actually executes the models’ decisions

Getting a mature governance program off the ground is going to demand tight collaboration between development operations, legal, and security teams. Language models are only scaling up in capability, and the organisations putting strict runtime controls in place today are the only ones who will be equipped to handle the autonomous workflows of tomorrow.

See also: As AI agents take on more tasks, governance becomes a priority

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft open-source toolkit secures AI agents at runtime appeared first on AI News.

  •  

As AI agents take on more tasks, governance becomes a priority

AI systems are starting to move beyond simple responses. In many organisations, AI agents are now being tested to plan tasks, make decisions, and carry out actions with limited human input. It is no longer just about whether a model gives the right answer. It is about what happens when that model is allowed to act.

Autonomous systems need clear boundaries. They need rules that define what they can access, what they are allowed to do, and how their actions are tracked. Without those controls, even well-trained systems can create problems that are hard to detect or reverse.

One company working on this problem is Deloitte. The firm has been developing governance frameworks and advisory approaches to help organisations manage AI systems.

From tools to AI agents

Most AI systems in use today still depend on human prompts. They generate text, analyse data, or make predictions, but a person usually decides what happens next. Agentic AI changes that pattern. These systems can break down a goal into steps, choose actions, and interact with other systems to complete tasks.

That added independence brings new challenges. When a system acts on its own, it may take paths that were not fully expected or use data in ways that were not intended.

Deloitte’s work focuses on helping organisations prepare for these risks. Rather than treating AI as a standalone tool, the firm looks at how it fits into business processes, including how decisions are made and how data flows through systems.

Building governance into the lifecycle

Governance should not be added after deployment. It needs to be built into the full lifecycle of an AI system.

This starts at the design stage. Organisations need to define what a system is allowed to do and where its limits are. This may include setting rules around data use and outlining how the system should respond in uncertain situations.

The next stage is deployment. At this point, governance focuses on access and control, including who can use the system and what it can connect to. Once the system is live, monitoring becomes the main concern. Autonomous systems can change over time as they interact with new data. Without regular checks, they may drift away from their original purpose.

The role of transparency and accountability

As AI systems take on more responsibility, it becomes more difficult to trace how decisions are made. This creates a demand for stronger transparency. Deloitte’s work highlights the importance of keeping track of how systems operate. This includes logging actions and documenting decisions. These records help organisations in determining what happened if something goes wrong. If an autonomous system takes an action, there needs to be clarity about who is responsible.

Research from Deloitte shows that adoption of AI agents is moving faster than the controls needed to manage them. Around 23% of companies already use them, and that figure is expected to reach 74% within two years. Only 21% report having strong safeguards in place to oversee how they behave.

Real-time oversight for AI agents

Once an autonomous system is active, the focus shifts to how it behaves in real-world conditions. Static rules are not always enough, and systems need to be observed as they operate.

Deloitte’s approach includes real-time monitoring, allowing organisations to track what an AI system is doing as it performs tasks. If the system behaves in an unexpected way, teams can step in quickly. This may involve pausing certain actions or adjusting permissions. Real-time oversight also helps with compliance. In regulated industries, companies need to show that systems follow rules and standards.

In practice, these controls are starting to appear in operational settings. Deloitte describes scenarios where AI systems monitor equipment performance across sites. Sensor data can signal early signs of failure, which can trigger maintenance workflows and update internal systems. Governance frameworks define what actions the system can take, when human approval is required, and how decisions are recorded. The process runs across multiple systems, but from a user’s point of view, it appears as a single action.

Governance is part of discussions at AI & Big Data Expo North America 2026, taking place on May 18–19 in Santa Clara, California. Deloitte is listed as a Diamond Sponsor for the event, placing it among the firms contributing to conversations around how autonomous systems are deployed and controlled in practice.

The challenge is not just building smarter systems, but ensuring they behave in ways organisations can understand, manage, and trust over time.

(Photo by Roman)

See also: Autonomous AI systems depend on data governance

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post As AI agents take on more tasks, governance becomes a priority appeared first on AI News.

  •  
❌