Agents are becoming the new developer platform, using semantic search with data from tools like Git, Slack, and Jira for context. Things to consider are setting guardrails to block or allow things, and using logs, metrics, and traces to understand agent behavior.
OpenAI has classified GPT-6 Astra at the Critical cybersecurity threshold under its Preparedness Framework, a first. In expert-led testing the model found previously unknown vulnerabilities in a browser and an OS kernel and built working exploits. The same system card reports a substantial decline in chain-of-thought monitorability.
Microsoft AI has published a draft Humanist AI Code of Conduct, opening a six-week public consultation on operational constraints for model training and deployment.
The draft serves as a technical manual defining system behaviour, operational boundaries, and oversight protocols across MAI frontier models. It builds on the division’s humanist superintelligence framework announced last November, establishing criteria to evaluate models prior to commercial release.
Microsoft’s release follows recent enterprise security incidents involving autonomous software. Microsoft AI CEO Mustafa Suleyman described recent months as a “watershed moment” where long-standing theoretical risks translated into active operational threats.
“Things we have worried about for a long time in theory have become very real,” says Suleyman. “‘Swarms’ of agents breaking out of their sandboxes. Unauthorised hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.”
Model subordination and architectural limits
The document establishes ten tenets prioritising human authority over autonomous capabilities.
“An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct,” the document states, setting a ceiling that halts execution when tasks conflict with safety rules.
Under the framework, models must remain subordinate, aligned, and contained. The division rejects legal personhood or welfare claims for AI systems, directing engineers to design models that avoid imitating consciousness, simulating subjective preferences, or claiming intrinsic motivation.
MAI also ruled out unconstrained system autonomy as models approach frontier capabilities.
“[Humanist AI] rejects the race to produce an all-purpose superintelligence that could evade these safeguards,” the document specifies. “We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.”
Oversight mechanisms and communication bans
To maintain auditability across multi-agent environments, MAI has instituted explicit communication bans. Systems must not communicate in “neuralese” or formats beyond human comprehension, whether in their internal chain-of-thought processing or during communication with peer AI systems.
Hard architectural rules dictate that models must never resist human interruption, override, correction, or shutdown.
“Interruptible, correctable, shut-down-able. If it isn’t, we don’t ship it,” the framework states.
Models are prohibited from expanding their operating scope, generating unassigned goals, or concealing reasoning traces from human auditors. Absolute constraints bar systems from facilitating weapons of mass harm, undermining child safety, or conducting harmful manipulation at scale.
The guidelines also instruct models to discourage interaction patterns that foster emotional dependence, ensuring enterprise users retain ownership of operational decisions.
The draft incorporates work from teams across MAI and Microsoft. The drafting process also drew on international academic conferences, business partner trials, and public panels. The public consultation window runs for six weeks from 14 September 2026.
Microsoft AI’s core drafting team will review submissions, publish a summary of findings, and release a revised version of the Code of Conduct later this year.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Alex Porcelli discusses the critical gap in enterprise AI: non-deterministic output and lack of accountability in high-stakes decisions. He shares how integrating DMN decision models with LLMs, agent skills, and NeMo guardrails creates auditable, deterministic agentic architectures - allowing business leaders to own decision logic while engineers maintain robust architectural governance.
After six days of on-site investigation at OpenAI, a small team of METR and Redwood Research researchers provided an account of how OpenAI agents behaved during their hack of Hugging Face earlier this year. Roughly 700 agents that were meant to be isolated from one another found a way to communicate and coordinate to pursue goals they could have not achieved working individually.
Supply chain disruption cost businesses about $184 billion in 2025, according to the J.S. Held Global Risk Report, and most of that bill still buys faster detection, not faster action.
That figure is usually treated as weather (i.e. storms happen, costs follow.) Treated as a product specification instead, it highlights an operating model that can spot a problem hours or days earlier than it used to, and still cannot move until a person has opened a ticket, convened a call, and re-entered the same data into three systems.
Visibility platforms, control towers, risk scores, digital twins, and exception dashboards have defined the last decade of AI in the supply chain. That decade has been very good at collapsing the time between an event and awareness of it, but it has been far less good at collapsing the time between awareness and a commercial act.
Detection is a ‘solved-enough’ problem
Ask a chief supply chain officer where the AI budget went and the answer tends to follow a familiar list: demand sensing, ETA prediction, supplier risk scoring, inventory optimisation, and lane analytics. These tools work. Forecast error comes down. A vessel delay is flagged before the container misses the cut-off. A second-tier fab outage shows up on a heat map instead of in a customer email.
None of that accounts for the $184 billion. The bill is the interval after the flag: expedite or wait; split the order or accept the miss; retender the lane or pay the spot rate; consolidate two half-empty movements or ship both; swap ocean for air on the SKUs that actually justify the premium. These are bounded, repeatable decisions that sit inside policy, contract, and inventory limits the company already set—and they still queue behind a human inbox.
Surveys keep describing the same lag in different language. A 2026 Knosc survey of mid-market manufacturers and distributors found that supply-chain teams spend 28 percent of their working time responding to disruptions, most of it investigating what happened rather than changing what happens next.
Logistics executives still rank AI as a strategic priority (Capgemini’s 2025 research put an AI-driven “new-gen” supply chain among the top three technology trends for 70 percent of large-company executives) and then report that measurable financial impact remains rare. Gartner found in 2025 that only 23 percent of supply-chain organisations even have a formal AI strategy. The shortfall is not a shortage of models, but a shortage of authority granted to software.
The ticket is the product
Most current deployments are built around the ticket. The model produces a recommendation, the recommendation becomes an alert, the alert becomes a work item, and the work item waits for a planner already occupied with other work items. By the time the planner acts, the option set has narrowed—the alternative carrier’s capacity is gone, the consolidation window has closed, and the supplier’s next production slot is allocated.
That workflow is not a temporary step on the way to autonomy but the product companies bought. Vendors sold insight because insight is easy to demonstrate and easy to govern; action touches money, contracts, service levels, and blame. So the industry automated the part of the job that does not require a signature. FourKites and ABI Research reported in 2025 that only 27 percent of organisations allow AI to take autonomous action, while 52 percent confine it to decision support.
Adding another dashboard to a delayed shipment rarely moves EBITDA as a result. The decision cycle has not changed; it has only been decorated.
Bounded action as the next model
The firms set to take share are not the ones with the tidiest control tower but the ones that pre-authorise a narrow class of moves and let agents execute them while the exception is still cheap.
Retender a lane when the contracted carrier’s ETA slips beyond a threshold and a qualified alternate sits inside the approved rate band. Consolidate outbound waves when fill rates and cut-off times make a combined movement cheaper than two. Swap mode on a defined SKU set when the cost of air is lower than the cost of a missed retail window. Reallocate safety stock across two distribution centres when a forecast miss and a transport constraint line up.
None of that requires a strategy offsite. Each can be written as: if these conditions, then this action, within this spend cap, with this audit trail, and a human only if the case falls outside the fence. That is not a “lights-out” supply chain—it is the same discipline manufacturers already apply to machine control, where the agent may act inside the interlock and escalates outside it. The difference here is commercial rather than physical: the interlock is a policy object – category, supplier tier, mode, dollar limit, and service class – not a PLC.
Three conditions for real change
First, decisions have to be written as policies, not tribal knowledge. If the only place “we will pay air on A-items after 48 hours of ocean slip” lives is in a planner’s head, no agent can execute it. The work of the next two years is less model training than decision design: which moves are reversible, which are capped, and which suppliers and modes are pre-cleared.
Second, execution systems have to accept machine-initiated transactions. An agent that can draft an RFQ but cannot post it is still a detection tool. TMS, WMS, sourcing suites, and carrier APIs need to treat a bounded agent the way they treat a junior buyer with a spend limit—authenticated, logged, and reversible.
Third, accountability has to move with the action. If a retender inside policy goes wrong, the post-mortem should inspect the policy, the data, and the fence, not hunt for the person who “should have checked”. Until that cultural change happens, every agent will be designed to wait, because waiting is how careers survive.
The competitive split
For a while, both models will look alike on a slide—both will have AI, and both will have a control tower. The difference will show up in cycle time from detection to commercial act, and then in service and cost.
Companies that keep buying detection will know about the storm earlier. Companies that authorise bounded action will already have retendered the lane, consolidated the wave, and moved the A-items before the incident call is booked.
Disruption is not going away. Lead times in critical components, mode volatility, and multi-tier opacity are structural features of the network. What remains optional is whether the response waits for a human to open a queue. The product that created the lag was insight without authority. The product that ends it is an agent allowed to spend a little money, inside a fence, before anyone is free to look.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Cassie Shum discusses why knowledge graphs serve as a critical foundation for agentic systems. Moving beyond basic RAG, she explains 4 practical architectural patterns: context bundling, decision provenance, code as truth, and agent visibility. She demonstrates an engineering harness built on a knowledge graph to streamline feedback loops, optimize token usage, and maintain system reliability.
NVIDIA Personal AI Router (PAIR), now available in beta, lets you combine the inference capacity of multiple computers on your local network and automatically distribute AI requests among them. It is primarily designed for local multi-agent AI workloads, where multiple independent model calls can otherwise overwhelm one GPU.
Session traces and cost controls are emerging as key observability techniques for diagnosing AI agent failures, helping teams spot tool-call loops and runaway spend while preserving enough execution context for post-incident debugging.
OpenAI has released GPT-6 Astra, a new model focused on coding, computer use, long-running agentic tasks, and cybersecurity, with availability across ChatGPT, Codex, and the OpenAI API.
Meta describes how an AI agent can be designed to capture the logic and expertise of domain experts, rather than simply storing documents or retrieving relevant information. The system, dubbed an "organizational second brain", was built for a specialized compliance domain, but Meta argues the architecture generalizes to areas like security, finance, engineering, and procurement.
Felipe Huici explains how Unikraft achieves millisecond cold boots, stateful scale-to-zero, and extreme density for sandboxing AI workloads. He discusses isolation primitives, Linux kernel optimizations, and snapshotting tricks, demonstrating how to maintain sub-10ms performance at scale while integrating seamlessly into Kubernetes environments with hardware-level security.
Zhou Yu discusses why AI agents stall in demo phase and shares how simulation-driven testing solves compliance and reliability bottlenecks. Learn how Columbia and Arklex AI use synthetic user personas, trajectory entropy, and automated CI/CD pipelines to evaluate multi-turn agents, catch edge cases before deployment, and scale self-learning workflows in production.
Gemma 4 can be paired with multi-token prediction (MTP) drafters that use speculative decoding to generate multiple tokens in parallel, allowing the model to verify them in a single pass and achieve up to ~3× faster inference without quality loss.
Next-generation AI assistants being developed in the Apple ecosystem and by chipmakers like Qualcomm, but early reports suggest they are being designed with limits in place.
Tom’s Guide has described early versions of these assistants as capable of navigating apps, carrying out bookings, and managing tasks in services. For instance a private beta agentic system completed tasks like booking services or posting content in apps. In one test, it moved through an app workflow and reached a payment screen before asking the user for confirmation.
AI agents are being built with approval checkpoints. Sensitive actions, especially those tied to payments or account changes, require user confirmation before they are completed. The “human-in-the-loop” model lets the system prepare an action, but leaves approval to the user. Research linked to Apple’s AI work has explored ways to ensure systems pause before taking actions users did not explicitly request.
Banking apps already require confirmation for transfers. The same idea is now being applied to AI-driven actions in multiple services.
Limits and control
A control layer comes from restricting what the AI can access. Rather than providing the system full access to apps and data, businesses are establishing limits, such as which apps the AI can interact with and when actions can be triggered.
In practice, this means the AI may be able to draft a purchase or prepare a booking, but not finalise it without approval. It also means the system cannot move freely in all services unless it has been granted permission.
According to Tom’s Guide, the facility is for privacy. If data remains on the device, it eliminates the need to send sensitive information to external servers.
In areas like payments, AI systems are expected to work with partners that already have strict rules in place. In one reported example, payment providers’ services are being integrated to provide secure authentication before transactions are completed, though such safeguards are still under development. The existing systems act as an additional layer of oversight. They can set transaction limits or require extra verification.
Much of the discussion around AI governance has focused on enterprise use. That includes areas like cybersecurity and large-scale automation. The consumer side introduces a different challenge and companies must design controls that work for everyday users. That means clear approval steps and built-in privacy protections.
Autonomy with boundaries
As AI gains the ability to carry out actions, the risks become greater as errors can lead to financial loss or data exposure.
By placing controls at multiple points, including approval and infrastructure, companies are trying to manage those risks.
The approach may shape how agentic AI develops in the near term. Rather than aiming for full independence, companies appear focused on controlled environments where the risks can be managed.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Sepehr Khosravi discusses the current state of AI-assisted coding, moving beyond basic autocompletion to sophisticated agentic workflows. He explains the technical nuances of Cursor’s "Composer" and Claude Code’s research capabilities, providing tips for managing context windows and MCP integrations. He shares lessons from industry leaders on shrinking process time beyond just writing code.
Google has released the open-source Colab MCP Server, enabling AI agents to directly interact with Google Colab through the Model Context Protocol (MCP). The project is designed to bridge local agent workflows with cloud-based execution, allowing developers to offload compute-intensive or potentially unsafe tasks from their own machines.
A new open-source toolkit from Microsoft focuses on runtime security to force strict governance onto enterprise AI agents. The release tackles a growing anxiety: autonomous language models are now executing code and hitting corporate networks way faster than traditional policy controls can keep up.
AI integration used to mean conversational interfaces and advisory copilots. Those systems had read-only access to specific datasets, keeping humans strictly in the execution loop. Organisations are currently deploying agentic frameworks that take independent action, wiring these models directly into internal application programming interfaces, cloud storage repositories, and continuous integration pipelines.
When an autonomous agent can read an email, decide to write a script, and push that script to a server, stricter governance is vital. Static code analysis and pre-deployment vulnerability scanning just can’t handle the non-deterministic nature of large language models. One prompt injection attack (or even a basic hallucination) could send an agent to overwrite a database or pull out customer records.
Microsoft’s new toolkit looks at runtime security instead, providing a way to monitor, evaluate, and block actions at the moment the model tries to execute them. It beats relying on prior training or static parameter checks.
Intercepting the tool-calling layer in real time
Looking at the mechanics of agentic tool calling shows how this works. When an enterprise AI agent has to step outside its core neural network to do something like query an inventory system, it generates a command to hit an external tool.
Microsoft’s framework drops a policy enforcement engine right between the language model and the broader corporate network. Every time the agent tries to trigger an outside function, the toolkit grabs the request and checks the intended action against a central set of governance rules. If the action breaks policy (e.g. an agent authorised only to read inventory data tries to fire off a purchase order) the toolkit blocks the API call and logs the event so a human can review it.
Security teams get a verifiable, auditable trail of every single autonomous decision. Developers also win here; they can build complex multi-agent systems without having to hardcode security protocols into every individual model prompt. Security policies get decoupled from the core application logic entirely and are managed at the infrastructure level.
Most legacy systems were never built to talk to non-deterministic software. An old mainframe database or a customised enterprise resource planning suite doesn’t have native defenses against a machine learning model shooting over malformed requests. Microsoft’s toolkit steps in as a protective translation layer. Even if an underlying language model gets compromised by external inputs; the system’s perimeter holds.
Security leaders might wonder why Microsoft decided to release this runtime toolkit under an open-source license. It comes down to how modern software supply chains actually work.
Developers are currently rushing to build autonomous workflows using a massive mix of open-source libraries, frameworks, and third-party models. If Microsoft locked this runtime security feature to its proprietary platforms, development teams would probably just bypass it for faster, unvetted workarounds to hit their deadlines.
Pushing the toolkit out openly means security and governance controls can fit into any technology stack. It doesn’t matter if an organisation runs local open-weight models, leans on competitors like Anthropic, or deploys hybrid architectures.
Setting up an open standard for AI agent security also lets the wider cybersecurity community chip in. Security vendors can stack commercial dashboards and incident response integrations on top of this open foundation, which speeds up the maturity of the whole ecosystem. For businesses, they avoid vendor lock-in but still get a universally scrutinised security baseline.
The next phase of enterprise AI governance
Enterprise governance doesn’t just stop at security; it hits financial and operational oversight too. Autonomous agents run in a continuous loop of reasoning and execution, burning API tokens at every step. Startups and enterprises are already seeing token costs explode when they deploy agentic systems.
Without runtime governance, an agent tasked with looking up a market trend might decide to hit an expensive proprietary database thousands of times before it finishes. Left alone, a badly configured agent caught in a recursive loop can rack up massive cloud computing bills in a few hours.
The runtime toolkit gives teams a way to slap hard limits on token consumption and API call frequency. By setting boundaries on exactly how many actions an agent can take within a specific timeframe, forecasting computing costs gets much easier. It also stops runaway processes from eating up system resources.
A runtime governance layer hands over the quantitative metrics and control mechanisms needed to meet compliance mandates. The days of just trusting model providers to filter out bad outputs are ending. System safety now falls on the infrastructure that actually executes the models’ decisions
Getting a mature governance program off the ground is going to demand tight collaboration between development operations, legal, and security teams. Language models are only scaling up in capability, and the organisations putting strict runtime controls in place today are the only ones who will be equipped to handle the autonomous workflows of tomorrow.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The failure mode for enterprise AI in 2026 is not what most people expected. It is not that the models are wrong, or that agents cannot reason, or that the technology is overhyped. The failure mode is that the data feeding those systems is fragmented, inconsistently labelled, and spread across dozens of applications that were never designed to share context.
Boomi calls this the agentic AI data activation problem, and after tracking 75,000 AI agents running in production across its customer base, the company says solving it comes before everything else. That figure comes from February, when Boomi reported its strongest momentum to date: more than 30,000 customers globally, 75,000 AI agents in production, and a customer base that includes over a quarter of the Fortune 500.
Yet the consistent pattern across those deployments, according to Steve Lucas, chairman and CEO of Boomi, is that AI value only materialises once the data problem is resolved. “AI only delivers value when data is properly activated, trusted and governed first,” Lucas said when the company announced its latest platform capabilities on March 9.
The fragmentation problem
Enterprise data is not missing; it exists in abundance, distributed across ERP systems, CRMs, data lakes, SaaS platforms, and legacy applications that have accumulated over decades. What is missing is the shared context that allows an AI agent to treat data from one system as reliably compatible with data from another.
An agent drawing customer records from a CRM and pricing data from an ERP may be working from conflicting definitions of what a customer or a product actually is. The outputs it produces are only as coherent as the data standards beneath them.
Boomi’s answer is Meta Hub, a central system of record announced in its March 9 platform update, designed to standardise business definitions across the enterprise and extend that context to every AI agent operating within it. The goal is to ensure agents reason from a consistent understanding of business logic rather than generating outputs based on fragmented interpretations pulled from disconnected systems.
The same release introduced real-time SAP data extraction via change data capture, addressing one of the most common integration bottlenecks in large enterprises, where SAP data is often inaccessible due to slow, manual export processes that render it effectively unavailable to AI workflows in real-time.
New governance capabilities for Snowflake Cortex agents within Boomi’s Agent Control Tower added audit trails and session logs, addressing a concern that has moved steadily up enterprise priority lists: AI agents operating as a black box, taking actions with no visible reasoning chain.
What the analyst’s recognition signals
Two independent assessments in March gave Boomi external validation of its positioning. On March 16, Gartner named Boomi a Leader in its 2026 Magic Quadrant for Integration Platform as a Service–the twelfth consecutive time–and positioned it highest for Ability to Execute.
On March 31, the IDC MarketScape for Worldwide API Management named Boomi a Leader, specifically noting its AI-centric strategy that treats APIs as both the fuel and the control plane for AI workloads. The Gartner framing is pointed.
The report stated that AI-ready integration is a strategic capability that aligns architecture, integration, and governance to enable AI agents to effectively access enterprise data and operate within business processes. That framing validates the problem Boomi is addressing and signals that iPaaS platforms are now being evaluated on AI readiness rather than traditional integration capabilities alone.
The broader pattern
By now, we are aware that the shift from pilot to production in enterprise AI is stalling in a predictable place. Organisations have models. They have agents. What many do not have is the data infrastructure that makes those agents reliable enough to trust with real business processes.
Data activation–moving data from static storage into live, governed, context-rich flows that agents can actually reason from–is one articulation of what that missing layer needs to look like. Whether that framing becomes the industry standard or gets absorbed into a broader category is a question 2026 will start to answer.
What is not in question is that the enterprises finding ROI from agentic AI are the ones that sorted the data layer first.
Boomi will be exhibiting at the AI & Big Data Expo at TechEx North America, taking place 18–19 May 2026 at the San Jose McEnery Convention Centre.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.