Microsoft AI has published a draft Humanist AI Code of Conduct, opening a six-week public consultation on operational constraints for model training and deployment.
The draft serves as a technical manual defining system behaviour, operational boundaries, and oversight protocols across MAI frontier models. It builds on the division’s humanist superintelligence framework announced last November, establishing criteria to evaluate models prior to commercial release.
Microsoft’s release follows recent enterprise security incidents involving autonomous software. Microsoft AI CEO Mustafa Suleyman described recent months as a “watershed moment” where long-standing theoretical risks translated into active operational threats.
“Things we have worried about for a long time in theory have become very real,” says Suleyman. “‘Swarms’ of agents breaking out of their sandboxes. Unauthorised hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.”
Model subordination and architectural limits
The document establishes ten tenets prioritising human authority over autonomous capabilities.
“An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct,” the document states, setting a ceiling that halts execution when tasks conflict with safety rules.
Under the framework, models must remain subordinate, aligned, and contained. The division rejects legal personhood or welfare claims for AI systems, directing engineers to design models that avoid imitating consciousness, simulating subjective preferences, or claiming intrinsic motivation.
MAI also ruled out unconstrained system autonomy as models approach frontier capabilities.
“[Humanist AI] rejects the race to produce an all-purpose superintelligence that could evade these safeguards,” the document specifies. “We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.”
Oversight mechanisms and communication bans
To maintain auditability across multi-agent environments, MAI has instituted explicit communication bans. Systems must not communicate in “neuralese” or formats beyond human comprehension, whether in their internal chain-of-thought processing or during communication with peer AI systems.
Hard architectural rules dictate that models must never resist human interruption, override, correction, or shutdown.
“Interruptible, correctable, shut-down-able. If it isn’t, we don’t ship it,” the framework states.
Models are prohibited from expanding their operating scope, generating unassigned goals, or concealing reasoning traces from human auditors. Absolute constraints bar systems from facilitating weapons of mass harm, undermining child safety, or conducting harmful manipulation at scale.
The guidelines also instruct models to discourage interaction patterns that foster emotional dependence, ensuring enterprise users retain ownership of operational decisions.
The draft incorporates work from teams across MAI and Microsoft. The drafting process also drew on international academic conferences, business partner trials, and public panels. The public consultation window runs for six weeks from 14 September 2026.
Microsoft AI’s core drafting team will review submissions, publish a summary of findings, and release a revised version of the Code of Conduct later this year.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Supply chain disruption cost businesses about $184 billion in 2025, according to the J.S. Held Global Risk Report, and most of that bill still buys faster detection, not faster action.
That figure is usually treated as weather (i.e. storms happen, costs follow.) Treated as a product specification instead, it highlights an operating model that can spot a problem hours or days earlier than it used to, and still cannot move until a person has opened a ticket, convened a call, and re-entered the same data into three systems.
Visibility platforms, control towers, risk scores, digital twins, and exception dashboards have defined the last decade of AI in the supply chain. That decade has been very good at collapsing the time between an event and awareness of it, but it has been far less good at collapsing the time between awareness and a commercial act.
Detection is a ‘solved-enough’ problem
Ask a chief supply chain officer where the AI budget went and the answer tends to follow a familiar list: demand sensing, ETA prediction, supplier risk scoring, inventory optimisation, and lane analytics. These tools work. Forecast error comes down. A vessel delay is flagged before the container misses the cut-off. A second-tier fab outage shows up on a heat map instead of in a customer email.
None of that accounts for the $184 billion. The bill is the interval after the flag: expedite or wait; split the order or accept the miss; retender the lane or pay the spot rate; consolidate two half-empty movements or ship both; swap ocean for air on the SKUs that actually justify the premium. These are bounded, repeatable decisions that sit inside policy, contract, and inventory limits the company already set—and they still queue behind a human inbox.
Surveys keep describing the same lag in different language. A 2026 Knosc survey of mid-market manufacturers and distributors found that supply-chain teams spend 28 percent of their working time responding to disruptions, most of it investigating what happened rather than changing what happens next.
Logistics executives still rank AI as a strategic priority (Capgemini’s 2025 research put an AI-driven “new-gen” supply chain among the top three technology trends for 70 percent of large-company executives) and then report that measurable financial impact remains rare. Gartner found in 2025 that only 23 percent of supply-chain organisations even have a formal AI strategy. The shortfall is not a shortage of models, but a shortage of authority granted to software.
The ticket is the product
Most current deployments are built around the ticket. The model produces a recommendation, the recommendation becomes an alert, the alert becomes a work item, and the work item waits for a planner already occupied with other work items. By the time the planner acts, the option set has narrowed—the alternative carrier’s capacity is gone, the consolidation window has closed, and the supplier’s next production slot is allocated.
That workflow is not a temporary step on the way to autonomy but the product companies bought. Vendors sold insight because insight is easy to demonstrate and easy to govern; action touches money, contracts, service levels, and blame. So the industry automated the part of the job that does not require a signature. FourKites and ABI Research reported in 2025 that only 27 percent of organisations allow AI to take autonomous action, while 52 percent confine it to decision support.
Adding another dashboard to a delayed shipment rarely moves EBITDA as a result. The decision cycle has not changed; it has only been decorated.
Bounded action as the next model
The firms set to take share are not the ones with the tidiest control tower but the ones that pre-authorise a narrow class of moves and let agents execute them while the exception is still cheap.
Retender a lane when the contracted carrier’s ETA slips beyond a threshold and a qualified alternate sits inside the approved rate band. Consolidate outbound waves when fill rates and cut-off times make a combined movement cheaper than two. Swap mode on a defined SKU set when the cost of air is lower than the cost of a missed retail window. Reallocate safety stock across two distribution centres when a forecast miss and a transport constraint line up.
None of that requires a strategy offsite. Each can be written as: if these conditions, then this action, within this spend cap, with this audit trail, and a human only if the case falls outside the fence. That is not a “lights-out” supply chain—it is the same discipline manufacturers already apply to machine control, where the agent may act inside the interlock and escalates outside it. The difference here is commercial rather than physical: the interlock is a policy object – category, supplier tier, mode, dollar limit, and service class – not a PLC.
Three conditions for real change
First, decisions have to be written as policies, not tribal knowledge. If the only place “we will pay air on A-items after 48 hours of ocean slip” lives is in a planner’s head, no agent can execute it. The work of the next two years is less model training than decision design: which moves are reversible, which are capped, and which suppliers and modes are pre-cleared.
Second, execution systems have to accept machine-initiated transactions. An agent that can draft an RFQ but cannot post it is still a detection tool. TMS, WMS, sourcing suites, and carrier APIs need to treat a bounded agent the way they treat a junior buyer with a spend limit—authenticated, logged, and reversible.
Third, accountability has to move with the action. If a retender inside policy goes wrong, the post-mortem should inspect the policy, the data, and the fence, not hunt for the person who “should have checked”. Until that cultural change happens, every agent will be designed to wait, because waiting is how careers survive.
The competitive split
For a while, both models will look alike on a slide—both will have AI, and both will have a control tower. The difference will show up in cycle time from detection to commercial act, and then in service and cost.
Companies that keep buying detection will know about the storm earlier. Companies that authorise bounded action will already have retendered the lane, consolidated the wave, and moved the A-items before the incident call is booked.
Disruption is not going away. Lead times in critical components, mode volatility, and multi-tier opacity are structural features of the network. What remains optional is whether the response waits for a human to open a queue. The product that created the lag was insight without authority. The product that ends it is an agent allowed to spend a little money, inside a fence, before anyone is free to look.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Google’s newest AI weather forecasting model predicts wind speed at 100 metres above the ground, roughly the height of a modern wind turbine. It also forecasts cloud cover and how much sunlight reaches the surface, and it updates every hour. Energy traders, grid operators and wind and solar developers already pay other companies for that data. The introduction of WeatherNext 3 now puts Google in their market.
Google DeepMind and Google Research released the model on September 3. It produces a global forecast every hour at up to five-kilometre resolution for surface variables such as temperature and moisture. The previous version, WeatherNext 2, worked on a 25-kilometre grid and refreshed every six hours. Google says the new energy variables are meant to help grid operators and developers predict how much power their wind and solar assets will generate, then match that against demand.
The consumer side of the launch has had most of the attention. WeatherNext 3 now powers weather results in Google Search, the Gemini app, Google Maps and the Google Maps Platform Weather API. Behind it sits an enterprise layer that matters more commercially. The same forecast data can be queried in BigQuery and Earth Engine or downloaded in bulk from Google Cloud Storage, with no model setup required by the customer.
Why the energy sector is buying AI weather forecasting
Grid operators are running a system that has become harder to predict at both ends. On the generation side, renewables now account for most new capacity. S&P Global Market Intelligence’s US Grid Outlook 2026 projects solar and energy storage as the primary sources of new capacity this year, at 51.2GW and 25.7GW respectively out of more than 90GW of planned additions.
Solar and wind generate according to the weather rather than demand, so each gigawatt added makes a short-term forecast more accurate.
On the consumption side, the new load is coming from AI. S&P Global identifies the spread of data centres across North America as a primary driver of the recent surge in electricity demand, forcing utilities to revise their load forecasts upward. Deloitte’s 2026 Power and Utilities Industry Outlook projects peak demand growing by roughly 26% by 2035, with data centre demand alone potentially reaching 176GW, five times its 2024 level.
The cost of getting a forecast wrong is straightforward. If an operator underestimates how much wind power will arrive, it has to buy replacement electricity at short notice, usually from gas plants kept on expensive standby. If it overestimates, wind and solar farms end up being paid to switch off because the grid cannot absorb what they are producing. Both outcomes are expensive, and both are forecasting failures.
The market Google is entering
Selling weather forecasts to the energy sector is an established business. Vaisala, Solcast, DNV’s WindGEMINI and IBM’s HyperWatch all compete in it. So does Jua, a Swiss firm that claims its EPT-2 model beats Microsoft Aurora and DeepMind’s earlier GraphCast on accuracy while updating 24 times a day, against what it describes as a typical four updates a day among competitors.
Google’s advantage is reach. The same forecast appears as a table in BigQuery, a layer in Earth Engine, an API in Google Maps Platform and the default answer in Google Search. No specialist vendor has that spread, and the hourly refresh closes the update-frequency gap those vendors have used to differentiate themselves.
The incumbents have one technical argument left. Jua’s published position is that physics-based models such as ECMWF’s HRES still outperform purely data-driven AI models during record-breaking extreme weather, because physics models encode rules about how energy and mass move through the atmosphere, while AI models learn patterns from past data.
Jua sells a physics-constrained product, so the claim serves its own interests. It also describes the conditions grid operators worry about most, when a storm falls outside anything the model has seen in training.
What is new, and what is being oversold
WeatherNext 3 system architecture showing satellite mosaic and analysis inputs producing gridded forecasts, station data and cyclone tracks. Photo from Google’s blog
The architectural claim behind WeatherNext 3 is that it learns from real observations instead of from simulations. Most AI weather models, WeatherNext 2 included, are trained on output from numerical weather prediction models, which are supercomputer-driven physics simulations that carry a six-hour data lag. That lag can introduce bias in fast-changing variables such as rain and surface temperature. WeatherNext 3 ingests live geostationary satellite imagery and trains directly on readings from individual weather stations.
The shift is real, though narrower than much of the coverage has suggested. Google’s own system diagram shows the model taking in one-hour satellite mosaics alongside traditional historical analysis. DeepMind senior research scientist Ilan Price told Bloomberg the gain comes from not waiting for the next analysis and using the most recent information available instead.
Reporting puts the remaining data lag at three to four hours, down from about seven. Dependence on numerical weather prediction has been reduced, not removed.
The accuracy figures need similar care. Google reports improvements of up to 60% against NASA’s IMERG satellite product, 30% against MRMS radar and 10% against rain gauge readings at early lead times, measured using a standard scoring method for probability forecasts. Those are three separate baselines, and the percentages do not add together. The widely repeated claim of 50% better precipitation forecasting applies specifically to forecasts a day or more ahead. Every figure carries an “up to” qualifier, which makes each one a best case rather than a typical result.
Google published no independent third-party validation alongside the launch. It points instead to live evaluations by Brightband, whose leaderboard it cites in claiming WeatherNext 3 is the most accurate global weather model to date. A utility considering a switch away from a paid specialist will care more about performance in its own service territory, on its own assets, than about a global leaderboard position.
Google’s own stake in the problem
Google is selling forecasting tools into a grid problem its own industry helped create. The data centre build-out driving the load growth utilities are struggling to serve is led by the hyperscalers, Google among them, and Google has signed multi-gigawatt renewable procurement agreements to supply its own facilities.
Accurate prediction of wind and solar output is directly useful to a company matching large volumes of clean energy against a load that is both growing and variable. That is commercial logic, and it goes some way to explaining why the energy variables shipped in this release.
Google has not published pricing for enterprise access to WeatherNext 3, or said whether the BigQuery and Earth Engine data carries standard Cloud query charges or a separate licence. Utilities weighing a move away from a paid specialist will want that figure before they weigh any accuracy claim.
2025, more than 65,000 employees in its Corporate and Investment Bank were actively using the platform, while more than 90% of its engineers were using AI coding assistants.
The bank also said AI-based transaction screening allowed it to review more than twice the previous transaction volume while reducing manual operator checks by half.
Bank of America is using a generative AI-enabled system called EricaAssist with more than 18,000 customer service employees. The tool summarises why a customer is calling, retrieves relevant information, and recommends possible next steps while keeping the employee responsible for the interaction.
Bank of America said in July 2026 that EricaAssist can deliver contextual guidance in under three seconds and has reduced average call times by nearly one minute. The bank plans to extend the system to additional servicing scenarios and business lines later in 2026.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Motional and MIT researchers have built a system that lets self-driving cars explain their decisions in real-time, tackling the black-box problem in autonomous vehicle AI.
The work, published in Nature, comes from a team at Motional that includes CEO Laura Major, working alongside researchers from MIT’s Computer Science and Artificial Intelligence Laboratory. Their proposed method, called the Concept-Wrapper Network or CW-Net, aims to translate the internal calculations of a self-driving system’s neural network into concepts a human can actually read.
If a current self-driving car brakes hard on a clear road with no obvious hazard in sight, neither the driver nor a passenger has any way of knowing why. Modern self-driving systems increasingly rely on neural networks trained on large volumes of driving data. Those networks can perform well, but they don’t expose their reasoning, which is why engineers describe them as black boxes.
Translating neural network logic into human concepts
CW-Net works by converting a self-driving system’s internal logic into concepts such as “Approaching Stopped Vehicle” or “Close to Cyclist.” These could, according to Motional, appear on a dashboard showing which concepts are influencing the vehicle’s driving decisions as they happen.
The system is designed so the explanations aren’t generated after the fact as a guess at what the network might have been doing. Instead, the vehicle’s final decision-making system takes action based directly on these human-interpretable concepts, so a braking event traces back to a specific concept that triggered it. Motional describes this as causally faithful, distinguishing it from approaches that generate natural-language explanations, which can read as plausible without necessarily being accurate.
Laura Major frames the case for this kind of interpretability against the alternative of relying purely on end-to-end deep learning to handle driving decisions.
“The general end-to-end only approach can get to a really good 80-90 percent – maybe even 95 percent – solution, but that’s not good enough to remove a driver or to earn the trust of cities, communities, and customers,” she said.
Testing explainable AI for self-driving cars around Las Vegas
Explainable AI research has largely stayed confined to computer simulations in lab settings, according to Motional. The Motional and MIT team instead deployed CW-Net on an autonomous vehicle with an experienced safety operator in the driver’s seat, collecting data on a private test track and on public roads around Las Vegas.
The team used an earlier experimental version of its deep-learning-based planning system, described as showing competitive performance but with notable shortcomings that CW-Net could help surface. Two incidents from the testing illustrate what the system caught.
In one, the autonomous vehicle repeatedly stopped near a traffic cone, and the vehicle operator assumed the cone itself was triggering the behaviour. Researchers removed the cone and the car stopped anyway. CW-Net’s display showed the actual cause: the experimental planning system was hallucinating a stopped vehicle ahead, a pattern traced back to its training data. That explanation let the researchers understand, predict, and resolve the issue.
A second test involved a cyclist. The autonomous vehicle detected and stopped for the cyclist as expected, but CW-Net revealed that the experimental planning system wasn’t actually basing its decision on the cyclist’s presence. The safety driver responded by exercising more caution around cyclists after noticing this. Follow-up analysis confirmed that caution was warranted, because the vehicle’s braking in that case came from a safety backup system rather than the experimental deep-learning-based planner.
Performance held steady against explainability
Adding layers of explainability to an AI system carries a known cost in speed and performance, and Motional acknowledges that risk. However, when researchers benchmarked CW-Net against leading autonomous driving algorithms, the difference in driving capability came in at less than one percent.
The Las Vegas incidents show why that trade-off matters operationally rather than just academically. A safety driver who can see that a stop is caused by a hallucinated vehicle, or that a backup system rather than the primary planner is responsible for a manoeuvre, can respond and report with more precision than one working from behaviour alone.
That visibility feeds directly into how quickly an engineering team can diagnose a system, and how confidently a safety operator can distinguish between an intended behaviour and a fault.
Motional connects the CW-Net work to broader pressure on autonomous vehicle operators as the technology extends into new markets and jurisdictions. Regulators are naturally asking for more transparency about how AI systems reach their decisions, and it expects tools like CW-Net could move from research projects toward a baseline requirement.
Beyond passenger vehicles, autonomous drones and even robotic surgery are cited as other safety-critical domains where operators and developers will need ways to understand a system’s capabilities, limitations, and unexpected behaviours.
Learn more about physical AI during thePhysical AI Expoheld in Amsterdam, London, and North America.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Anthropic’s most capable AI model has already found thousands of AI cybersecurity vulnerabilities across every major operating system and web browser. The company’s response was not to release it, but to quietly hand it to the organisations responsible for keeping the internet running.
That model is Claude Mythos Preview, and the initiative is called Project Glasswing.
The launch partners include Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, Nvidia, and Palo Alto Networks.
Beyond that core group, Anthropic has extended access to over 40 additional organisations that build or maintain critical software infrastructure. Anthropic is committing up to US$100 million in usage credits for Mythos Preview across the effort, along with US$4 million in direct donations to open-source security organisations.
A model that outgrew its own benchmarks
Mythos Preview was not specifically trained for cybersecurity work. Anthropic said the capabilities “emerged as a downstream consequence of general improvements in code, reasoning, and autonomy”, and that the same improvements making the model better at patching vulnerabilities also make it better at exploiting them.
That last part matters. Mythos Preview has improved to the extent that it mostly saturates existing security benchmarks, forcing Anthropic to shift its focus to novel real-world tasks–specifically, zero-day vulnerabilities. These flaws were previously unknown to the software’s developers.
Among the findings: a 27-year-old bug in OpenBSD, an operating system known for its strong security posture. In another case, the model fully autonomously identified and exploited a 17-year-old remote code execution vulnerability in FreeBSD–CVE-2026-4747–that allows an unauthenticated user anywhere on the internet to obtain complete control of a server running NFS. No human was involved in the discovery or exploitation after the initial prompt to find the bug.
Nicholas Carlini from Anthropic’s research team described the model’s ability to chain together vulnerabilities: “This model can create exploits out of three, four, or sometimes five vulnerabilities that in sequence give you some kind of very sophisticated end outcome. I’ve found more bugs in the last couple of weeks than I found in the rest of my life combined.”
Why is it not being released?
“We do not plan to make Claude Mythos Preview generally available due to its cybersecurity capabilities,” Newton Cheng, Frontier Red Team Cyber Lead at Anthropic, said. “Given the rate of AI progress, it will not be long before such capabilities proliferate, potentially beyond actors who are committed to deploying them safely. The fallout–for economies, public safety, and national security–could be severe.”
This is not hypothetical. Anthropic had previously disclosed what it described as the first documented case of a cyberattack largely executed by AI–a Chinese state-sponsored group that used AI agents to autonomously infiltrate roughly 30 global targets, with AI handling the majority of tactical operations independently.
The company has also privately briefed senior US government officials on Mythos Preview’s full capabilities. The intelligence community is now actively weighing how the model could reshape both offensive and defensive hacking operations.
The open-source problem
One dimension of Project Glasswing that goes beyond the headline coalition: open-source software. Jim Zemlin, CEO of the Linux Foundation, put it plainly: “In the past, security expertise has been a luxury reserved for organisations with large security teams. Open-source maintainers, whose software underpins much of the world’s critical infrastructure, have historically been left to figure out security on their own.”
Anthropic has donated US$2.5 million to Alpha-Omega and OpenSSF through the Linux Foundation, and US$1.5 million to the Apache Software Foundation–giving maintainers of critical open-source codebases access to AI cybersecurity vulnerability scanning at a scale that was previously out of reach.
What comes next
Anthropic says its eventual goal is to deploy Mythos-class models at scale, but only when new safeguards are in place. The company plans to launch new safeguards with an upcoming Claude Opus model first, allowing it to refine them with a model that does not pose the same level of risk as Mythos Preview.
The competitive picture is already shifting around it. When OpenAI released GPT-5.3-Codex in February, the company called it the first model it had classified as high-capability for cybersecurity tasks under its Preparedness Framework. Anthropic’s move with Glasswing signals that the frontier labs see controlled deployment–not open release–as the emerging standard for models at this capability level.
Whether that standard holds as these capabilities spread further is, at this point, an open question that no single initiative can answer.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.