❌

Reading view

GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity

OpenAI has classified GPT-6 Astra at the Critical cybersecurity threshold under its Preparedness Framework, a first. In expert-led testing the model found previously unknown vulnerabilities in a browser and an OS kernel and built working exploits. The same system card reports a substantial decline in chain-of-thought monitorability.

By Steef-Jan Wiggers
  •  

Deloitte: Scale ‘autonomous intelligence’ for real growth

Enterprise leaders must progress past generative applications and scale “autonomous intelligence” to capture real growth.

Generating text or summarising internal communications offers localised productivity improvements, yet these abilities rarely alter the core cost or revenue structure of a large organisation. Enterprises are now focused on deploying systems capable of independent execution. Leaders are demanding applications that can traverse internal networks, execute multi-step logic, and finalise transactions without constant human prompting.

Prakul Sharma, principal and AI & Insights Practice Leader at Deloitte Consulting LLP, said: “At Deloitte, we view this as the third stage on an intelligence maturity curve, from ‘assisted intelligence,’ in which AI and analytics help people interpret information, through ‘artificial intelligence,’ with machine learning augmenting human decisions, to ‘autonomous intelligence,’ where AI decides and executes in defined boundaries.

“Today’s GenAI-era abilities – like chatbots and conversational AI – sit in the middle of that curve. Agentic AI acts as the bridge into autonomy, and it is where the centre of gravity is changing now. The difference we are seeing is agency: GenAI produces an answer, while autonomous intelligence pursues an outcome by reasoning over a goal, invoking tools and data, and adapting as conditions change, with humans setting guardrails not driving every step.

“We’re seeing this show up in industries, and in every case, the unlock isn’t the agent itself, but the surrounding governance architecture of identity and human-in-the-loop checkpoints, making autonomy safe to scale.”

Forensic audits for targeted margin improvement

To extract actual economic value, these autonomous systems must integrate directly into revenue-generating or cost-heavy workflows.

Consider a scenario in enterprise procurement: an agentic application continuously cross-references supply chain inventory against live vendor pricing in an enterprise resource planning system. It can then independently authorise purchase orders in predefined financial parameters, halting only for human approval when deviations occur.

The same system must also carry a verifiable identity in the ERP, read pricing data that is current enough to be contractually binding, and operate in approval thresholds that legal and compliance have formally endorsed. Any one of those dependencies, left unresolved, collapses the case for autonomous execution entirely. Achieving this level of automation therefore requires a forensic examination of existing operations before allocating any compute resources.

Sharma outlines the method Deloitte uses to initiate this operational overhaul and locate areas where autonomy can generate tangible revenue:

“The first step we advise is starting with a decision audit and the process. We ask leaders to pick one or two value chains where outcomes are bottlenecked by decisions not by tasks in that process, and to map how those decisions get made today. We ask questions like who has the data, who has the authority, where the handoffs break, what actions are needed, and where judgement is being applied.

“Asking these questions surfaces the process workflows where autonomy will create real economic value, while simultaneously exposing any data and governance gaps that may have derailed a pilot. From there, we help leaders sequence the rewire: stand up the foundational layers with AI and agentic fabric, data, evals, agent identity, and human-in-the-loop patterns against that first value chain, prove it works, and then use it as the template to scale.”

Integrating the right data infrastructure and upstream architecture

Once the operational target is isolated, the technological execution frequently stalls owing to upstream friction. The underlying foundation models from major providers have advanced quickly enough to handle complex reasoning tasks, becoming largely interchangeable commodities. The friction point lies in connecting these reasoning engines to legacy data architectures.

Sharma observes that the true technical barriers emerge long before the prompt reaches the large language model:

“Based on what we are seeing, the model is rarely the bottleneck, since frontier ability is now rapidly becoming a commodity. Where enterprises trip up in the design phase is upstream of the model. They select a use case before mapping the underlying workflow, resulting in the agent automating a process that was already broken or poorly instrumented.

“The second pattern is data: clients may underestimate that autonomous systems need decision-grade data, not reporting-grade data, meaning lineage and access controls that most enterprise data estates were not built to support.”

The distinction matters because most enterprise data estates were built for human analysts, not autonomous systems. Reporting-grade data – aggregated on a nightly or weekly batch cycle, structured for dashboard consumption, and stripped of the lineage that records how a value was derived – is adequate when a person applies judgement before acting on it. An autonomous agent has no such backstop. When it retrieves a contract price or a stock level to execute a transaction, that figure must carry a timestamp current enough to be binding, a traceable provenance, and access controls that confirm the agent is authorised to read and act on it.

Providing this decision-grade data involves integrating autonomous agents with right event stores and databases designed to manage both structured and unstructured enterprise information. When an agent retrieves data to execute a task, the enterprise must guarantee its freshness. Relying on stale batch-processed data introduces extreme risk, potentially causing the system to act on obsolete pricing tiers or outdated compliance frameworks.

The financial model for scaling these systems also requires forecasting variable compute expenses. Because agentic workflows involve multiple interactions with large language models to reason through a single goal, API costs can escalate unpredictably. Mitigating hallucination risks through retrieval-augmented generation processes also increases the necessary compute overhead, requiring strict financial controls before enterprise-deployment.

Reconciling governance debt and enterprise ecosystems

Transitioning from controlled testing environments to live enterprise deployment is a very different proposition. A small-scale test might perform perfectly using carefully selected data sets, but deploying that ability in thousands of employees and interconnected software platforms exposes vulnerabilities.

Navigating modern enterprise security environments means integrating the agentic architecture deeply with existing identity providers and cloud-native security controls across hybrid cloud ecosystems.

Sharma identifies this integration failure and the resulting governance debt that halts progress:

“The main roadblock we see is what we call the production gap. A pilot can succeed with a clever prompt, a curated dataset, and a champion team running it manually, but enterprise deployment requires continuous evaluations, identity and authorisation that work in systems the pilot never touched, change management for the users, and a financial model that can absorb use-based costs at scale.

“Tied to that is governance debt: the controls, audit trails, and risk frameworks waived to accelerate a pilot often become the gating items once legal and compliance evaluate a production rollout. The clients that break through are ones that don’t treat pilots as experiments but instead treat them as the first production instance of a reusable platform – with the same evals, identity model, and governance. Instead of starting over, this allows the second and third use cases to build on the first.”

Compliance frameworks applied during initial testing are often completely insufficient for live deployment. Teams eager to prove a concept frequently bypass standard corporate security protocols, creating the very gating items that prevent future scaling.

What unites all three failure modes – the production gap, governance debt, and upstream data friction – is that each one is invisible during a well-run pilot. A champion team with a curated dataset and management cover can paper over missing identity controls, stale data, and deferred compliance reviews for long enough to produce a convincing demonstration. It is only when the system must operate in the full enterprise, with real users, live data, and legal scrutiny, that the gaps become structural blockers not known workarounds.

Building a reusable platform from the outset – with identity verification, continuous model evaluations, and financial monitoring treated as first-class requirements not post-launch additions – is what allows organisations to avoid rebuilding those foundations for every subsequent deployment.

Prakul Sharma’s interview was conducted ahead of the AI & Big Data Expo North America, where Deloitte is a important sponsor. Be sure to swing by Deloitte’s booth at stand #272 to hear more directly from the organisation’s experts. Prakul Sharma will be sharing more of his insights during a panel session on day one and day two of the industry-leading event.

 

(Image source: Pixabay, under licence.)

 

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Deloitte: Scale ‘autonomous intelligence’ for real growth appeared first on AI News.

  •  

Meta has a competitive AI model but loses its open-source identity

The open-source AI movement has never lacked for options. Mistral, Falcon, and a growing field of open-weight models have been available to developers for years. But when Meta threw its weight behind Llama, something shifted. A company with three billion users, vast compute resources, and the credibility of a tech giant was now building openly, and the developer community responded.

By early 2026, the Llama ecosystem had reached 1.2 billion downloads, averaging about 1 million per day. That is the context for what happened on April 8, 2026. Meta launched Muse Spark, its first major new Meta AI model in a year, and the first product from its newly formed Meta Superintelligence Labs.

It is capable in ways Llama 4 never was, benchmarks well against the current frontier, and is completely proprietary. No free download. No open weights. No building on it unless Meta decides you can.

The companyspentUS$14.3 billion, brought in Alexandr Wang from Scale AI to lead its AI rebuild, then spent nine months tearing down its entire AI stack and starting over. Muse Spark is what came out the other side. The developer community that made Llama what it was is now being asked to wait for a future open-source version that may or may not arrive on any predictable timeline.

What is Muse Spark?

Muse Spark is a natively multimodal reasoning model with tool-use, visual chain of thought, and multi-agent orchestration built in. It now powers Meta AI, which reaches over three billion users in Meta’s apps. Meta rebuilt its technology infrastructure from scratch, letting the company create a model that is as capable as its older midsize Llama 4 variant for an order of magnitude less compute.

That efficiency number is worth noting. At the scale Meta operates, compute costs compound fast, and running a frontier-class Meta AI model at a fraction of the cost of its predecessors changes the economics of deploying it in billions of interactions daily.

On benchmarks, the picture is genuinely mixed. Muse Spark scores 52 on the Artificial Intelligence Index v4.0, placing it fourth overall behind Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Meta has not claimed to have built the best model in the world, which is itself a departure from the over-claiming that damaged Llama 4’s credibility.

Where Muse Spark leads is health. On HealthBench Hard – open-ended health queries – it scores 42.8, substantially ahead of Gemini 3.1 Pro at 20.6, GPT-5.4 at 40.1, and Grok 4.2 at 20.3. Health is a stated priority for Meta; the company says it worked with over 1,000 physicians to curate training data for the model.

Muse Spark also offers three modes of interaction: Instant mode for quick answers, Thinking mode for multi-step reasoning tasks, and Contemplating mode, which orchestrates multiple agents’ reasoning in parallel to compete with the most demanding reasoning modes from Gemini Deep Think and GPT Pro.

The open-source retreat

This is the part of the Muse Spark story that the benchmark tables do not capture. Unlike Meta’s previous models, which were released as open-weight models – meaning anyone could download and run them on their own equipment – Muse Spark is entirely proprietary. The company said it will offer the model in a private preview to select partners through an API, making Muse Spark even more proprietary than the paid models offered by Meta’s rivals.

Wang addressed the change directly, stating: “Nine months ago, we rebuilt our AI stack from scratch. New infrastructure, new architecture, new data pipelines. This is step one. Bigger models are already in development with plans to open-source future versions.”

The developer community’s response has been sceptical. Some see this as a necessary pivot after Llama 4 failed to gain expected traction. Others view it as Meta closing the gates once it has something worth protecting. That is the community now being asked to wait while competitors without that open-source legacy continue shipping freely available weights.

Distribution over benchmarks

Meanwhile, Meta is not waiting for the developer community to come around. Muse Spark will debut in the coming weeks inside Facebook, Instagram, WhatsApp, and Messenger, as well as in Meta’s Ray-Ban AI glasses. That rollout path is arguably more consequential than any benchmark result. OpenAI and Anthropic sell to developers and enterprises. Meta deploys directly to over three billion people already inside its apps daily.

Meta’s push into health does raise privacy questions worth watching. Muse Spark users will need to log in with an existing Meta account to use it, and while Meta does not explicitly say personal account information will be used by the AI, the company has generally trained on public user data and has positioned Muse Spark as a personal superintelligence product.

Meta stock rose more than 9% on the day of the launch, a signal that investors read the Muse Spark release as proof that the US$14.3 billion bet on Wang and the nine-month rebuild produced something real. Whether the promised open-source versions actually materialise is a question the developer community will press every quarter. The answer will define how this chapter of Meta’s AI story is remembered.

See Also: The Meta-Manus review: What enterprise AI buyers need to know about cross-border compliance risk

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Meta has a competitive AI model but loses its open-source identity appeared first on AI News.

  •  
❌