NVIDIA is using Palantir Foundry and cuOpt to automate its hardware supply chain allocation decisions across global manufacturing sites.
The company measures operational delivery from wafer-out to first token. This window splits into time-to-rack (the transit from fab output to an assembled data centre system) and time-to-token (which covers power, cooling, networking, and day-one software readiness.)
Managing NVL72 and Vera Rubin component flows
Hardware scaling has magnified supply constraints. An NVIDIA Grace Blackwell NVL72 rack contains 18 compute trays, with each tray requiring two Grace CPUs, four Blackwell GPUs, and 32 HBM3e memory packages sourced across thousands of suppliers, OEMs, and contract design partners.
The upcoming supply chain constructed for NVIDIA’s Vera Rubin architecture is twice as large as the network supporting Grace Blackwell.
Assembly cannot proceed until parts arrive from three designated channels: direct inventory, consignment stock, and external suppliers. Early shipments must wait on delayed components, extending the metric NVIDIA terms ‘Time of Ownership’ (the duration from when a facility receives materials to when finished sub-assemblies depart.)
Factory allocations are reworked weekly over rolling two-quarter horizons to resolve part availability, throughput limits, and customer fulfilment schedules.
Mixed-integer linear programming via cuOpt
To coordinate these dependencies, the NVIDIA operations team built the ‘Digital Supply Chain Intelligence’ command centre using Palantir Foundry. Foundry’s Ontology models facilities, supplier commits, component stocks, and production targets as interconnected objects and links.
NVIDIA cuOpt, an open-source library for GPU-accelerated decision optimisation, reads this operational layer directly. Formulating distribution as a mixed-integer linear program designed to minimise TOO, the solver evaluates parts constraints across every tier of the bill of materials.
Beyond outputting weekly delivery schedules, cuOpt identifies active factory limits, such as regional assembly capacity caps versus raw memory availability.
Training Nemotron on qualitative operational records
Mathematical optimisation alone failed to capture unstructured operational variables observed by human planners, including supplier call transcripts, regional weather forecasts, partner email exchanges, and geopolitical events.
NVIDIA addressed this by post-training Nemotron 3.5 Lightning, an open-weight mixture-of-experts model featuring 30 billion total parameters and approximately three billion active parameters per forward pass.
The engineering pipeline processes historical records through NeMo Anonymizer to redact sensitive operational fields, NeMo Data Designer to balance training examples with synthetic capacity disruption scenarios, and NeMo AutoModel to apply low-rank adaptation (LoRA) parameters while keeping base model weights frozen. Palantir Autopilot manages data lineage, model tracking, and recommendation delivery.
Production benchmarks and future reinforcement learning
Evaluated on historical allocation records, the post-trained Nemotron 3.5 Lightning model achieved 86.7 percent decision accuracy, compared to 55.5 percent for the larger Nemotron 3 Ultra model and 17.5 percent for the un-tuned Lightning base model.
The post-trained model achieved a 58.6 percent balanced accuracy and a 57.5 percent macro-F1 score, outperforming Nemotron 3 Ultra’s 42 percent balanced accuracy and 39.5 percent macro-F1 score.
Fine-tuning completed on two NVIDIA B200 GPUs within minutes. Domain fine-tuning improved allocation decisions, though production risk forecasting further into the future remained difficult.
Operational choices, planner revisions, overrides, and observed factory outputs are continuously written back to the Palantir Ontology.
NVIDIA confirmed this dataset will form preference pairs for reinforcement learning routines – scoring recommendations on allocation precision, policy compliance, and evidence grounding – with production models remaining strictly isolated from live and unmonitored retraining.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
JD.com is expanding AI and robotics across its logistics network under a new Physical AI Acceleration Plan, while reiterating a five-year target to procure 3 million robots, 1 million autonomous vehicles, and 100,000 delivery drones.
The company launched the plan at JDDiscovery 2026 in Beijing. JD Logistics also unveiled its industrial Wolf Robot series, designed for tasks across warehousing, sorting, transport, and delivery.
Specialised systems include equipment designed to operate in temperatures as low as minus 20 degrees Celsius, automated pharmacy dispatch systems, autonomous delivery vehicles, and drones.
The five-year procurement plan builds on automation that JD Logistics already has in operation. As of June 30, its LangzuTech Goods-to-Person automated warehousing system had been deployed in more than 30 warehouses across China, with deployments also launched in the UK and Germany.
JD Logistics’ broader warehouse network included more than 1,800 self-operated warehouses and more than 2,000 third-party cloud warehouses on its Open Warehouse Platform as of June 30. The network covered more than 36 million square metres in aggregate.
The company also had thousands of unmanned vehicles in regular operation across more than 20 Chinese provinces by the end of June. More than 100 domestic drone routes were operating across applications including parcel and food delivery, emergency medicine transport, and disaster relief.
From AI decisions to physical execution
JD Logistics is connecting its physical equipment with Meta Brain, an AI system used across warehousing, transportation, and delivery. JD said Meta Brain 3.0 can calculate optimal routes for hundreds of millions of parcels in seconds, compared with minutes previously.
Meta Brain also powers JD Logistics’ LangzuTech Packer robotic arm, which combines the model with multimodal sensor data to track, grasp, and place parcels with different shapes.
JD Logistics said in a first-quarter regulatory filing that the Packer uses parallel reinforcement learning in simulated environments to optimise parcel-placement sequences and loading layouts. The company said the system is designed to improve sorting efficiency and the use of available carrier space.
JD upgraded the robotic arm’s force-control technology during the second quarter to support more precise cage-loading operations. By June, the Packer was operating around the clock at multiple JD Logistics parks, according to the company’s interim report.
JD is also adding computing capacity to support AI development. JD Cloud plans to work with Chinese chipmaker Moore Threads on a cluster containing 100,000 GPUs for large-model training, inference, and embodied-AI workloads.
The two companies have previously worked on a 10,000-GPU cluster, according to Data Center Dynamics. Details of which Moore Threads GPU models will be used in the planned 100,000-GPU system have not been disclosed.
JD Cloud also plans to collect more than 10 million hours of video showing real-world human activities over the next two years for embodied-AI training.
Beyond warehouse operations, JD Logistics has expanded its autonomous vehicle network into night-time delivery. Its interim report said the company had launched night-time autonomous routes in Shenzhen, allowing vehicles to operate around the clock.
JD is also using drones in rural logistics. In June, JD Logistics launched a drone delivery network in Zizhong, Sichuan province, covering 78 administrative villages, and said deliveries to some mountain villages could be completed in as little as seven minutes.
Scaling automation across JD’s logistics network
JD did not disclose the total expected cost of the five-year procurement programme at JDDiscovery or provide a network-wide return-on-investment target.
JD Logistics spent RMB2.3 billion on research and development during the first half of 2026, up 23.7% from RMB1.9 billion a year earlier. The company attributed the increase to continued investment in technology and innovation but did not provide a breakdown showing how much was spent specifically on AI or robotics.
Depreciation of property and equipment and amortisation of other intangible assets rose 18.7% to RMB2.6 billion during the first half of 2026, from RMB2.2 billion a year earlier. JD Logistics attributed the increase mainly to additional logistics equipment and vehicles.
Purchases of property and equipment and investment properties totalled RMB3.09 billion over the same six-month period, compared with RMB2.70 billion a year earlier. Those figures cover the wider logistics business and are not disclosed as spending specifically associated with the new physical AI programme.
JD’s automation plans also come as China’s major ecommerce platforms expand fulfilment infrastructure. Reuters reported on September 3 that competition between JD.com, Alibaba, and Meituan had moved from heavy spending on delivery subsidies towards logistics infrastructure, broader supply, and order-level economics.
Alibaba and JD have been opening dark stores and fast-fulfilment “lightning warehouses” in densely populated areas to support deliveries within an hour, while Meituan has been building supermarkets to expand its grocery operations. Ministry of Commerce research cited by Reuters estimates China’s instant-retail market will reach RMB1.2 trillion, or about $178 billion, by the end of 2026.
JD and companies within its ecosystem employ around 700,000 delivery and logistics personnel, according to the South China Morning Post.
JD founder Richard Liu said earlier this year that robots would eventually take over parcel-delivery work now carried out by human couriers. The Financial Times reported in June that JD had signed agreements with around 120 educational institutions to retrain workers for roles including robot repair and maintenance.
JD said JD Logistics currently operates eight robot repair centres in China and plans to expand its robotics after-sales capabilities over the next five years. The company expects the expansion to support more than 100,000 robotics service engineer jobs.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Most notably, Apple is taking aim at the growing wave of AI wearables with new "Audio Intelligence" features that let you rewind moments and remember details from daily conversations.
Samsung has partnered with Mistral AI to deploy on-premises models across its semiconductor manufacturing and engineering operations.
The agreement was announced during the bilateral state summit held in Paris between South Korea and France. Samsung will integrate Mistral’s software suite – including its flagship Mistral Large model – into internal semiconductor facilities to build customised models for intelligence-driven factory infrastructure.
On-premises AI models for semiconductor fab infrastructure
The deployment relies on private enterprise installations to process sensitive engineering and operational records within Samsung’s computing perimeter. This architecture keeps proprietary technical data contained within company infrastructure, avoiding external cloud exposure while maintaining control over operational assets.
“Increasing complexities involved in AI chip design and manufacturing requires continuous innovation in semiconductor technologies,” says Young Hyun Jun, Vice Chairman and CEO of the Device Solutions (DS) Division at Samsung Electronics.
Mistral will provide Samsung with a specialised stack of software tools to assist in how processors are designed and produced.
“AI is reshaping how we build complex technologies, from silicon to software,” says Arthur Mensch, co-founder and CEO of Mistral.
“We are proud to support Samsung Electronics with our expertise in electronics and semiconductors, helping to improve how chips are designed and manufactured, and to accelerate technical progress across the global semiconductor and AI value chain.”
Defect detection and yield stabilisation
Samsung plans to deploy the targeted models directly to automated defect detection and fab machinery tuning. As semiconductor production processes advance, rapid data analysis inside the fab becomes necessary to maintain factory throughput.
The company expects targeted AI models to accelerate development cycles, improve manufacturing precision, and stabilise production yields across advanced memory and logic chips. The operational scope covers Samsung’s memory division, logic design units, and contract foundry business.
Samsung also led Mistral AI’s Series D funding round, securing a strategic equity stake to support long-term technical cooperation.
The lead investment expands cross-industry collaboration between silicon manufacturers and AI developers across advanced memory, logic, and foundry operations.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Arm has launched Arm Total Design for Physical AI alongside a new robotics framework to establish common standards across automated systems.
Physical industries – spanning mining, agriculture, manufacturing, and global transport – account for trillions of dollars in economic activity and an estimated $200 billion annual compute opportunity by the 2030s.
To address engineering fragmentation across these sectors, Arm is convening more than 80 partner organisations spanning software, hardware, and AI. Initial ecosystem participants include AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, and Unitree Robotics.
The initiative targets physical systems that combine AI models, runtime software, compute silicon, sensors, and actuators to sense, reason, and act in operational environments. Hardware manufacturers and software developers require standardised baselines to reduce integration risk, optimise compute workloads, and move from proof-of-concept testing to deployment at scale.
Arm standardises capability tiers for robotics systems
Robotics currently lacks a common method to describe, compare, and communicate system capabilities, according to an architectural manifesto (PDF) published by Arm chief architect Richard Grisenthwaite. This fragmentation makes robotic systems harder to design, integrate, and scale across industrial deployments.
In response, Arm has introduced the Robotics Capability Framework as a collaborative starting point for a shared technical vocabulary, patterned after the SAE Levels used for driving automation.
Arm’s new framework categorises robotic systems across progressing tiers of operational sophistication, mapping machines from reactive setups to context-aware, cognitive, and self-improving systems.
Each capability tier links real-world use cases to machine behaviours, outputs, and hardware constraints. These criteria establish parameters for system latency, compute placement, memory allocation, power constraints, determinism, and safety standards.
Arm developed the initial baseline using feedback from across the robotics sector. Participating organisations contributing to the framework include Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, and Robotec.ai.
Virtual platforms accelerate pre-silicon automotive physical AI development
Arm Total Design for Physical AI extends a collaborative development structure previously used for cloud AI infrastructure. The programme brings together AI models, virtual platforms, digital twins, sensors, compute silicon, and software stacks to enable earlier development and testing cycles.
Autonomous transport and robotics face common technical requirements across sensory perception, AI processing, real-time control, safety, and power-efficient compute. Arm demonstrated this collaborative methodology in the automotive sector alongside AWS, Google, HERE, RemotiveLabs, and Siemens.
The participating automotive companies developed an integrated digital cockpit reference solution. This environment enabled software engineering teams to develop, test, and validate complex automotive code on the Arm Zena CSS platform prior to physical silicon availability.
Arm is now soliciting technical contributions from the wider engineering community to expand the Robotics Capability Framework as physical AI implementations progress.
Learn more about physical AI during thePhysical AI Expoheld in Amsterdam, London, and North America.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
NVIDIA has agreed to acquire Hugging Face for $12.93 billion to scale the open-source model repository’s platform and infrastructure.
The transaction targets platform growth and infrastructure investment, aiming to expand AI access for enterprise developers, software engineers, and research institutions globally.
Built over the past decade by Clem Delangue, Julien Chaumond, Thomas Wolf, and their engineering team, Hugging Face serves as the primary home for the open model developer community.
Platform metrics show more than 18 million developers, researchers, and creators share more than three million models, 500,000 datasets, and one million applications. Commercial adoption includes more than 200,000 companies using the environment to discover, evaluate, customise, and deploy AI models.
Hardware neutrality and multi-cloud commitments
NVIDIA stated that Hugging Face will remain an open platform for the entire AI sector. Developers will retain full control over their selection of models, software frameworks, cloud providers, inference services, and computing hardware.
NVIDIA hardware will not be mandatory to build on or deploy software through the platform. The service will maintain operational support for alternative accelerators, multi-cloud architectures, and open-weight models from all third-party builders.
Jensen Huang, Founder and CEO of NVIDIA, said: “Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want, and the computing platforms they want.
“NVIDIA compute will not be required to build on or deploy through Hugging Face.”
Open-weight models and distributed development
Huang noted a recent open letter he co-authored regarding the role of open weights in the AI economy. The position argued that open weights broaden AI access and ensure technical leadership remains distributed across companies, academic institutions, and developer communities.
Under this operational model, commercial businesses, startups, universities, and public bodies can build on advanced capabilities without the expense of training baseline models from scratch.
The approach allows organisations to match specific models to operational tasks across factories, hospitals, farms, classrooms, and commercial businesses, while addressing cybersecurity and data sovereignty requirements.
Julien Chaumond, Co-Founder and CEO of Hugging Face, commented: “AI is at an inflection point. Open-source AI can become less relevant in the coming years if the big closed labs run away with it, or it can become the foundational fabric of the next phase of human civilisation.
“Those are vastly different outcomes, and we need the critical mass to ensure we give our collective best shot to the second outcome. Given Jensen Huang’s stance on open source AI and how he stepped up to defend it when it was under threat earlier in the summer, NVIDIA was the only partner we truly considered.”
Infrastructure expansion and brand preservation
NVIDIA stands as the largest contributor of open models and data to Hugging Face, with a portfolio of more than 500 open models and more than 250 open datasets. The company builds its libraries, tools, and models openly to allow external engineers to inspect, modify, and build atop the software.
Technical integration will focus on applying NVIDIA infrastructure and engineering resources to improve repository reliability, safety controls, model evaluation tooling, inference execution, and deployment pipelines.
“I am honored that Clem came to me as he considered the next chapter of Hugging Face and believed NVIDIA would be a great home for the company, its community and the future of open models,” Huang stated.
Hugging Face will retain its independent brand identity following the completion of the transaction, with the existing team continuing operations across multi-cloud and multi-accelerator environments.
“This gives fuel to our long-term vision and mission of unlocking the community’s progress to ensure that AI, which is the greatest breakthrough of our lifetime, is accessible to as many people as possible,” Chaumond concludes.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Save up to $410 on your TechCrunch Disrupt 2026 pass before prices increase on May 29 at 11:59 p.m. PT. Register here to join the tech epicenter in San Francisco.
The Nvidia Vera chip is rarely the headline when earnings beat estimates, but it should be. When Nvidia reported Q1 revenue of US$81.62 billion on Wednesday, beating analyst estimates of US$78.86 billion, and guided Q2 at US$91 billion–well above Wall Street’s US$86.84 billion forecast–the numbers did what Nvidia numbers always do: dominate the room.
But buried in CEO Jensen Huang’s conference call with analysts was something more strategically interesting than another quarterly beat. Huang told analysts that Nvidia’s new Vera central processors unlock access to a US$200 billion market, one that sits entirely outside the US$1 trillion the company has already forecast from its Blackwell and Rubin AI GPU lineup between 2025 and 2027.
He expects Vera chip revenue to hit US$20 billion by the end of this fiscal year. “I expect (Vera) to be the second largest” sales contributor, Huang said during the call.
That’s not a footnote. That’s a second front.
The Vera chip and the inference pivot
The reason Nvidia needs a second front is straightforward: its biggest customers are building their own. Google, Amazon, and Microsoft–collectively expected to pour more than US$700 billion into AI infrastructure this year, up sharply from around US$400 billion in 2025, are simultaneously pouring funds into custom silicon to run AI models. Intel and AMD are also touting CPUs as a credible play for inference workloads.
The narrative in the chip industry has shifted from who can train the biggest model to who can serve it cheapest and fastest. Inference is where Nvidia’s GPU dominance is most exposed. Training large models is still firmly Nvidia territory, but inference, generating answers at scale, in real time, is increasingly where custom chips from Google’s TPU line, Amazon’s Trainium and others are making their case.
Nvidia’s answer is Vera. The chip, developed in part using technology from Groq, a startup specialising in inference that Nvidia licensed in a deal reportedly worth around US$17 billion, targets exactly this workload. The full Vera Rubin platform, which combines the Vera CPU with Rubin GPUs, is set to launch later this year.
Supply is already the constraint
Huang was candid about one problem: supply. “My sense is that we’ll be supply-constrained through the entire life of Vera Rubin,” he said on the call. It’s a telling admission for a product Nvidia is positioning as a major growth pillar. To get ahead of disruptions, Nvidia is spending heavily on the supply chain. The company disclosed that its supply commitments rose to US$119 billion in Q1, up from US$95.2 billion the previous quarter, a significant jump that reflects both confidence in demand and anxiety about a global memory chip crunch.
Nvidia also announced a US$80 billion share repurchase programme and raised its quarterly cash dividend to 25 cents per share, from 1 cent, moves that signal financial confidence even as Huang warned of tightening supply.
The question investors are asking
Despite the beats, Nvidia shares fell 1.6% in extended trading after the results. eMarketer analyst Jacob Bourne captured the mood: “Nvidia delivered another beat, but at this point that’s essentially priced in as it keeps beating quarter after quarter. The lingering question is whether it can convince investors the AI buildout has durability into 2027 and 2028, especially as the narrative shifts toward inference workloads and competing silicon from Google, Amazon, AMD, and Intel.”
Huang pushed back with numbers of his own. He pointed to a growing sub-segment of AI-specific cloud customers whose spend is now roughly equal to the hyperscalers, but growing faster quarter-over-quarter. “We should be growing faster than hyperscale capex,” he said.
The Vera chip is central to that argument. Whether the supply chain cooperates is a different question entirely.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Alibaba has unveiled a new AI processor built specifically for AI agents, pairing the chip announcement with a multi-year silicon roadmap and a new large language model, signalling that the company is building an integrated AI stack rather than just filling a gap left by US export controls.
The Zhenwu M890, developed by Alibaba’s semiconductor subsidiary T-Head, delivers three times the performance of its predecessor, the Zhenwu 810E, according to the company, as per Reuters report. But the performance jump is less notable than the architectural intent behind the chip: the M890 is purpose-built for AI agents, where software systems must retain long stretches of context, coordinate with other models in real time, and execute complex multi-step tasks with limited human intervention.
Those demands, heavy on memory bandwidth and inter-model communication, are meaningfully different from what standard inference chips are optimised for. The difference matters because it tells you something about where Alibaba thinks AI compute is heading. The company isn’t designing around today’s dominant use case; it’s building for the workload profile it expects to define enterprise AI over the next several years.
Built for AI agents, not just inference
More significant than the chip itself is the roadmap Alibaba put alongside it. The M890 will be followed by the V900 in the third quarter of 2027, expected to deliver another roughly threefold performance gain, followed by the J900 in the third quarter of 2028. That’s a deliberate, sustained cadence of in-house silicon upgrades that mirrors the kind of tick-tock product cycles Nvidia has used to maintain its lead in AI accelerators.
The parallel to Huawei is worth noting. Huawei laid out a similar chip roadmap for its Ascend line last year, and both announcements reflect the same underlying reality: Chinese technology companies have concluded that depending on foreign silicon, even in scenarios where export restrictions might ease, is a structural risk they cannot accept. The response has been to treat semiconductor development as a long-term capability-building exercise rather than a procurement problem.
Alibaba’s commitment to that exercise is not shallow. The company pledged more than 380 billion yuan, roughly US$53 billion, on cloud and AI infrastructure over three years last year, its largest-ever investment commitment to the sector. The M890 and its successors are downstream of that spending.
Traction that predates the announcement
T-Head said it has shipped more than 560,000 Zhenwu units to date, with over 400 external customers across 20 industries deploying the chips, including automakers and financial services firms. That is a material production footprint, not lab hardware, and it provides Alibaba with real-world deployment data at scale ahead of the M890’s rollout.
The new chip will be available to Chinese enterprise customers through Alibaba Cloud’s domestic model platform, Bailian, packaged inside the Panjiu AL128, a server system that stacks 128 M890 accelerators into a single rack.
The software side of the stack
Alongside the hardware, Alibaba announced Qwen 3.7-Max, the latest version of its flagship large language model, described as engineered for advanced coding and long-running agent tasks. The company said the model can operate continuously for up to 35 hours without performance degradation, a capability specification that only makes sense if you are designing for extended autonomous operation.
The timing is deliberate. Releasing a chip and a model optimised for the same workload class on the same day is a platform play. Alibaba is building a closed loop: its own silicon in T-Head, its own model in Qwen, its own cloud delivery in Bailian. Each component reinforces the others, and the combined stack is designed to reduce enterprise customers’ dependence on any external vendor.
More than half a million chips have been shipped. A successor is arriving in 2027, with another planned for 2028. T-Head is not hedging. At some point, building around US export controls stops being a workaround and starts being a strategy. Alibaba appears to have crossed that line.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure.
When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely.
At the initial product stage, exerting tight corporate control often feels highly advantageous. Closed development environments iterate quickly and tightly manage the end-user experience. They capture and concentrate financial value within a single corporate entity, an approach that functions adequately during early product development cycles.
However, IBM’s analysis highlights that expectations change entirely when a technology solidifies into a foundational layer. Once other institutional frameworks, external markets, and broad operational systems rely on the software, the prevailing standards adapt to a new reality. At infrastructure scale, embracing openness ceases to be an ideological stance and becomes a highly practical necessity.
AI is currently crossing this threshold within the enterprise architecture stack. Models are increasingly embedded directly into the ways organisations secure their networks, author source code, execute automated decisions, and generate commercial value. AI functions less as an experimental utility and more as core operational infrastructure.
The recent limited preview of Anthropic’s Claude Mythos model brings this reality into sharper focus for enterprise executives managing risk. Anthropic reports that this specific model can discover and exploit software vulnerabilities at a level matching few human experts.
In response to this power, Anthropic launched Project Glasswing, a gated initiative designed to place these advanced capabilities directly into the hands of network defenders first. From IBM’s perspective, this development forces technology officers to confront immediate structural vulnerabilities. If autonomous models possess the capability to write exploits and shape the overall security environment, Thomas notes that concentrating the understanding of these systems within a small number of technology vendors invites severe operational exposure.
With models achieving infrastructure status, IBM argues the primary issue is no longer exclusively what these machine learning applications can execute. The priority becomes how these systems are constructed, governed, inspected, and actively improved over extended periods.
As underlying frameworks grow in complexity and corporate importance, maintaining closed development pipelines becomes exceedingly difficult to defend. No single vendor can successfully anticipate every operational requirement, adversarial attack vector, or system failure mode.
Implementing opaque AI structures introduces heavy friction across existing network architecture. Connecting closed proprietary models with established enterprise vector databases or highly sensitive internal data lakes frequently creates massive troubleshooting bottlenecks. When anomalous outputs occur or hallucination rates spike, teams lack the internal visibility required to diagnose whether the error originated in the retrieval-augmented generation pipeline or the base model weights.
Integrating legacy on-premises architecture with highly gated cloud models also introduces severe latency into daily operations. When enterprise data governance protocols strictly prohibit sending sensitive customer information to external servers, technology teams are left attempting to strip and anonymise datasets before processing. This constant data sanitisation creates enormous operational drag.
Furthermore, the spiralling compute costs associated with continuous API calls to locked models erode the exact profit margins these autonomous systems are supposed to enhance. The opacity prevents network engineers from accurately sizing hardware deployments, forcing companies into expensive over-provisioning agreements to maintain baseline functionality.
Why open-source AI is essential for operational resilience
Restricting access to powerful applications is an understandable human instinct that closely resembles caution. Yet, as Thomas points out, at massive infrastructure scale, security typically improves through rigorous external scrutiny rather than through strict concealment.
This represents the enduring lesson of open-source software development. Open-source code does not eliminate enterprise risk. Instead, IBM maintains it actively changes how organisations manage that risk. An open foundation allows a wider base of researchers, corporate developers, and security defenders to examine the architecture, surface underlying weaknesses, test foundational assumptions, and harden the software under real-world conditions.
Within cybersecurity operations, broad visibility is rarely the enemy of operational resilience. In fact, visibility frequently serves as a strict prerequisite for achieving that resilience. Technologies deemed highly important tend to remain safer when larger populations can challenge them, inspect their logic, and contribute to their continuous improvement.
Thomas addresses one of the oldest misconceptions regarding open-source technology: the belief that it inevitably commoditises corporate innovation. In practical application, open infrastructure typically pushes market competition higher up the technology stack. Open systems transfer financial value rather than destroying it.
As common digital foundations mature, the commercial value relocates toward complex implementation, system orchestration, continuous reliability, trust mechanics, and specific domain expertise. IBM’s position asserts that the long-term commercial winners are not those who own the base technological layer, but rather the organisations that understand how to apply it most effectively.
We have witnessed this identical pattern play out across previous generations of enterprise tooling, cloud infrastructure, and operating systems. Open foundations historically expanded developer participation, accelerated iterative improvement, and birthed entirely new, larger markets built on top of those base layers. Enterprise leaders increasingly view open-source as highly important for infrastructure modernisation and emerging AI capabilities. IBM predicts that AI is highly likely to follow this exact historical trajectory.
Looking across the broader vendor ecosystem, leading hyperscalers are adjusting their business postures to accommodate this reality. Rather than engaging in a pure arms race to build the largest proprietary black boxes, highly profitable integrators are focusing heavily on orchestration tooling that allows enterprises to swap out underlying open-source models based on specific workload demands. Highlighting its ongoing leadership in this space, IBM is a key sponsor of this year’s AI & Big Data Expo North America, where these evolving strategies for open enterprise infrastructure will be a primary focus.
This approach completely sidesteps restrictive vendor lock-in and allows companies to route less demanding internal queries to smaller and highly efficient open models, preserving expensive compute resources for complex customer-facing autonomous logic. By decoupling the application layer from the specific foundation model, technology officers can maintain operational agility and protect their bottom line.
The future of enterprise AI demands transparent governance
Another pragmatic reason for embracing open models revolves around product development influence. IBM emphasises that narrow access to underlying code naturally leads to narrow operational perspectives. In contrast, who gets to participate directly shapes what applications are eventually built.
Providing broad access enables governments, diverse institutions, startups, and varied researchers to actively influence how the technology evolves and where it is commercially applied. This inclusive approach drives functional innovation while simultaneously building structural adaptability and necessary public legitimacy.
As Thomas argues, once autonomous AI assumes the role of core enterprise infrastructure, relying on opacity can no longer serve as the organising principle for system safety. The most reliable blueprint for secure software has paired open foundations with broad external scrutiny, active code maintenance, and serious internal governance.
As AI permanently enters its infrastructure phase, IBM contends that identical logic increasingly applies directly to the foundation models themselves. The stronger the corporate reliance on a technology, the stronger the corresponding case for demanding openness.
If these autonomous workflows are truly becoming foundational to global commerce, then transparency ceases to be a subject of casual debate. According to IBM, it is an absolute, non-negotiable design requirement for any modern enterprise architecture.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The news follows a report from Nikkei Asia on Tuesday that raised concerns the company’s foldable iPhone could be delayed due to challenges during the phone’s engineering test phase.
Madelyn Olson discusses the evolution of Valkey's data structures, moving away from "textbook" pointer-chasing HashMaps to more cache-aware designs. She explains the implementation of "Swedish" tables to maximize memory density. She shares insights on systems intuition, memory prefetching, and the rigorous testing needed for mission-critical caches.