❌

Reading view

Palantir Foundry and cuOpt drive NVIDIA supply chain allocation

NVIDIA is using Palantir Foundry and cuOpt to automate its hardware supply chain allocation decisions across global manufacturing sites.

The company measures operational delivery from wafer-out to first token. This window splits into time-to-rack (the transit from fab output to an assembled data centre system) and time-to-token (which covers power, cooling, networking, and day-one software readiness.)

Managing NVL72 and Vera Rubin component flows

Hardware scaling has magnified supply constraints. An NVIDIA Grace Blackwell NVL72 rack contains 18 compute trays, with each tray requiring two Grace CPUs, four Blackwell GPUs, and 32 HBM3e memory packages sourced across thousands of suppliers, OEMs, and contract design partners.

The upcoming supply chain constructed for NVIDIA’s Vera Rubin architecture is twice as large as the network supporting Grace Blackwell.

Assembly cannot proceed until parts arrive from three designated channels: direct inventory, consignment stock, and external suppliers. Early shipments must wait on delayed components, extending the metric NVIDIA terms ‘Time of Ownership’ (the duration from when a facility receives materials to when finished sub-assemblies depart.)

Factory allocations are reworked weekly over rolling two-quarter horizons to resolve part availability, throughput limits, and customer fulfilment schedules.

Mixed-integer linear programming via cuOpt

To coordinate these dependencies, the NVIDIA operations team built the ‘Digital Supply Chain Intelligence’ command centre using Palantir Foundry. Foundry’s Ontology models facilities, supplier commits, component stocks, and production targets as interconnected objects and links.

NVIDIA cuOpt, an open-source library for GPU-accelerated decision optimisation, reads this operational layer directly. Formulating distribution as a mixed-integer linear program designed to minimise TOO, the solver evaluates parts constraints across every tier of the bill of materials.

Beyond outputting weekly delivery schedules, cuOpt identifies active factory limits, such as regional assembly capacity caps versus raw memory availability.

Training Nemotron on qualitative operational records

Mathematical optimisation alone failed to capture unstructured operational variables observed by human planners, including supplier call transcripts, regional weather forecasts, partner email exchanges, and geopolitical events.

NVIDIA addressed this by post-training Nemotron 3.5 Lightning, an open-weight mixture-of-experts model featuring 30 billion total parameters and approximately three billion active parameters per forward pass.

The engineering pipeline processes historical records through NeMo Anonymizer to redact sensitive operational fields, NeMo Data Designer to balance training examples with synthetic capacity disruption scenarios, and NeMo AutoModel to apply low-rank adaptation (LoRA) parameters while keeping base model weights frozen. Palantir Autopilot manages data lineage, model tracking, and recommendation delivery.

Production benchmarks and future reinforcement learning

Evaluated on historical allocation records, the post-trained Nemotron 3.5 Lightning model achieved 86.7 percent decision accuracy, compared to 55.5 percent for the larger Nemotron 3 Ultra model and 17.5 percent for the un-tuned Lightning base model.

The post-trained model achieved a 58.6 percent balanced accuracy and a 57.5 percent macro-F1 score, outperforming Nemotron 3 Ultra’s 42 percent balanced accuracy and 39.5 percent macro-F1 score.

Accuracy score results for the post-trained NVIDIA Nemotron 3.5 Lightning AI model.

Fine-tuning completed on two NVIDIA B200 GPUs within minutes. Domain fine-tuning improved allocation decisions, though production risk forecasting further into the future remained difficult.

Operational choices, planner revisions, overrides, and observed factory outputs are continuously written back to the Palantir Ontology.

NVIDIA confirmed this dataset will form preference pairs for reinforcement learning routines – scoring recommendations on allocation precision, policy compliance, and evidence grounding – with production models remaining strictly isolated from live and unmonitored retraining.

See also: Supply chains detect fast, act slow: How AI agents fix it

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Palantir Foundry and cuOpt drive NVIDIA supply chain allocation appeared first on AI News.

  •  

Samsung taps Mistral AI models for semiconductor manufacturing

Samsung has partnered with Mistral AI to deploy on-premises models across its semiconductor manufacturing and engineering operations.

The agreement was announced during the bilateral state summit held in Paris between South Korea and France. Samsung will integrate Mistral’s software suite – including its flagship Mistral Large model – into internal semiconductor facilities to build customised models for intelligence-driven factory infrastructure.

On-premises AI models for semiconductor fab infrastructure

The deployment relies on private enterprise installations to process sensitive engineering and operational records within Samsung’s computing perimeter. This architecture keeps proprietary technical data contained within company infrastructure, avoiding external cloud exposure while maintaining control over operational assets.

“Increasing complexities involved in AI chip design and manufacturing requires continuous innovation in semiconductor technologies,” says Young Hyun Jun, Vice Chairman and CEO of the Device Solutions (DS) Division at Samsung Electronics.

Mistral will provide Samsung with a specialised stack of software tools to assist in how processors are designed and produced.

“AI is reshaping how we build complex technologies, from silicon to software,” says Arthur Mensch, co-founder and CEO of Mistral.

“We are proud to support Samsung Electronics with our expertise in electronics and semiconductors, helping to improve how chips are designed and manufactured, and to accelerate technical progress across the global semiconductor and AI value chain.”

Defect detection and yield stabilisation

Samsung plans to deploy the targeted models directly to automated defect detection and fab machinery tuning. As semiconductor production processes advance, rapid data analysis inside the fab becomes necessary to maintain factory throughput.

The company expects targeted AI models to accelerate development cycles, improve manufacturing precision, and stabilise production yields across advanced memory and logic chips. The operational scope covers Samsung’s memory division, logic design units, and contract foundry business.

Samsung also led Mistral AI’s Series D funding round, securing a strategic equity stake to support long-term technical cooperation.

The lead investment expands cross-industry collaboration between silicon manufacturers and AI developers across advanced memory, logic, and foundry operations.

See also: Arm launches Total Design for Physical AI and robotics framework

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Samsung taps Mistral AI models for semiconductor manufacturing appeared first on AI News.

  •  

Arm launches Total Design for Physical AI and robotics framework

Arm has launched Arm Total Design for Physical AI alongside a new robotics framework to establish common standards across automated systems.

Physical industries – spanning mining, agriculture, manufacturing, and global transport – account for trillions of dollars in economic activity and an estimated $200 billion annual compute opportunity by the 2030s.

To address engineering fragmentation across these sectors, Arm is convening more than 80 partner organisations spanning software, hardware, and AI. Initial ecosystem participants include AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, and Unitree Robotics.

The initiative targets physical systems that combine AI models, runtime software, compute silicon, sensors, and actuators to sense, reason, and act in operational environments. Hardware manufacturers and software developers require standardised baselines to reduce integration risk, optimise compute workloads, and move from proof-of-concept testing to deployment at scale.

Arm standardises capability tiers for robotics systems

Robotics currently lacks a common method to describe, compare, and communicate system capabilities, according to an architectural manifesto (PDF) published by Arm chief architect Richard Grisenthwaite. This fragmentation makes robotic systems harder to design, integrate, and scale across industrial deployments.

In response, Arm has introduced the Robotics Capability Framework as a collaborative starting point for a shared technical vocabulary, patterned after the SAE Levels used for driving automation.

Arm’s new framework categorises robotic systems across progressing tiers of operational sophistication, mapping machines from reactive setups to context-aware, cognitive, and self-improving systems.

Each capability tier links real-world use cases to machine behaviours, outputs, and hardware constraints. These criteria establish parameters for system latency, compute placement, memory allocation, power constraints, determinism, and safety standards.

The robotics capability framework for physical AI by Arm.

Arm developed the initial baseline using feedback from across the robotics sector. Participating organisations contributing to the framework include Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, and Robotec.ai.

Virtual platforms accelerate pre-silicon automotive physical AI development

Arm Total Design for Physical AI extends a collaborative development structure previously used for cloud AI infrastructure. The programme brings together AI models, virtual platforms, digital twins, sensors, compute silicon, and software stacks to enable earlier development and testing cycles.

Autonomous transport and robotics face common technical requirements across sensory perception, AI processing, real-time control, safety, and power-efficient compute. Arm demonstrated this collaborative methodology in the automotive sector alongside AWS, Google, HERE, RemotiveLabs, and Siemens.

The participating automotive companies developed an integrated digital cockpit reference solution. This environment enabled software engineering teams to develop, test, and validate complex automotive code on the Arm Zena CSS platform prior to physical silicon availability.

Arm is now soliciting technical contributions from the wider engineering community to expand the Robotics Capability Framework as physical AI implementations progress.

Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America.

See also: NVIDIA Jetson Orin Nano 2 brings physical AI to drones and robots

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Arm launches Total Design for Physical AI and robotics framework appeared first on AI News.

  •  

Nvidia’s Vera chip is the US$200 billion bet Jensen Huang doesn’t want you to overlook

The Nvidia Vera chip is rarely the headline when earnings beat estimates, but it should be. When Nvidia reported Q1 revenue of US$81.62 billion on Wednesday, beating analyst estimates of US$78.86 billion, and guided Q2 at US$91 billion–well above Wall Street’s US$86.84 billion forecast–the numbers did what Nvidia numbers always do: dominate the room. 

But buried in CEO Jensen Huang’s conference call with analysts was something more strategically interesting than another quarterly beat. Huang told analysts that Nvidia’s new Vera central processors unlock access to a US$200 billion market, one that sits entirely outside the US$1 trillion the company has already forecast from its Blackwell and Rubin AI GPU lineup between 2025 and 2027. 

He expects Vera chip revenue to hit US$20 billion by the end of this fiscal year. “I expect (Vera) to be the second largest” sales contributor, Huang said during the call.

That’s not a footnote. That’s a second front.

The Vera chip and the inference pivot

The reason Nvidia needs a second front is straightforward: its biggest customers are building their own. Google, Amazon, and Microsoft–collectively expected to pour more than US$700 billion into AI infrastructure this year, up sharply from around US$400 billion in 2025, are simultaneously pouring funds into custom silicon to run AI models. Intel and AMD are also touting CPUs as a credible play for inference workloads. 

The narrative in the chip industry has shifted from who can train the biggest model to who can serve it cheapest and fastest. Inference is where Nvidia’s GPU dominance is most exposed. Training large models is still firmly Nvidia territory, but inference, generating answers at scale, in real time, is increasingly where custom chips from Google’s TPU line, Amazon’s Trainium and others are making their case.

Nvidia’s answer is Vera. The chip, developed in part using technology from Groq, a startup specialising in inference that Nvidia licensed in a deal reportedly worth around US$17 billion, targets exactly this workload. The full Vera Rubin platform, which combines the Vera CPU with Rubin GPUs, is set to launch later this year.

Supply is already the constraint

Huang was candid about one problem: supply. “My sense is that we’ll be supply-constrained through the entire life of Vera Rubin,” he said on the call. It’s a telling admission for a product Nvidia is positioning as a major growth pillar. To get ahead of disruptions, Nvidia is spending heavily on the supply chain. The company disclosed that its supply commitments rose to US$119 billion in Q1, up from US$95.2 billion the previous quarter, a significant jump that reflects both confidence in demand and anxiety about a global memory chip crunch.

Nvidia also announced a US$80 billion share repurchase programme and raised its quarterly cash dividend to 25 cents per share, from 1 cent, moves that signal financial confidence even as Huang warned of tightening supply.

The question investors are asking

Despite the beats, Nvidia shares fell 1.6% in extended trading after the results. eMarketer analyst Jacob Bourne captured the mood: “Nvidia delivered another beat, but at this point that’s essentially priced in as it keeps beating quarter after quarter. The lingering question is whether it can convince investors the AI buildout has durability into 2027 and 2028, especially as the narrative shifts toward inference workloads and competing silicon from Google, Amazon, AMD, and Intel.”

Huang pushed back with numbers of his own. He pointed to a growing sub-segment of AI-specific cloud customers whose spend is now roughly equal to the hyperscalers, but growing faster quarter-over-quarter. “We should be growing faster than hyperscale capex,” he said.

The Vera chip is central to that argument. Whether the supply chain cooperates is a different question entirely.

(Image source: Nvidia’s Newsroom)

See Also: The Nvidia H200 China deal survived the Trump-Xi summit–just not in the way anyone expected

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Nvidia’s Vera chip is the US$200 billion bet Jensen Huang doesn’t want you to overlook appeared first on AI News.

  •  

Alibaba is designing AI chips around agents, and that changes what the race is actually about

Alibaba has unveiled a new AI processor built specifically for AI agents, pairing the chip announcement with a multi-year silicon roadmap and a new large language model, signalling that the company is building an integrated AI stack rather than just filling a gap left by US export controls.

The Zhenwu M890, developed by Alibaba’s semiconductor subsidiary T-Head, delivers three times the performance of its predecessor, the Zhenwu 810E, according to the company, as per Reuters report. But the performance jump is less notable than the architectural intent behind the chip: the M890 is purpose-built for AI agents, where software systems must retain long stretches of context, coordinate with other models in real time, and execute complex multi-step tasks with limited human intervention. 

Those demands, heavy on memory bandwidth and inter-model communication, are meaningfully different from what standard inference chips are optimised for. The difference matters because it tells you something about where Alibaba thinks AI compute is heading. The company isn’t designing around today’s dominant use case; it’s building for the workload profile it expects to define enterprise AI over the next several years.

Built for AI agents, not just inference

More significant than the chip itself is the roadmap Alibaba put alongside it. The M890 will be followed by the V900 in the third quarter of 2027, expected to deliver another roughly threefold performance gain, followed by the J900 in the third quarter of 2028. That’s a deliberate, sustained cadence of in-house silicon upgrades that mirrors the kind of tick-tock product cycles Nvidia has used to maintain its lead in AI accelerators.

The parallel to Huawei is worth noting. Huawei laid out a similar chip roadmap for its Ascend line last year, and both announcements reflect the same underlying reality: Chinese technology companies have concluded that depending on foreign silicon, even in scenarios where export restrictions might ease, is a structural risk they cannot accept. The response has been to treat semiconductor development as a long-term capability-building exercise rather than a procurement problem.

Alibaba’s commitment to that exercise is not shallow. The company pledged more than 380 billion yuan, roughly US$53 billion, on cloud and AI infrastructure over three years last year, its largest-ever investment commitment to the sector. The M890 and its successors are downstream of that spending.

Traction that predates the announcement

T-Head said it has shipped more than 560,000 Zhenwu units to date, with over 400 external customers across 20 industries deploying the chips, including automakers and financial services firms. That is a material production footprint, not lab hardware, and it provides Alibaba with real-world deployment data at scale ahead of the M890’s rollout.

The new chip will be available to Chinese enterprise customers through Alibaba Cloud’s domestic model platform, Bailian, packaged inside the Panjiu AL128, a server system that stacks 128 M890 accelerators into a single rack.

The software side of the stack

Alongside the hardware, Alibaba announced Qwen 3.7-Max, the latest version of its flagship large language model, described as engineered for advanced coding and long-running agent tasks. The company said the model can operate continuously for up to 35 hours without performance degradation, a capability specification that only makes sense if you are designing for extended autonomous operation.

The timing is deliberate. Releasing a chip and a model optimised for the same workload class on the same day is a platform play. Alibaba is building a closed loop: its own silicon in T-Head, its own model in Qwen, its own cloud delivery in Bailian. Each component reinforces the others, and the combined stack is designed to reduce enterprise customers’ dependence on any external vendor.

More than half a million chips have been shipped. A successor is arriving in 2027, with another planned for 2028. T-Head is not hedging. At some point, building around US export controls stops being a workaround and starts being a strategy. Alibaba appears to have crossed that line.

(Image source: The White House)

See Also: Alibaba Qwen is challenging proprietary AI model economics

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Alibaba is designing AI chips around agents, and that changes what the race is actually about appeared first on AI News.

  •  
❌