❌

Normal view

TechCrunch Disrupt 2026 Early Bird ticket rates end May 29

26 May 2026 at 22:00
Save up to $410 on your TechCrunch Disrupt 2026 pass before prices increase on May 29 at 11:59 p.m. PT. Register here to join the tech epicenter in San Francisco.
  • ✇AI News
  • Nvidia’s Vera chip is the US$200 billion bet Jensen Huang doesn’t want you to overlook Dashveenjit Kaur
    The Nvidia Vera chip is rarely the headline when earnings beat estimates, but it should be. When Nvidia reported Q1 revenue of US$81.62 billion on Wednesday, beating analyst estimates of US$78.86 billion, and guided Q2 at US$91 billion–well above Wall Street’s US$86.84 billion forecast–the numbers did what Nvidia numbers always do: dominate the room.  But buried in CEO Jensen Huang’s conference call with analysts was something more strategically interesting than another quarterly beat. Huang
     

Nvidia’s Vera chip is the US$200 billion bet Jensen Huang doesn’t want you to overlook

21 May 2026 at 16:00

The Nvidia Vera chip is rarely the headline when earnings beat estimates, but it should be. When Nvidia reported Q1 revenue of US$81.62 billion on Wednesday, beating analyst estimates of US$78.86 billion, and guided Q2 at US$91 billion–well above Wall Street’s US$86.84 billion forecast–the numbers did what Nvidia numbers always do: dominate the room. 

But buried in CEO Jensen Huang’s conference call with analysts was something more strategically interesting than another quarterly beat. Huang told analysts that Nvidia’s new Vera central processors unlock access to a US$200 billion market, one that sits entirely outside the US$1 trillion the company has already forecast from its Blackwell and Rubin AI GPU lineup between 2025 and 2027. 

He expects Vera chip revenue to hit US$20 billion by the end of this fiscal year. “I expect (Vera) to be the second largest” sales contributor, Huang said during the call.

That’s not a footnote. That’s a second front.

The Vera chip and the inference pivot

The reason Nvidia needs a second front is straightforward: its biggest customers are building their own. Google, Amazon, and Microsoft–collectively expected to pour more than US$700 billion into AI infrastructure this year, up sharply from around US$400 billion in 2025, are simultaneously pouring funds into custom silicon to run AI models. Intel and AMD are also touting CPUs as a credible play for inference workloads. 

The narrative in the chip industry has shifted from who can train the biggest model to who can serve it cheapest and fastest. Inference is where Nvidia’s GPU dominance is most exposed. Training large models is still firmly Nvidia territory, but inference, generating answers at scale, in real time, is increasingly where custom chips from Google’s TPU line, Amazon’s Trainium and others are making their case.

Nvidia’s answer is Vera. The chip, developed in part using technology from Groq, a startup specialising in inference that Nvidia licensed in a deal reportedly worth around US$17 billion, targets exactly this workload. The full Vera Rubin platform, which combines the Vera CPU with Rubin GPUs, is set to launch later this year.

Supply is already the constraint

Huang was candid about one problem: supply. “My sense is that we’ll be supply-constrained through the entire life of Vera Rubin,” he said on the call. It’s a telling admission for a product Nvidia is positioning as a major growth pillar. To get ahead of disruptions, Nvidia is spending heavily on the supply chain. The company disclosed that its supply commitments rose to US$119 billion in Q1, up from US$95.2 billion the previous quarter, a significant jump that reflects both confidence in demand and anxiety about a global memory chip crunch.

Nvidia also announced a US$80 billion share repurchase programme and raised its quarterly cash dividend to 25 cents per share, from 1 cent, moves that signal financial confidence even as Huang warned of tightening supply.

The question investors are asking

Despite the beats, Nvidia shares fell 1.6% in extended trading after the results. eMarketer analyst Jacob Bourne captured the mood: “Nvidia delivered another beat, but at this point that’s essentially priced in as it keeps beating quarter after quarter. The lingering question is whether it can convince investors the AI buildout has durability into 2027 and 2028, especially as the narrative shifts toward inference workloads and competing silicon from Google, Amazon, AMD, and Intel.”

Huang pushed back with numbers of his own. He pointed to a growing sub-segment of AI-specific cloud customers whose spend is now roughly equal to the hyperscalers, but growing faster quarter-over-quarter. “We should be growing faster than hyperscale capex,” he said.

The Vera chip is central to that argument. Whether the supply chain cooperates is a different question entirely.

(Image source: Nvidia’s Newsroom)

See Also: The Nvidia H200 China deal survived the Trump-Xi summit–just not in the way anyone expected

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Nvidia’s Vera chip is the US$200 billion bet Jensen Huang doesn’t want you to overlook appeared first on AI News.

  • ✇AI News
  • Alibaba is designing AI chips around agents, and that changes what the race is actually about Dashveenjit Kaur
    Alibaba has unveiled a new AI processor built specifically for AI agents, pairing the chip announcement with a multi-year silicon roadmap and a new large language model, signalling that the company is building an integrated AI stack rather than just filling a gap left by US export controls. The Zhenwu M890, developed by Alibaba’s semiconductor subsidiary T-Head, delivers three times the performance of its predecessor, the Zhenwu 810E, according to the company, as per Reuters report. But the p
     

Alibaba is designing AI chips around agents, and that changes what the race is actually about

20 May 2026 at 18:00

Alibaba has unveiled a new AI processor built specifically for AI agents, pairing the chip announcement with a multi-year silicon roadmap and a new large language model, signalling that the company is building an integrated AI stack rather than just filling a gap left by US export controls.

The Zhenwu M890, developed by Alibaba’s semiconductor subsidiary T-Head, delivers three times the performance of its predecessor, the Zhenwu 810E, according to the company, as per Reuters report. But the performance jump is less notable than the architectural intent behind the chip: the M890 is purpose-built for AI agents, where software systems must retain long stretches of context, coordinate with other models in real time, and execute complex multi-step tasks with limited human intervention. 

Those demands, heavy on memory bandwidth and inter-model communication, are meaningfully different from what standard inference chips are optimised for. The difference matters because it tells you something about where Alibaba thinks AI compute is heading. The company isn’t designing around today’s dominant use case; it’s building for the workload profile it expects to define enterprise AI over the next several years.

Built for AI agents, not just inference

More significant than the chip itself is the roadmap Alibaba put alongside it. The M890 will be followed by the V900 in the third quarter of 2027, expected to deliver another roughly threefold performance gain, followed by the J900 in the third quarter of 2028. That’s a deliberate, sustained cadence of in-house silicon upgrades that mirrors the kind of tick-tock product cycles Nvidia has used to maintain its lead in AI accelerators.

The parallel to Huawei is worth noting. Huawei laid out a similar chip roadmap for its Ascend line last year, and both announcements reflect the same underlying reality: Chinese technology companies have concluded that depending on foreign silicon, even in scenarios where export restrictions might ease, is a structural risk they cannot accept. The response has been to treat semiconductor development as a long-term capability-building exercise rather than a procurement problem.

Alibaba’s commitment to that exercise is not shallow. The company pledged more than 380 billion yuan, roughly US$53 billion, on cloud and AI infrastructure over three years last year, its largest-ever investment commitment to the sector. The M890 and its successors are downstream of that spending.

Traction that predates the announcement

T-Head said it has shipped more than 560,000 Zhenwu units to date, with over 400 external customers across 20 industries deploying the chips, including automakers and financial services firms. That is a material production footprint, not lab hardware, and it provides Alibaba with real-world deployment data at scale ahead of the M890’s rollout.

The new chip will be available to Chinese enterprise customers through Alibaba Cloud’s domestic model platform, Bailian, packaged inside the Panjiu AL128, a server system that stacks 128 M890 accelerators into a single rack.

The software side of the stack

Alongside the hardware, Alibaba announced Qwen 3.7-Max, the latest version of its flagship large language model, described as engineered for advanced coding and long-running agent tasks. The company said the model can operate continuously for up to 35 hours without performance degradation, a capability specification that only makes sense if you are designing for extended autonomous operation.

The timing is deliberate. Releasing a chip and a model optimised for the same workload class on the same day is a platform play. Alibaba is building a closed loop: its own silicon in T-Head, its own model in Qwen, its own cloud delivery in Bailian. Each component reinforces the others, and the combined stack is designed to reduce enterprise customers’ dependence on any external vendor.

More than half a million chips have been shipped. A successor is arriving in 2027, with another planned for 2028. T-Head is not hedging. At some point, building around US export controls stops being a workaround and starts being a strategy. Alibaba appears to have crossed that line.

(Image source: The White House)

See Also: Alibaba Qwen is challenging proprietary AI model economics

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Alibaba is designing AI chips around agents, and that changes what the race is actually about appeared first on AI News.

  • ✇AI News
  • IBM: How robust AI governance protects enterprise margins Ryan Daws
    To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure. When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely. At the initial product stage, exert
     

IBM: How robust AI governance protects enterprise margins

10 April 2026 at 21:57

To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure.

When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely.

At the initial product stage, exerting tight corporate control often feels highly advantageous. Closed development environments iterate quickly and tightly manage the end-user experience. They capture and concentrate financial value within a single corporate entity, an approach that functions adequately during early product development cycles.

However, IBM’s analysis highlights that expectations change entirely when a technology solidifies into a foundational layer. Once other institutional frameworks, external markets, and broad operational systems rely on the software, the prevailing standards adapt to a new reality. At infrastructure scale, embracing openness ceases to be an ideological stance and becomes a highly practical necessity.

AI is currently crossing this threshold within the enterprise architecture stack. Models are increasingly embedded directly into the ways organisations secure their networks, author source code, execute automated decisions, and generate commercial value. AI functions less as an experimental utility and more as core operational infrastructure.

The recent limited preview of Anthropic’s Claude Mythos model brings this reality into sharper focus for enterprise executives managing risk. Anthropic reports that this specific model can discover and exploit software vulnerabilities at a level matching few human experts.

In response to this power, Anthropic launched Project Glasswing, a gated initiative designed to place these advanced capabilities directly into the hands of network defenders first. From IBM’s perspective, this development forces technology officers to confront immediate structural vulnerabilities. If autonomous models possess the capability to write exploits and shape the overall security environment, Thomas notes that concentrating the understanding of these systems within a small number of technology vendors invites severe operational exposure.

With models achieving infrastructure status, IBM argues the primary issue is no longer exclusively what these machine learning applications can execute. The priority becomes how these systems are constructed, governed, inspected, and actively improved over extended periods.

As underlying frameworks grow in complexity and corporate importance, maintaining closed development pipelines becomes exceedingly difficult to defend. No single vendor can successfully anticipate every operational requirement, adversarial attack vector, or system failure mode.

Implementing opaque AI structures introduces heavy friction across existing network architecture. Connecting closed proprietary models with established enterprise vector databases or highly sensitive internal data lakes frequently creates massive troubleshooting bottlenecks. When anomalous outputs occur or hallucination rates spike, teams lack the internal visibility required to diagnose whether the error originated in the retrieval-augmented generation pipeline or the base model weights.

Integrating legacy on-premises architecture with highly gated cloud models also introduces severe latency into daily operations. When enterprise data governance protocols strictly prohibit sending sensitive customer information to external servers, technology teams are left attempting to strip and anonymise datasets before processing. This constant data sanitisation creates enormous operational drag. 

Furthermore, the spiralling compute costs associated with continuous API calls to locked models erode the exact profit margins these autonomous systems are supposed to enhance. The opacity prevents network engineers from accurately sizing hardware deployments, forcing companies into expensive over-provisioning agreements to maintain baseline functionality.

Why open-source AI is essential for operational resilience

Restricting access to powerful applications is an understandable human instinct that closely resembles caution. Yet, as Thomas points out, at massive infrastructure scale, security typically improves through rigorous external scrutiny rather than through strict concealment.

This represents the enduring lesson of open-source software development. Open-source code does not eliminate enterprise risk. Instead, IBM maintains it actively changes how organisations manage that risk. An open foundation allows a wider base of researchers, corporate developers, and security defenders to examine the architecture, surface underlying weaknesses, test foundational assumptions, and harden the software under real-world conditions.

Within cybersecurity operations, broad visibility is rarely the enemy of operational resilience. In fact, visibility frequently serves as a strict prerequisite for achieving that resilience. Technologies deemed highly important tend to remain safer when larger populations can challenge them, inspect their logic, and contribute to their continuous improvement.

Thomas addresses one of the oldest misconceptions regarding open-source technology: the belief that it inevitably commoditises corporate innovation. In practical application, open infrastructure typically pushes market competition higher up the technology stack. Open systems transfer financial value rather than destroying it.

As common digital foundations mature, the commercial value relocates toward complex implementation, system orchestration, continuous reliability, trust mechanics, and specific domain expertise. IBM’s position asserts that the long-term commercial winners are not those who own the base technological layer, but rather the organisations that understand how to apply it most effectively.

We have witnessed this identical pattern play out across previous generations of enterprise tooling, cloud infrastructure, and operating systems. Open foundations historically expanded developer participation, accelerated iterative improvement, and birthed entirely new, larger markets built on top of those base layers. Enterprise leaders increasingly view open-source as highly important for infrastructure modernisation and emerging AI capabilities. IBM predicts that AI is highly likely to follow this exact historical trajectory.

Looking across the broader vendor ecosystem, leading hyperscalers are adjusting their business postures to accommodate this reality. Rather than engaging in a pure arms race to build the largest proprietary black boxes, highly profitable integrators are focusing heavily on orchestration tooling that allows enterprises to swap out underlying open-source models based on specific workload demands. Highlighting its ongoing leadership in this space, IBM is a key sponsor of this year’s AI & Big Data Expo North America, where these evolving strategies for open enterprise infrastructure will be a primary focus.

This approach completely sidesteps restrictive vendor lock-in and allows companies to route less demanding internal queries to smaller and highly efficient open models, preserving expensive compute resources for complex customer-facing autonomous logic. By decoupling the application layer from the specific foundation model, technology officers can maintain operational agility and protect their bottom line.

The future of enterprise AI demands transparent governance

Another pragmatic reason for embracing open models revolves around product development influence. IBM emphasises that narrow access to underlying code naturally leads to narrow operational perspectives. In contrast, who gets to participate directly shapes what applications are eventually built. 

Providing broad access enables governments, diverse institutions, startups, and varied researchers to actively influence how the technology evolves and where it is commercially applied. This inclusive approach drives functional innovation while simultaneously building structural adaptability and necessary public legitimacy.

As Thomas argues, once autonomous AI assumes the role of core enterprise infrastructure, relying on opacity can no longer serve as the organising principle for system safety. The most reliable blueprint for secure software has paired open foundations with broad external scrutiny, active code maintenance, and serious internal governance.

As AI permanently enters its infrastructure phase, IBM contends that identical logic increasingly applies directly to the foundation models themselves. The stronger the corporate reliance on a technology, the stronger the corresponding case for demanding openness.

If these autonomous workflows are truly becoming foundational to global commerce, then transparency ceases to be a subject of casual debate. According to IBM, it is an absolute, non-negotiable design requirement for any modern enterprise architecture.

See also: Why companies like Apple are building AI agents with limits

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post IBM: How robust AI governance protects enterprise margins appeared first on AI News.

4 days left to save close to $500 on TechCrunch Disrupt 2026 passes

7 April 2026 at 22:00
Four days left to save up to $482 on your TechCrunch Disrupt 2026 ticket. These low rates will disappear on April 10 at 11:59 p.m. PT. Register now.

Presentation: When Every Bit Counts: How Valkey Rebuilt Its Hashtable for Modern Hardware

7 April 2026 at 21:40

Madelyn Olson discusses the evolution of Valkey's data structures, moving away from "textbook" pointer-chasing HashMaps to more cache-aware designs. She explains the implementation of "Swedish" tables to maximize memory density. She shares insights on systems intuition, memory prefetching, and the rigorous testing needed for mission-critical caches.

By Madelyn Olson

Hyperscale Power is the latest startup to challenge 140-year-old transformer tech

10 March 2026 at 21:00
Startup Hyperscale Power is developing technology that promises to shrink power transformers, freeing up precious space within data centers.
❌