❌

Normal view

The AI skills gap is here, says AI company, and power users are pulling ahead

26 March 2026 at 05:44
Anthropic finds AI isn’t replacing jobs yet, but early data shows growing inequality as experienced users gain an edge, raising concerns about future displacement and workforce divides.
  • βœ‡MIT Technology Review
  • Why this battery company is pivoting to AI Casey Crownhart
    Qichao Hu doesn’t mince words about how he sees the state of the battery industry. β€œAlmost every Western battery company has either died or is going to die. It’s kind of the reality,” he says. Hu is the CEO of SES AI, a Massachusetts-based battery company. It once had aims of making huge amounts of advanced lithium metal batteries for major industries like electric vehiclesβ€”but now the company is placing its bets on AI materials discovery. Hu sees the pivot as an essential one. β€œIt’s just
     

Why this battery company is pivoting to AI

25 March 2026 at 23:02

Qichao Hu doesn’t mince words about how he sees the state of the battery industry. β€œAlmost every Western battery company has either died or is going to die. It’s kind of the reality,” he says.

Hu is the CEO of SES AI, a Massachusetts-based battery company. It once had aims of making huge amounts of advanced lithium metal batteries for major industries like electric vehiclesβ€”but now the company is placing its bets on AI materials discovery.

Hu sees the pivot as an essential one. β€œIt’s just not possible for a Western company to build a sustainable business,” he says. The company is still making some batteries, but only for smaller markets like drones rather than those that would require higher volumes, like EVs. The new focus is the company’s battery materials discovery platformβ€”which it can either license to other battery companies or use to develop materials to sell.Β 

Some leading US EV battery companies have folded in recent months, and others, like SES AI, are making dramatic changes in strategy. This shift in who’s building batteries and where they’re doing it could shape the future geopolitics of energy.Β 

The work that would eventually evolve into SES AI began at MIT, where Hu completed his graduate research. His battery work was aimed at applications in oil and gas exploration. The industry uses sensors that go deep underground, where temperatures can top 120 Β°C (about 250 Β°F). The team hoped to develop a battery that could withstand those high temperatures and last longer on a single charge.Β 

The chosen technology was a solid polymer lithium metal battery. These cells use lithium metal for their anode and a polymer for their electrolyte (the material that ions move through in a battery cell). Together, these components can increase the energy density of a cell significantly, relative to the lithium-ion batteries that are common in personal devices and EVs today. (Lithium-ion batteries generally use a graphite material for their anode and a liquid for the electrolyte.)

That solid-state battery technology became the foundation of Solid Energy, a startup Hu founded that spun out from MIT in 2012 and raised its first private investment in 2013.

The team eventually realized that underground oil exploration was a small market, so after several years of operation they began to focus on electric vehicles, which were starting to come into the mainstream. After the team tweaked the chemistry to work better at lower temperatures, the company built its first pilot facility in Massachusetts and eventually another facility in Shanghai.

By 2021, the battery industry was booming, Hu recalls, and EVs were the hottest industry to be in. There was a ton of interest in next-generation battery technology from major automakers at the time, and Solid Energy started developing technology with GM, Hyundai, and Honda.

Larger vehicles, like SUVs and trucks, seemed like a good fit for next-generation batteries, Hu says. Massive vehicles like the ones Americans like to drive would need lighter batteries so they could have a reasonable range without being prohibitively heavy.

The company also shifted its chemistry focus, and in 2022 it announced a battery with a silicon anode rather than a lithium metal one. That shift could help make the battery easier to manufacture.

Since then, growth in the EV market has slowed, at least in the US, partly because of major pullbacks in funding from the Trump administration. EV tax credits for drivers, a key piece of support pushing Americans toward electric options, ended in late 2025. With the market for large electric cars in trouble, Hu says, β€œnow we have to look at every market.”  

The AI materials discovery platform on which it’s pinning many of its hopes is called Molecular Universe. The company seeks not only to provide its software to other battery companies but also to identify new battery materials and either license them or sell them to those companies.

vials of electrolytes inside a machine at the synthesis foundry
COURTESY OF SES AI

The platform has already identified six new electrolyte materials, according to the company. Hu says one is an additive that could help improve the lifetime of batteries with silicon anodes.Β 

One of the challenges with silicon anodes is that they tend to swell a lot during use, which can cause physical damage and prevent efficient charging and discharging. To address the problem, the industry typically uses a material called fluoroethylene carbonate (FEC), which can help form an elastic film on the anode so the battery can still charge effectively. That additive can degrade at high temperatures, though, producing gases that can harm a battery’s lifetime. The SES platform identified a compound that works like FEC but doesn’t release those gases.

The company’s long history and deep battery knowledge could help make its platform a useful tool, Hu says. He sees the actual model as less crucial than SES’s domain expertise and data from years of making and testing batteries.Β 

β€œBy not actually making the physical battery, we’re actually able to scale and then generate revenue faster,” he says.Β 

But some experts are skeptical about the near-term prospects for AI materials discovery to revive the industry. β€œNew materials development, as much as we thought that was what people wanted (and, frankly, it should be what the cell makers want)β€”I don’t know that that seems to be the real linchpin of the battery industry’s progress,” says Kara Rodby, a technical principal at Volta Energy Technologies, a venture capital firm that focuses on the energy storage industry.

Investors are pulling back, and a slowdown in public support is making things difficult for some parts of the battery industry, she adds: β€œI don’t know that the ability to discover any new material is going to unlock anything new for the battery industry at this point in time.”

  • βœ‡MIT Technology Review
  • The Download: reawakening frozen brains, and the AI Hype Index returns Thomas Macaulay
    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. This scientist rewarmed and studied pieces of his friend’s cryopreserved brainΒ  L. StephenΒ Coles’sΒ brain sits in a vat at a storage facility in Arizona. It has been held there at a temperature of around βˆ’146 degrees Β°C for over a decade,Β largely undisturbed. Before he died in 2014, Coles had the brain frozen with an ambitious goal in mind: reanimation.Β 
     

The Download: reawakening frozen brains, and the AI Hype Index returns

25 March 2026 at 20:47

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

This scientist rewarmed and studied pieces of his friend’s cryopreserved brainΒ 

L. StephenΒ Coles’sΒ brain sits in a vat at a storage facility in Arizona. It has been held there at a temperature of around βˆ’146 degrees Β°C for over a decade,Β largely undisturbed. Before he died in 2014, Coles had the brain frozen with an ambitious goal in mind: reanimation.Β 

His friend, cryobiologist Greg Fahy, believes it could be revived one day. But other experts are less optimistic.Β Β 

Still, Fahy’s research could lead to new ways to study the brain. And using cryopreservation for organ transplantation is becomingΒ a viableΒ reality.Β Β 

Read the full story to find out what the future holds for the technology.Β 

β€”JessicaΒ HamzelouΒ 

The AI Hype IndexΒ 

Separating AI reality from hyped-up fictionΒ isn’tΒ always easy.Β That’sΒ whyΒ we’veΒ created the AI Hype Indexβ€”a simple, at-a-glance summary of everything you need to know about the state of the industry. TakeΒ a look at this month’s edition.Β 
Β 

MIT Technology Review Narrated: how PokΓ©mon Go is giving delivery robots an inch-perfect view of the worldΒ Β 

PokΓ©mon Go was the world’s first augmented-reality megahit. Released in 2016 by Niantic, the AR twist on the juggernaut PokΓ©mon franchise fast became a global phenomenon. β€œ500 million people installed that app in 60 days,” says Brian McClendon, CTO at Niantic Spatial, an AI company that Niantic spun out last year.Β Β 

Now Niantic Spatial is using that vast trove of crowdsourced dataΒ to build a kind of world modelβ€”a buzzyΒ new technologyΒ that grounds the smarts of LLMs in real environments. The firm wants to use it to help robots navigate more precisely.Β 

β€”Will Douglas HeavenΒ 

This is our latestΒ storyΒ to be turned into an MIT Technology Review Narrated podcast, whichΒ we’reΒ publishing each week onΒ SpotifyΒ andΒ Apple Podcasts. Just navigate to MIT Technology Review Narrated on either platform,Β andΒ follow us to get all ourΒ new contentΒ asΒ it’sΒ released.Β 

The next era of space explorationΒ 

Our footprint in the solar system is rapidly expanding. Programs to build permanent Moon bases and find life on Mars have transitioned from science fiction to active space agency missions. The scientists behind them will not only shed new light on theΒ cosmos, butΒ also reveal where humanity is headed.Β 

To examine what the future holds in store, MIT Technology Review features editor Amanda Silverman will sit down today with award-winning science journalist and author Robin George Andrews for an exclusive subscriber-only Roundtable conversation about β€œThe Next Era of Space Exploration.” Register hereΒ to join the session at 16:00 GMT / 12:00 PM ET / 9:00 AM PT.Β 

The must-readsΒ 

I’veΒ combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.Β 

1 OpenAI is shutting down AI video generator SoraΒ Β 
The app attracted at least as much controversy as acclaim. (CNBC)Β 
+ Closing it means saying goodbye to $1 billion from Disney.Β (BBC)Β 
+ OpenAI is cutting back on side projects ahead of an expected IPO.Β (WSJΒ $)Β 
+ ButΒ it’sΒ focusing its efforts on building a fully automated researcher.Β (MIT Technology Review)Β 

2 A judge suspects the Pentagon is illegally punishing AnthropicΒ 
She labelled the DoD’s ban β€œtroubling.” (Bloomberg)Β 
+ Anthropic and the Pentagon are facing off in court.Β (Guardian)Β 
+ The DoD wants AI companies to train on classified data.Β (MIT Technology Review)Β 

3 Meta has been ordered to pay $375 million for endangering children onlineΒ 
Prosecutors said the company knew it put children at risk. (Engadget)Β 
+ Meta is offering its top talent stock options as incentives for its AI push.Β (CNBC)Β 

4 Arm will sell its own computer chips for the first timeΒ 
It’sΒ aimed at dataΒ centersΒ that run AI tasks. (NYTΒ $)Β 
+ Arm stock jumped 13% on the news.Β (CNBC)Β 

5 Manus’s founders have been barred from leaving China following Meta’s takeoverΒ 
Beijing is reviewing the $2 billion acquisition of the AI startup.Β (FTΒ $)Β 

6 Baltimore has suedΒ xAIΒ overΒ Grok’sΒ fake nude imagesΒ Β 
The chatbot allegedly violated consumer protections. (Guardian)Β 
+Β There’sΒ a big market for pornographic deepfakes of real women.Β (MIT Technology Review)Β 

7 NASA plans to send a nuclear-powered spacecraft to Mars in 2028Β 
It’llΒ take a payload of Ingenuity-class helicopters to the Red Planet. (NYTΒ $)Β 
+ NASA also wants to put a $20 billionΒ baseΒ on the Moon.Β (The Verge)Β 

8 A company is secretly turning Zoom meetings into AI-generated podcastsΒ 
WebinarTVΒ turns the calls into content without telling anyone. (404 Media)Β 

9 Iranian volunteers have built their own missile warning mapΒ 
It fills the gap left by Iran’s lack of a public emergency alert tool. (WiredΒ $)Β 
+ Here’s where OpenAI’s tech could show up in Iran.Β (MIT Technology Review)Β 

10 A nonprofit is sending basic income payments to AI-impacted workersΒ 
It’sΒ starting by giving 25-50 people $1,000 per month. (Gizmodo)Β 

Quote of the dayΒ 

β€œI amΒ first and foremostΒ a scientist. My goal is to understand nature. But doing science is, sort of, like reading the mind of God.” 

β€”DeepMind CEO Demis Hassabis shares his approach to AI strategy withΒ the FT.Β 

One More ThingΒ 

many ui windows framing different views of an asteroid on the way to Earth
EVA REDAMONTI

Inside the hunt for the most dangerous asteroid everΒ Β 

As asteroid 2024 YR4 hurtled toward Earth, astronomersΒ determinedΒ that this massive rock posed a higher risk of impact than any object of its size in recorded history. Then, just as quickly as history was made, experts declared that the danger had passed.Β 

This is the inside story of the network of global scientists who found, followed, planned for, and finally dismissed the most dangerous asteroid ever foundβ€”all under the tightest of timelines and with the highest of stakes.Β Find out how they did it.Β 

β€”Robin George AndrewsΒ 

We can still haveΒ nice thingsΒ 

A place for comfort, fun and distraction to brighten up your day. (Got any ideas?Β Drop me a line.)Β 
Β 
+ Soothe subscription fatigue withΒ this simple cancellation tool.Β 
+Β Takashi Murakami’sΒ reimagined MonetsΒ are pop-art magic.Β 
+ Jump into a rabbit hole with this app thatΒ visualizes linksΒ between Wikipedia pages.Β 
+ ThisΒ playful lynxΒ that snatched the top prize in a photo competition is a delight.Β 

  • βœ‡cs.AI, q-bio.NC updates on arXiv.org
  • Black Hole-Inspired Horizon Model for Neural Signal Dynamics E. Canessa
    arXiv:2603.22297v1 Announce Type: new Abstract: Electroencephalographic (EEG) signals provide macroscopic observables of complex neural dynamics. We introduce a horizon-inspired framework in which measured EEG signals are modeled as projections of a complex wave-like representation constrained by an effective boundary analogous to an event horizon. In this formulation the signal amplitude obeys a renormalization-group scaling relation while EEG spectral entropy parameterizes the accessibility o
     

Black Hole-Inspired Horizon Model for Neural Signal Dynamics

arXiv:2603.22297v1 Announce Type: new Abstract: Electroencephalographic (EEG) signals provide macroscopic observables of complex neural dynamics. We introduce a horizon-inspired framework in which measured EEG signals are modeled as projections of a complex wave-like representation constrained by an effective boundary analogous to an event horizon. In this formulation the signal amplitude obeys a renormalization-group scaling relation while EEG spectral entropy parameterizes the accessibility of observable modes. The resulting solutions generate oscillatory structures whose geometry and spectral signatures can be explored through signal analysis and sonification. This mapping between entropy-based neural observables and wave-like signal representations provides a physically motivated framework linking entropy measures, scale-dependent dynamics, and observable neural oscillations, and suggests testable connections between spectral entropy and the amplitude scaling of EEG modes.

Ca2+ transient detection and segmentation with the Astronomically motivated algorithm for Background Estimation And Transient Segmentation (Astro-BEATS)

arXiv:2603.22311v1 Announce Type: new Abstract: Fluorescence-based Ca$^{2+}$-imaging is a powerful tool for studying localized neuronal activity, including miniature Synaptic Calcium Transients, providing real-time insights into synaptic activity. These transients induce only subtle changes in the fluorescence signal, often barely above baseline, which poses a significant challenge for automated synaptic transient detection and segmentation. Detecting astronomical transients similarly requires efficient algorithms that will remain robust over a large field of view with varying noise properties. We leverage techniques used in astronomical transient detection for miniature Synaptic Calcium Transient detection in fluorescence microscopy. We present Astro-BEATS, an automatic miniature Synaptic Calcium Transient segmentation algorithm that incorporates image estimation and source-finding techniques used in astronomy and designed for Ca$^{2+}$-imaging videos. Astro-BEATS outperforms current threshold-based approaches for synaptic Ca$^{2+}$ transient detection and segmentation. The produced segmentation masks can be used to train a supervised deep learning algorithm for improved synaptic Ca$^{2+}$ transient detection in Ca$^{2+}$-imaging data. The speed of Astro-BEATS and its applicability to previously unseen datasets without re-optimization makes it particularly useful for generating training datasets for deep learning-based approaches.

Computational Arbitrage in AI Model Markets

arXiv:2603.22404v1 Announce Type: new Abstract: Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a budget for a verifiable solution. An arbitrageur efficiently allocates inference budget across providers to undercut the market, thus creating a competitive offering with no model-development risk. In this work, we initiate the study of arbitrage in AI model markets, empirically demonstrating the viability of arbitrage and illustrating its economic consequences. We conduct an in-depth case study of SWE-bench GitHub issue resolution using two representative models, GPT-5 mini and DeepSeek v3.2. In this verifiable domain, simple arbitrage strategies generate net profit margins of up to 40%. Robust arbitrage strategies that generalize across different domains remain profitable. Distillation further creates strong arbitrage opportunities, potentially at the expense of the teacher model's revenue. Multiple competing arbitrageurs drive down consumer prices, reducing the marginal revenue of model providers. At the same time, arbitrage reduces market segmentation and facilitates market entry for smaller model providers by enabling earlier revenue capture. Our results suggest that arbitrage can be a powerful force in AI model markets with implications for model development, distillation, and deployment.

Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length

arXiv:2603.22608v1 Announce Type: new Abstract: Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances. For example, analysing the overall sentiment of a number of movie reviews requires an LLM to process the sentiment of each review individually in order to provide a final aggregated answer. While LLM performance on such individual tasks is generally high, there has been little research on how LLMs perform when dealing with multi-instance inputs. In this paper, we perform a comprehensive evaluation of the multi-instance processing (MIP) ability of LLMs for tasks in which they excel individually. The results show that all LLMs follow a pattern of slight performance degradation for small numbers of instances (approximately 20-100), followed by a performance collapse on larger instance counts. Crucially, our analysis shows that while context length is associated with this degradation, the number of instances has a stronger effect on the final results. This finding suggests that when optimising LLM performance for MIP, attention should be paid to both context length and, in particular, instance count.

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

arXiv:2603.22813v1 Announce Type: new Abstract: Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.

CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models

By: Youzhi Liu Β· Li Gao Β· Liu Liu Β· Mingyang Lv Β· Yang Cai
25 March 2026 at 12:00
arXiv:2603.22846v1 Announce Type: new Abstract: Embodied Visual Tracking (EVT), a core dynamic task in embodied intelligence, requires an agent to precisely follow a language-specified target. Yet most existing methods rely on single-agent imitation learning, suffering from costly expert data and limited generalization due to static training environments. Inspired by competition-driven capability evolution, we propose CoMaTrack, a competitive game-theoretic multi-agent reinforcement learning framework that trains agents in a dynamic adversarial setting with competitive subtasks, yielding stronger adaptive planning and interference-resilient strategies. We further introduce CoMaTrack-Bench, the first benchmark for competitive EVT, featuring game scenarios between a tracker and adaptive opponents across diverse environments and instructions, enabling standardized robustness evaluation under active adversarial interactions. Experiments show that CoMaTrack achieves state-of-the-art results on both standard benchmarks and CoMaTrack-Bench. Notably, a 3B VLM trained with our framework surpasses previous single-agent imitation learning methods based on 7B models on the challenging EVT-Bench, achieving 92.1% in STT, 74.2% in DT, and 57.5% in AT. The benchmark code will be available at https://github.com/wlqcode/CoMaTrack-Bench

On the use of Aggregation Operators to improve Human Identification using Dental Records

arXiv:2603.23003v1 Announce Type: new Abstract: The comparison of dental records is a standardized technique in forensic dentistry used to speed up the identification of individuals in multiple-comparison scenarios. Specifically, the odontogram comparison is a procedure to compute criteria that will be used to perform a ranking. State-of-the-art automatic methods either make use of simple techniques, without utilizing the full potential of the information obtained from a comparison, or their internal behavior is not known due to the lack of peer-reviewed publications. This work aims to design aggregation mechanisms to automatically compare pairs of dental records that can be understood and validated by experts, improving the current methods. To do so, we introduce different aggregation approaches using the state-of-the-art codification, based on seven different criteria. In particular, we study the performance of i) data-driven lexicographical order-based aggregations, ii) well-known fuzzy logic aggregation methods and iii) machine learning techniques as aggregation mechanisms. To validate our proposals, 215 forensic cases from two different populations have been used. The results obtained show how the use of white-box machine learning techniques as aggregation models (average ranking from 2.02 to 2.21) are able to improve the state-of-the-art (average ranking of 3.91) without compromising the explainability and interpretability of the method.

Minibal: Balanced Game-Playing Without Opponent Modeling

arXiv:2603.23059v1 Announce Type: new Abstract: Recent advances in game AI, such as AlphaZero and Ath\'enan, have achieved superhuman performance across a wide range of board games. While highly powerful, these agents are ill-suited for human-AI interaction, as they consistently overwhelm human players, offering little enjoyment and limited educational value. This paper addresses the problem of balanced play, in which an agent challenges its opponent without either dominating or conceding. We introduce Minibal (Minimize & Balance), a variant of Minimax specifically designed for balanced play. Building on this concept, we propose several modifications of the Unbounded Minimax algorithm explicitly aimed at discovering balanced strategies. Experiments conducted across seven board games demonstrate that one variant consistently achieves the most balanced play, with average outcomes close to perfect balance. These results establish Minibal as a promising foundation for designing AI agents that are both challenging and engaging, suitable for both entertainment and serious games.

Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models

arXiv:2603.23149v1 Announce Type: new Abstract: Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, current approaches relying on visual simulation incur prohibitive latencies, often exceeding several seconds per step. In this work, we challenge the assumption that visual processing is necessary for failure prevention. We show that a trained policy's latent state, combined with its planned actions, already encodes sufficient information to anticipate action outcomes, making visual simulation redundant for failure prevention. To this end, we introduce DILLO (DIstiLLed Language-ActiOn World Model), a fast steering layer that shifts the paradigm from "simulate-then-act" to "describe-then-act." DILLO is trained via cross-modal distillation, where a privileged Vision Language Model teacher annotates offline trajectories and a latent-conditioned Large Language Model student learns to predict semantic outcomes. This creates a text-only inference path, bypassing heavy visual generation entirely, achieving a 14x speedup over baselines. Experiments on MetaWorld and LIBERO demonstrate that DILLO produces high-fidelity descriptions of the next state and is able to steer the policy, improving episode success rate by up to 15 pp and 9.3 pp on average across tasks.

Automated Microservice Pattern Instance Detection Using Infrastructure-as-Code Artifacts and Large Language Models

arXiv:2502.04188v1 Announce Type: cross Abstract: Documenting software architecture is essential to preserve architecture knowledge, even though it is frequently costly. Architecture pattern instances, including microservice pattern instances, provide important structural software information. Practitioners should document this information to prevent knowledge vaporization. However, architecture patterns may not be detectable by analyzing source code artifacts, requiring the analysis of other types of artifacts. Moreover, many existing pattern detection instance approaches are complex to extend. This article presents our ongoing PhD research, early experiments, and a prototype for a tool we call MicroPAD for automating the detection of microservice pattern instances. The prototype uses Large Language Models (LLMs) to analyze Infrastructure-as-Code (IaC) artifacts to aid detection, aiming to keep costs low and maximize the scope of detectable patterns. Early experiments ran the prototype thrice in 22 GitHub projects. We verified that 83\% of the patterns that the prototype identified were in the project. The costs of detecting the pattern instances were minimal. These results indicate that the approach is likely viable and, by lowering the entry barrier to automating pattern instance detection, could help democratize developer access to this category of architecture knowledge. Finally, we present our overall research methodology, planned future work, and an overview of MicroPAD's potential industrial impact.

Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning

arXiv:2603.22317v1 Announce Type: cross Abstract: Graph-structured data typically exhibits complex topological heterogeneity, making it difficult to model accurately within a single Riemannian manifold. While emerging mixed-curvature methods attempt to capture such diversity, they often rely on implicit, task-driven routing that lacks fundamental geometric grounding. To address this challenge, we propose a Geometric Mixture-of-Experts framework (GeoMoE) that adaptively fuses node representations across diverse Riemannian spaces to better accommodate multi-scale topological structures. At its core, GeoMoE leverages Ollivier-Ricci Curvature (ORC) as an intrinsic geometric prior to orchestrate the collaboration of specialized experts. Specifically, we design a graph-aware gating network that assigns node-specific fusion weights, regularized by a curvature-guided alignment loss to ensure interpretable and geometry-consistent routing. Additionally, we introduce a curvature-aware contrastive objective that promotes geometric discriminability by constructing positive and negative pairs according to curvature consistency. Extensive experiments on six benchmark datasets demonstrate that GeoMoE outperforms state-of-the-art baselines across diverse graph types.

From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs

arXiv:2603.22321v1 Announce Type: cross Abstract: The recent advancements introduced by Large Language Models (LLMs) have transformed how Artificial Intelligence (AI) can support complex, real world tasks, pushing research outside the text boundaries towards multi modal contexts and leading to Multimodal Large Language Models (MLMs). Given the current adoption of LLM based assistants in solving technical or domain specific problems, the natural continuation of this trend is to extend the input domains of these assistants exploiting MLMs. Ideally, these MLMs should be used as real time assistants in procedural tasks, hopefully integrating a view of the environment where the user being assisted is, or even better sharing the same point of view via Virtual Reality (VR) or Augmented Reality (AR) supports, to reason over the same scenario the user is experiencing. With this work, we aim at evaluating the quality of currently openly available MLMs to provide this kind of assistance on technical tasks. To this end, we annotated a data set of furniture assembly with step by step labels and manual references: the Manual to Action Dataset (M2AD). We used this dataset to assess (1) to which extent the reasoning abilities of MLMs can be used to reduce the need for detailed labelling, allowing for more efficient, cost effective annotation practices, (2) whether MLMs are able to track the progression of assembly steps (3) and whether MLMs can refer correctly to the instruction manual pages. Our results showed that while some models understand procedural sequences, their performance is limited by architectural and hardware constraints, highlighting the need for multi image and interleaved text image reasoning.

A Direct Classification Approach for Reliable Wind Ramp Event Forecasting under Severe Class Imbalance

arXiv:2603.22326v1 Announce Type: cross Abstract: Decision support systems are essential for maintaining grid stability in low-carbon power systems, such as wind power plants, by providing real-time alerts to control room operators regarding potential events, including Wind Power Ramp Events (WPREs). These early warnings enable the timely initiation of more detailed system stability assessments and preventive actions. However, forecasting these events is challenging due to the inherent class imbalance in WPRE datasets, where ramp events are less frequent (typically less than 15\% of observed events) compared to normal conditions. Ignoring this characteristic undermines the performance of conventional machine learning models, which often favor the majority class. This paper introduces a novel methodology for WPRE forecasting as a multivariate time series classification task and proposes a data preprocessing strategy that extracts features from recent power observations and masks unavailable ramp information, making it integrable with traditional real-time ramp identification tools. Particularly, the proposed methodology combines majority-class undersampling and ensemble learning to enhance wind ramp event forecasting under class imbalance. Numerical simulations conducted on a real-world dataset demonstrate the superiority of our approach, achieving over 85% accuracy and 88% weighted F1 score, outperforming benchmark classifiers.

AgentSLR: Automating Systematic Literature Reviews in Epidemiology with Agentic AI

arXiv:2603.22327v1 Announce Type: cross Abstract: Systematic literature reviews are essential for synthesizing scientific evidence but are costly, difficult to scale and time-intensive, creating bottlenecks for evidence-based policy. We study whether large language models can automate the complete systematic review workflow, from article retrieval, article screening, data extraction to report synthesis. Applied to epidemiological reviews of nine WHO-designated priority pathogens and validated against expert-curated ground truth, our open-source agentic pipeline (AgentSLR) achieves performance comparable to human researchers while reducing review time from approximately 7 weeks to 20 hours (a 58x speed-up). Our comparison of five frontier models reveals that performance on SLR is driven less by model size or inference cost than by each model's distinctive capabilities. Through human-in-the-loop validation, we identify key failure modes. Our results demonstrate that agentic AI can substantially accelerate scientific evidence synthesis in specialised domains.

Beyond the Mean: Distribution-Aware Loss Functions for Bimodal Regression

arXiv:2603.22328v1 Announce Type: cross Abstract: Despite the strong predictive performance achieved by machine learning models across many application domains, assessing their trustworthiness through reliable estimates of predictive confidence remains a critical challenge. This issue arises in scenarios where the likelihood of error inferred from learned representations follows a bimodal distribution, resulting from the coexistence of confident and ambiguous predictions. Standard regression approaches often struggle to adequately express this predictive uncertainty, as they implicitly assume unimodal Gaussian noise, leading to mean-collapse behavior in such settings. Although Mixture Density Networks (MDNs) can represent different distributions, they suffer from severe optimization instability. We propose a family of distribution-aware loss functions integrating normalized RMSE with Wasserstein and Cram\'er distances. When applied to standard deep regression models, our approach recovers bimodal distributions without the volatility of mixture models. Validated across four experimental stages, our results show that the proposed Wasserstein loss establishes a new Pareto efficiency frontier: matching the stability of standard regression losses like MSE in unimodal tasks while reducing Jensen-Shannon Divergence by 45% on complex bimodal datasets. Our framework strictly dominates MDNs in both fidelity and robustness, offering a reliable tool for aleatoric uncertainty estimation in trustworthy AI systems.

Large Language Models for Missing Data Imputation: Understanding Behavior, Hallucination Effects, and Control Mechanisms

arXiv:2603.22332v1 Announce Type: cross Abstract: Data imputation is a cornerstone technique for handling missing values in real-world datasets, which are often plagued by missingness. Despite recent progress, prior studies on Large Language Models-based imputation remain limited by scalability challenges, restricted cross-model comparisons, and evaluations conducted on small or domain-specific datasets. Furthermore, heterogeneous experimental protocols and inconsistent treatment of missingness mechanisms (MCAR, MAR, and MNAR) hinder systematic benchmarking across methods. This work investigates the robustness of Large Language Models for missing data imputation in tabular datasets using a zero-shot prompt engineering approach. To this end, we present a comprehensive benchmarking study comparing five widely used LLMs against six state-of-the-art imputation baselines. The experimental design evaluates these methods across 29 datasets (including nine synthetic datasets) under MCAR, MAR, and MNAR mechanisms, with missing rates of up to 20\%. The results demonstrate that leading LLMs, particularly Gemini 3.0 Flash and Claude 4.5 Sonnet, consistently achieve superior performance on real-world open-source datasets compared to traditional methods. However, this advantage appears to be closely tied to the models' prior exposure to domain-specific patterns learned during pre-training on internet-scale corpora. In contrast, on synthetic datasets, traditional methods such as MICE outperform LLMs, suggesting that LLM effectiveness is driven by semantic context rather than purely statistical reconstruction. Furthermore, we identify a clear trade-off: while LLMs excel in imputation quality, they incur significantly higher computational time and monetary costs. Overall, this study provides a large-scale comparative analysis, positioning LLMs as promising semantics-driven imputers for complex tabular data.
❌