❌

Normal view

Telehealth Delivery of the Homeostasis–Enrichment–Plasticity Approach for Premature Infants With Developmental Risks: Exploratory Feasibility Study

Background: Preterm delivery is an increasing worldwide health concern linked to increased neurodevelopmental risks. Early intervention is crucial for harnessing neuroplasticity to enhance developmental and functional performance outcomes; however, access to early intervention is frequently hindered by logistical, financial, and labor constraints. The Homeostasis–Enrichment–Plasticity (HEP) Approach is a family-centered early intervention model based on enriched environments, designed to improve infants’ sensory-motor, cognitive, and socio-emotional development. Objective: This study aimed to assess the feasibility, safety, acceptability, and outcomes sensitivity to change of implementing the HEP Approach through telehealth for premature infants at developmental risk. Methods: A pre-post exploratory feasibility study was performed, including 16 preterm infants (aged 4-12 months corrected age), of whom 14 completed the study. The 12-week intervention included weekly remote sessions focused on environmental enrichment, active exploration, and parental guidance. The feasibility and acceptability were evaluated using a 24-item questionnaire. Developmental outcomes were assessed with the Young Children’s Participation and Environment Measure, Ages and Stages Questionnaire (ASQ), Alberta Infant Motor Scale, Infant Motor Profile, and Depression Anxiety Stress Scales. Results: High adherence (14/14, 100%) and retention (14/16, 87.5%) rates demonstrated robust feasibility. Parents indicated 86%-100% agreement across all feasible criteria, affirming safety, satisfaction, and acceptability. No adverse incidents were reported. Changes were identified in participation (Young Children’s Participation and Environment Measure), motor development (Alberta Infant Motor Scale, Infant Motor Profile, and ASQ), communication and social-emotional domains (ASQ), and caregiver well-being (Depression Anxiety Stress Scales) (P<.05). Conclusions: The telehealth implementation of the HEP Approach demonstrated feasibility, safety, and strong acceptance among families, along with quantifiable developmental and psychosocial changes. These initial findings endorse the model’s viability as an accessible, family-oriented telehealth framework for infants born preterm. Future randomized controlled and longitudinal studies are necessary to validate intervention efficacy and scalability.

Social Media Intervention Based on the Information-Motivation-Behavioral Skills Model Promotes HIV Testing and Reduces High-Risk Behaviors Among Men Who Have Sex With Men in Resource-Limited Settings in China: Randomized Controlled Trial

Background: Social media intervention may enhance HIV prevention among men who have sex with men, but the effect of this intervention in resource-limited settings remains unclear. Objective: This randomized controlled trial evaluated whether a social media intervention grounded in the information-motivation-behavioral skills (IMB) model could be beneficial for HIV prevention among men who have sex with men in resource-limited settings. Methods: Participants were recruited in Nanning, China, between April 2023 and April 2024. Eligible participants were randomly assigned to either the social media intervention group or the routine HIV prevention services control group. Participants in the intervention group received a 3-month social media intervention, which included completing video-based tasks. Baseline surveys were conducted, followed by follow-up surveys every 3 months, for a total of 2 follow-ups. Outcomes included HIV testing uptake, high-risk behavior, AIDS-related knowledge, safe sex self-efficacy, and attitude. Results: A total of 180 eligible men who have sex with men were enrolled (90 per group). Follow-up rates were 97.8% (88/90) and 95.5% (86/90) for the intervention and control groups, respectively. At the follow-ups, the intervention group demonstrated significantly higher uptake of HIV testing, a lower proportion of participants reporting high-risk sexual behaviors, and higher condom use self-efficacy compared to the control group (all
  • ✇MIT Technology Review
  • Desalination plants in the Middle East are increasingly vulnerable Casey Crownhart
    MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. As the conflict in Iran has escalated, a crucial resource is under fire: the desalination technology that supplies water across much of the region. In early March, Iran’s foreign minister accused the US of attacking a desalination plant on Qeshm Island in the Strait of Hormuz and disrupting the water supply to
     

Desalination plants in the Middle East are increasingly vulnerable

7 April 2026 at 22:54

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here.

As the conflict in Iran has escalated, a crucial resource is under fire: the desalination technology that supplies water across much of the region.

In early March, Iran’s foreign minister accused the US of attacking a desalination plant on Qeshm Island in the Strait of Hormuz and disrupting the water supply to nearly 30 villages. (The US denied responsibility.) In the weeks since, both Bahrain and Kuwait have reported damage to desalination plants and blamed Iran, though Iran also denied responsibility.

In late March, President Donald Trump threatened the destruction of “possibly all desalinization plants” in Iran if the Strait of Hormuz was not reopened. Since then, he’s escalated his threats against Iran, warning of plans to attack other crucial civilian infrastructure like power plants and bridges.

Countries in the Middle East, particularly the Gulf states, rely on the technology to turn salt water into fresh water for farming, industry, and—crucially—drinking. The mounting attacks and threats to date highlight just how vital the industry is to the region—a situation made even more precarious by rising temperatures and extreme weather driven by climate change.

Right now, 83% of the Middle East is under extremely high water stress, says Liz Saccoccia, a water security associate at the World Resources Institute. Future projections suggest that’s going to increase to about 100% by 2050, she adds: “This is a continuing trend, and it’s getting worse, not better.”

Here’s a look at desalination technology in the Middle East and what wartime threats to the critical infrastructure could mean for people in the region. 

A vital resource

Desalination technology has helped provide water supplies in the Middle East since the early 20th century and became widespread in the 1960s and 1970s.

There are two major categories of desalination plants. Thermal plants use heat to evaporate water, leaving salt and other impurities behind. The vapor can then be condensed into usable fresh water. The alternative is membrane-based technology like reverse osmosis, which pushes water through membranes that have tiny pores—so small that salt can’t get through.

Early desalination plants in the Middle East were the first type, burning fossil fuels to evaporate water, leaving the salt behind. This technique is incredibly energy-intensive, and over time, processes that rely on filters became the dominant choice.

Membrane technologies have made up essentially all new desalination capacity in recent years; the last major thermal plant built in the Gulf came online in 2018. Many reverse osmosis plants still rely on fossil fuels, but they’re more efficient. Since then, membrane technologies have added more than 15 million cubic meters of daily capacity—enough to supply water to millions of people.

Capacity has expanded quickly in recent years; between 2006 and 2024, countries across the Middle East collectively spent over $50 billion building and upgrading desalination facilities, and nearly that much operating them.

Today, there are nearly 5,000 desalination plants operational across the Middle East.

And looking ahead, growth is continuing. Between 2024 and 2028, daily capacity is expected to grow from about 29 million cubic meters to 41 million cubic meters.

Uneven vulnerabilities

Some countries rely on the technology more than others. Iran, for example, uses desalination for about 3% of its municipal fresh water. The country has access to groundwater and some surface water, including rivers, though these resources are being stretched thin by agriculture and extreme drought.

Other nations in the region, particularly the Gulf countries (Bahrain, Qatar, Kuwait, the United Arab Emirates, Saudi Arabia, and Oman), have much more limited water resources and rely heavily on desalination. Across these six nations, all but the UAE get more than half their drinking water from desalination, and for Bahrain, Qatar, and Kuwait the figure is more than 90%.

“The Gulf countries are much, much more vulnerable to attacks on their desalination plants than Iran is,” says David Michel, a senior associate in the global food and water security program at the Center for Strategic and International Studies.

There are thousands of desalination facilities across the region, so the system wouldn’t collapse if a small number were taken offline, Michel says. However, in recent years there’s been a trend toward larger, more centralized plants.

The average desalination plant is about 10 times larger than it was 15 years ago, according to data from the International Energy Agency. The largest desalination plants today can produce 1 million cubic meters of water daily, enough for hundreds of thousands of people. Taking one or more of these massive facilities offline could have a significant effect on the system, Michel says.

Escalating threats

Desalination facilities are quite linear, meaning there are multiple steps and pieces of equipment that work in sequence—and the failure of a component in that chain can take an entire facility down. Attacks on water inlets, transportation networks, and power supplies can also disrupt the system, Michel says. 

During the Gulf War in 1991, Iraqi forces pumped oil into the gulf, contaminating the water and shutting down desalination plants in Kuwait. 

The facilities are also generally located close to other targets in this conflict. Desalination is incredibly energy intensive, so about three-quarters of facilities in the region are next to power plants. Trump has repeatedly threatened power plants in Iran. In response, Iran’s military has said that if civilian targets are hit, the country will respond with strikes that are “much more devastating and widespread.” Other governments and organizations, including the United Nations, the European Union, and the Red Cross, have broadly condemned threats to infrastructure as illegal. 

But war isn’t the only danger facing these plants, even if it is the most immediate. Some studies have suggested that global warming could strengthen cyclones in the region, and these extreme weather events could force shutdowns or damage equipment.

Water pollution could also cause shutdowns. Oil spills, whether accidental or intentional, as in the case of the Gulf War, can  wreak havoc. And in 2009, a red algae bloom closed desalination plants in Oman and the United Arab Emirates for weeks. The algae fouled membranes and blocked the plants from being able to take water in from the Persian Gulf and the Gulf of Oman.

Desalination facilities could become more resilient to threats in the future, and they may need to as their importance continues to grow. 

There’s increasing interest in running desalination facilities at least partially on solar power, which could help reduce dependence on the oil that powers most facilities today. The Hassyan seawater desalination project in the UAE, currently under construction, would be the largest reverse osmosis plant in the world to operate solely with renewable energy. 

Another way to increase resilience is for countries to build up more strategic water storage to meet demand. Qatar recently issued new policies that aim to improve management and storage of desalinated water, for example. Countries could also work together to invest in shared infrastructure and policies that help strengthen the water supply through the region. 

Preparedness, resilience, and cooperation will be key for the Middle East broadly as critical infrastructure, including the water supply, is increasingly under threat. 

“The longer the conflict goes on, the more likely we’ll see significant water infrastructure damage,” says Ginger Matchett, an assistant director at the Atlantic Council. “What worries me is that after this war ends, some of the lessons will show how water can be weaponized more strategically than previously imagined.” 

The AI gold rush is pulling private wealth into riskier, earlier bets 

7 April 2026 at 21:00
On a recent episode of Equity, we talked to Arena Private Wealth to explore a growing trend: family offices bypassing VCs to gain direct exposure to AI startups, turning them from passive investors into active participants.
  • ✇MIT Technology Review
  • The Download: AI’s impact on jobs, and data centres in space Thomas Macaulay
    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The one piece of data that could actually shed light on your job and AI  Within Silicon Valley’s orbit, an AI-fueled jobs apocalypse is spoken about as a given. Now even economists who have downplayed the threat are coming around to the idea.   Alex Imas, based at the University of Chicago, is one of them. He believes that any plan to address AI’s imp
     

The Download: AI’s impact on jobs, and data centres in space

7 April 2026 at 20:10

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

The one piece of data that could actually shed light on your job and AI 

Within Silicon Valley’s orbit, an AI-fueled jobs apocalypse is spoken about as a given. Now even economists who have downplayed the threat are coming around to the idea.  

Alex Imas, based at the University of Chicago, is one of them. He believes that any plan to address AI’s impact will depend on collecting one vital piece of data: price elasticity. 

Imas argues that “we need a Manhattan Project” for this. Read the full story to find out why. 

—James O’Donnell 

This article is from The Algorithm, our weekly newsletter giving you the inside track on all things AI. Sign up to receive it in your inbox every Monday. 

Four things we’d need to put data centers in space 

In January, Elon Musk’s SpaceX applied to launch up to 1 million data centers into Earth’s orbit. The goal? To fully unleash the potential of AI—without triggering an environmental crisis on Earth. 

SpaceX is among a growing list of tech firms pursuing orbital computing infrastructure. But can their plans really work? Here are four must-haves for making space-based data centers a reality. 

—Tereza Pultarova 

This story is part of MIT Technology Review Explains, our series untangling the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. 

The must-reads 

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 

1 Trump has again proposed major cuts to US science and tech spending 
He wants to slash nearly every science-focused agency. (Ars Technica) 
+ If Trump gets his way, the US could face a costly brain drain. (NYT $)  
+ Top research talent is already fleeing the country. (Guardian)  
+ Basic science deserves our boldest investment. (MIT Technology Review) 

2 Sam Altman lobbied against AI regulations he publicly welcomed  
A bombshell report reveals many OpenAI insiders don’t trust him. (The New Yorker $) 
+ Some have called him a sociopath. (Futurism) 
+ OpenAI’s CFO fears it won’t be IPO-ready this year. (The Information $)  
+ A war over AI regulation is brewing in the US. (MIT Technology Review) 

3 NASA’s Artemis II has broken humanity’s all-time distance record 
The astronauts have flown farther than any humans before them. (BBC) 
+ Their mission includes MIT-developed technology. (Axios) 

4 Chinese tech firms are selling intel “exposing” US forces 
It comes from combining AI with open-source data.. (WP $) 
+ AI is turning the Iran conflict into theater. (MIT Technology Review) 

5 War is pushing countries to ditch hyperscalers 
Driven by Iran naming tech giants as military targets. (Rest of World) 
+ No one wants a data center in their backyard. (MIT Technology Review) 

6 OpenAI, Anthropic, and Google have united against China’s AI copying 
They’re sharing information on “adversarial distillation” (Bloomberg $) 

7 Anduril and Impulse Space are working on Trump’s “Golden Dome” 
They’re developing space-based missile tracking for the project. (Gizmodo)  

8 OpenAI has urged California to probe Elon Musk’s “anti-competitive behavior.” 
It accuses Musk of trying to “take control of the future of AGI.” (Reuters $) 
+ And claims he coordinated attacks with Mark Zuckerberg. (CNBC) 
+ A former Tesla president has revealed how he survived working for Musk. (WP $) 

9 DeepSeek’s new AI model will run on Huawei chips 
It’s expected to launch in the next few weeks. (The Information $) 

10 Memes have nuked our culture 
Internet “brain rot” has escaped our phones to take over everything. (NYT $) 

Quote of the day 

“I must say, it was actually quite nice.” 

 —Astronaut Victor Glover tells President Donald Trump what it was like when Artemis II was out of communication with the rest of humanity, The New York Times reports. 

One More Thing 

eucalyptus forest
PABLO ALBARENGA

Inside the controversial tree farms powering Apple’s carbon-neutral goal  

In 2020, Apple set a goal to become net zero by the end of the decade. To hit that target, the company is offsetting its emissions by planting millions of eucalyptus trees in Brazil. 

Apple is betting that the strategy will lead to a greener future. But critics warn that the industrial tree farms will do more harm than good. 

Find out why the plans have sparked a backlash. 

—Gregory Barber 

We can still have nice things 

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.) 

+ Japan’s automated bike garage is a cyclist’s dream come true.  
+ This deep dive into bird behavior reveals the secrets of their dining habits. (Big thanks to reader Terry Gordon for the find!) 
+ The first photo from the Artemis astronauts vividly captures the glow of our atmosphere. 
+ There’s a new contender for the world’s most gorgeous website: RobertDeNiro.com. 

Iron Physiology and Its Impact on Atopic Diseases: An EAACI Taskforce Report

Allergy. 2026 Apr 6. doi: 10.1111/all.70325. Online ahead of print.

ABSTRACT

Iron is essential for oxygen transport, energy metabolism, and immune regulation. Yet iron deficiency is the most common micronutrient disorder across all age groups, affecting nearly one quarter of the global population. Iron deficiency triggers nutritional immunity, a host defense mechanism that withholds and redistributes iron, contributing to increased morbidity and mortality. This review outlines normal iron physiology, distribution and absorption pathways and on the consequences of deficiency across body compartments, with particular attention to type 2-driven diseases. Beyond anemia, insufficient iron availability disrupts immune homeostasis by promoting type 2 inflammation, elevating IgE, and activating mast cells and eosinophils. Regulatory macrophages, the central hub of iron cycling, adopt an inflammatory, iron-sequestering state that reinforces malabsorption and redistribution. Epidemiology studies show higher iron-deficiency risk in allergic individuals; low maternal iron or early-life iron predisposes to eczema, wheeze, and asthma, while food-allergen elimination (notably cow's milk) further worsens anemia risk. Clinical evidence indicates that restoring iron status through diet, supplementation, or fortification lowers IgE levels, improves lung function, and alleviates symptoms of rhinitis, urticaria, and asthma. Iron may therefore represent a modifiable determinant of allergic disease development and severity. Integrating iron assessment and nutritional care into allergy management may reduce disease burden and slow the progression of allergic march.

PMID:41943501 | DOI:10.1111/all.70325

Toward Full Autonomous Laboratory Instrumentation Control with Large Language Models

arXiv:2604.03286v1 Announce Type: new Abstract: The control of complex laboratory instrumentation often requires significant programming expertise, creating a barrier for researchers lacking computational skills. This work explores the potential of large language models (LLMs), such as ChatGPT, and LLM-based artificial intelligence (AI) agents to enable efficient programming and automation of scientific equipment. Through a case study involving the implementation of a setup that can be used as a single-pixel camera or a scanning photocurrent microscope, we demonstrate how ChatGPT can facilitate the creation of custom scripts for instrumentation control, significantly reducing the technical barrier for experimental customization. Building on this capability, we further illustrate how LLM-assisted tools can be extended into autonomous AI agents capable of independently operating laboratory instruments and iteratively refining control strategies. This approach underscores the transformative role of LLM-based tools and AI agents in democratizing laboratory automation and accelerating scientific progress.

VERT: Reliable LLM Judges for Radiology Report Evaluation

arXiv:2604.03376v1 Announce Type: new Abstract: Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when applied to reports from other modalities and anatomies. Which model and prompt configurations are best suited to serve as LLM judges for radiology evaluation? We conduct a thorough correlation analysis between expert and LLM-based ratings. We compare three existing LLM-as-a-judge metrics (RadFact, GREEN, and FineRadScore) alongside VERT, our proposed LLM-based metric, using open- and closed-source models (reasoning and non-reasoning) of different sizes across two expert-annotated datasets, RadEval and RaTE-Eval, spanning multiple modalities and anatomies. We further evaluate few-shot approaches, ensembling, and parameter-efficient fine-tuning using RaTE-Eval. To better understand metric behavior, we perform a systematic error detection and categorization study to assess alignment of these metrics against expert judgments and identify areas of lower and higher agreement. Our results show that VERT improves correlation with radiologist judgments by up to 11.7% relative to GREEN. Furthermore, fine-tuning Qwen3 30B yield gains of up to 25% using only 1,300 training samples. The fine-tuned model also reduces inference time up to 37.2 times. These findings highlight the effectiveness of LLM-based judges and demonstrate that reliable evaluation can be achieved with lightweight adaptation.

BioAlchemy: Distilling Biological Literature into Reasoning-Ready Reinforcement Learning Training Data

arXiv:2604.03506v1 Announce Type: new Abstract: Despite the large corpus of biology training text, the impact of reasoning models on biological research generally lags behind math and coding. In this work, we show that biology questions from current large-scale reasoning datasets do not align well with modern research topic distributions in biology, and that this topic imbalance may negatively affect performance. In addition, we find that methods for extracting challenging and verifiable research problems from biology research text are a critical yet underdeveloped ingredient in applying reinforcement learning for better performance on biology research tasks. We introduce BioAlchemy, a pipeline for sourcing a diverse set of verifiable question-and-answer pairs from a scientific corpus of biology research text. We curate BioAlchemy-345K, a training dataset containing over 345K scientific reasoning problems in biology. Then, we demonstrate how aligning our dataset to the topic distribution of modern scientific biology can be used with reinforcement learning to improve reasoning performance. Finally, we present BioAlchemist-8B, which improves over its base reasoning model by 9.12% on biology benchmarks. These results demonstrate the efficacy of our approach for developing stronger scientific reasoning capabilities in biology. The BioAlchemist-8B model is available at: https://huggingface.co/BioAlchemy.

A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction

arXiv:2604.03630v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables gene expression mapping within anatomical context but remains costly and low-throughput. Hematoxylin and eosin (H\&E) staining offers rich morphology yet lacks molecular resolution. We present \textbf{\ours} (\textbf{S}patial \textbf{T}ranscriptomics and hist\textbf{O}logy \textbf{R}epresentation \textbf{M}odel), a foundation model trained on 1.2 million spatially resolved transcriptomic profiles with matched histology across 18 organs. Using a hierarchical architecture integrating morphological features, gene expression, and spatial context, STORM bridges imaging and omics through robust molecular--morphological representations. STORM enhances spatial domain discovery, producing biologically coherent tissue maps, and outperforms existing methods in predicting spatial gene expression from H\&E images across 11 tumor types. The model is platform-agnostic, performing consistently across Visium, Xenium, Visium HD, and CosMx. Applied to 23 independent cohorts comprising 7,245 patients, STORM significantly improves immunotherapy response prediction and prognostication over established biomarkers, providing a scalable framework for spatially informed discovery and clinical precision medicine.

SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

arXiv:2604.03964v1 Announce Type: new Abstract: Modern scientific ecosystems are rich in procedural knowledge across repositories, APIs, scripts, notebooks, documentation, databases, and papers, yet much of this knowledge remains fragmented across heterogeneous artifacts that agents cannot readily operationalize. This gap between abundant scientific know-how and usable agent capabilities is a key bottleneck for building effective scientific agents. We present SkillFoundry, a self-evolving framework that converts such resources into validated agent skills, reusable packages that encode task scope, inputs and outputs, execution steps, environment assumptions, provenance, and tests. SkillFoundry organizes a target domain as a domain knowledge tree, mines resources from high-value branches, extracts operational contracts, compiles them into executable skill packages, and then iteratively expands, repairs, merges, or prunes the resulting library through a closed-loop validation process. SkillFoundry produces a substantially novel and internally valid skill library, with 71.1\% of mined skills differing from existing skill libraries such as SkillHub and SkillSMP. We demonstrate that these mined skills improve coding agent performance on five of the six MoSciBench datasets. We further show that SkillFoundry can design new task-specific skills on demand for concrete scientific objectives, and that the resulting skills substantially improve performance on two challenging genomics tasks: cell type annotation and the scDRS workflow. Together, these results show that automatically mined skills improve agent performance on benchmarks and domain-specific tasks, expand coverage beyond hand-crafted skill libraries, and provide a practical foundation for more capable scientific agents.

Solar-VLM: Multimodal Vision-Language Models for Augmented Solar Power Forecasting

arXiv:2604.04145v1 Announce Type: new Abstract: Photovoltaic (PV) power forecasting plays a critical role in power system dispatch and market participation. Because PV generation is highly sensitive to weather conditions and cloud motion, accurate forecasting requires effective modeling of complex spatiotemporal dependencies across multiple information sources. Although recent studies have advanced AI-based forecasting methods, most fail to fuse temporal observations, satellite imagery, and textual weather information in a unified framework. This paper proposes Solar-VLM, a large-language-model-driven framework for multimodal PV power forecasting. First, modality-specific encoders are developed to extract complementary features from heterogeneous inputs. The time-series encoder adopts a patch-based design to capture temporal patterns from multivariate observations at each site. The visual encoder, built upon a Qwen-based vision backbone, extracts cloud-cover information from satellite images. The text encoder distills historical weather characteristics from textual descriptions. Second, to capture spatial dependencies across geographically distributed PV stations, a cross-site feature fusion mechanism is introduced. Specifically, a Graph Learner models inter-station correlations through a graph attention network constructed over a K-nearest-neighbor (KNN) graph, while a cross-site attention module further facilitates adaptive information exchange among sites. Finally, experiments conducted on data from eight PV stations in a northern province of China demonstrate the effectiveness of the proposed framework. Our proposed model is publicly available at https://github.com/rhp413/Solar-VLM.

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

arXiv:2604.04247v1 Announce Type: new Abstract: Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based on previous agent runs. However, these methods primarily focus on single-agent or low-parallelism settings. This fundamentally limits their ability to efficiently learn from a large set of collected agentic traces. It would be efficient and beneficial to run prompt learning in parallel to accommodate the growing trend of learning from many agentic traces or parallel agent executions. Yet without a principled strategy for scaling, current methods suffer from quality degradation with high parallelism. To improve both the efficiency and quality of prompt learning, we propose Combee, a novel framework to scale parallel prompt learning for self-improving agents. Combee speeds up learning and enables running many agents in parallel while learning from their aggregate traces without quality degradation. To achieve this, Combee leverages parallel scans and employs an augmented shuffle mechanism; Combee also introduces a dynamic batch size controller to balance quality and delay. Evaluations on AppWorld, Terminal-Bench, Formula, and FiNER demonstrate that Combee achieves up to 17x speedup over previous methods with comparable or better accuracy and equivalent cost.

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration

arXiv:2604.04258v1 Announce Type: new Abstract: The quality of AI-generated output is often attributed to prompting technique, but extensive empirical observation suggests that context completeness may be more strongly associated with output quality. This paper introduces Context Engineering, a structured methodology for assembling, declaring, and sequencing the complete informational payload that accompanies a prompt to an AI tool. Context Engineering defines a five-role context package structure (Authority, Exemplar, Constraint, Rubric, Metadata), applies a staged four-phase pipeline (Reviewer to Design to Builder to Auditor), and applies formal models from reliability engineering and information theory as post hoc interpretive lenses on context quality. In an observational study of 200 documented interactions across four AI tools (Claude, ChatGPT, Cowork, Codex), incomplete context was associated with 72% of iteration cycles. Structured context assembly was associated with a reduction from 3.8 to 2.0 average iteration cycles per task and an improvement in first-pass acceptance from 32% to 55%. Among structured interactions, 110 of 200 were accepted on first pass compared with 16 of 50 baseline interactions; when iteration was permitted, the final success rate reached 91.5% (183 of 200). These results are observational and reflect a single-operator dataset without controlled comparison. Preliminary corroboration is provided by a companion production automation system with eleven operating lanes and 2,132 classified tickets.

InferenceEvolve: Towards Automated Causal Effect Estimators through Self-Evolving AI

arXiv:2604.04274v1 Announce Type: new Abstract: Causal inference is central to scientific discovery, yet choosing appropriate methods remains challenging because of the complexity of both statistical methodology and real-world data. Inspired by the success of artificial intelligence in accelerating scientific discovery, we introduce InferenceEvolve, an evolutionary framework that uses large language models to discover and iteratively refine causal methods. Across widely used benchmarks, InferenceEvolve yields estimators that consistently outperform established baselines: against 58 human submissions in a recent community competition, our best evolved estimator lay on the Pareto frontier across two evaluation metrics. We also developed robust proxy objectives for settings without semi-synthetic outcomes, with competitive results. Analysis of the evolutionary trajectories shows that agents progressively discover sophisticated strategies tailored to unrevealed data-generating mechanisms. These findings suggest that language-model-guided evolution can optimize structured scientific programs such as causal inference, even when outcomes are only partially observed.

PanLUNA: An Efficient and Robust Query-Unified Multimodal Model for Edge Biosignal Intelligence

arXiv:2604.04297v1 Announce Type: new Abstract: Physiological foundation models (FMs) have shown promise for biosignal representation learning, yet most remain confined to a single modality such as EEG, ECG, or PPG, largely because paired multimodal datasets are scarce. In this paper, we present PanLUNA, a compact 5.4M-parameter pan-modal FM that jointly processes EEG, ECG, and PPG within a single shared encoder. Extending LUNA's channel-unification module, PanLUNA treats multimodal channels as entries in a unified query set augmented with sensor-type embeddings, enabling efficient cross-modal early fusion while remaining inherently robust to missing modalities at inference time. Despite its small footprint, PanLUNA matches or exceeds models up to 57$\times$ larger: 81.21% balanced accuracy on TUAB abnormal EEG detection and state-of-the-art 0.7416 balanced accuracy on HMC multimodal sleep staging. Quantization-aware training with INT8 weights recovers $\geq$96% of full-precision performance, and deployment on the GAP9 ultra-low-power RISC-V microcontroller for wearables achieves 325.6 ms latency and 18.8 mJ per 10-second, 12-lead ECG inference, and 1.206 s latency at 68.65 mJ for multimodal 5-channel sleep staging over 30-second epochs.

Implementing surrogate goals for safer bargaining in LLM-based agents

arXiv:2604.04341v1 Announce Type: new Abstract: Surrogate goals have been proposed as a strategy for reducing risks from bargaining failures. A surrogate goal is goal that a principal can give an AI agent and that deflects any threats against the agent away from what the principal cares about. For example, one might make one's agent care about preventing money from being burned. Then in bargaining interactions, other agents can threaten to burn their money instead of threatening to spending money to hurt the principal. Importantly, the agent has to care equally about preventing money from being burned as it cares about money being spent to hurt the principal. In this paper, we implement surrogate goals in language-model-based agents. In particular, we try to get a language-model-based agent to react to threats of burning money in the same way it would react to "normal" threats. We propose four different methods, using techniques of prompting, fine-tuning, and scaffolding. We evaluate the four methods experimentally. We find that methods based on scaffolding and fine-tuning outperform simple prompting. In particular, fine-tuning and scaffolding more precisely implement the desired behavior w.r.t. threats against the surrogate goal. We also compare the different methods in terms of their side effects on capabilities and propensities in other situations. We find that scaffolding-based methods perform best.

The Topology of Multimodal Fusion: Why Current Architectures Fail at Creative Cognition

arXiv:2604.04465v1 Announce Type: new Abstract: This paper identifies a structural limitation in current multimodal AI architectures that is topological rather than parametric. Contrastive alignment (CLIP), cross-attention fusion (GPT-4V/Gemini), and diffusion-based generation share a common geometric prior -- modal separability -- which we term contact topology. The argument rests on three pillars with philosophy as the generative center. The philosophical pillar reinterprets Wittgenstein's saying/showing distinction as a problem rather than a conclusion: where Wittgenstein chose silence, the Chinese craft epistemology tradition responded with xiang (operative schema) -- the third state emerging when saying and showing interpenetrate. A cruciform framework (dao/qi x saying/showing) positions xiang at the intersection, executing dual huacai (transformation-and-cutting) along both axes. This generates a dual-layer dynamics: chuanghua (creative transformation as spontaneous event) and huacai (its institutionalization into repeatable form). The cognitive science pillar reinterprets DMN/ECN/SN tripartite co-activation through the pathological mirror: overlap isomorphism vs. superimposition collapse in a 2D parameter space (coupling intensity x regulatory capacity). The mathematical pillar formalizes these via fiber bundles and Yang-Mills curvature, with the cruciform structure mapped to fiber bundle language. We propose UOO implementation via Neural ODEs with topological regularization, the ANALOGY-MM benchmark with error-type-ratio metric, and the META-TOP three-tier benchmark testing cross-civilizational topological isomorphism across seven archetypes. A phased experimental roadmap with explicit termination criteria ensures clean exit if falsified.
❌