❌

Reading view

Telehealth Delivery of the Homeostasis–Enrichment–Plasticity Approach for Premature Infants With Developmental Risks: Exploratory Feasibility Study

Background: Preterm delivery is an increasing worldwide health concern linked to increased neurodevelopmental risks. Early intervention is crucial for harnessing neuroplasticity to enhance developmental and functional performance outcomes; however, access to early intervention is frequently hindered by logistical, financial, and labor constraints. The Homeostasis–Enrichment–Plasticity (HEP) Approach is a family-centered early intervention model based on enriched environments, designed to improve infants’ sensory-motor, cognitive, and socio-emotional development. Objective: This study aimed to assess the feasibility, safety, acceptability, and outcomes sensitivity to change of implementing the HEP Approach through telehealth for premature infants at developmental risk. Methods: A pre-post exploratory feasibility study was performed, including 16 preterm infants (aged 4-12 months corrected age), of whom 14 completed the study. The 12-week intervention included weekly remote sessions focused on environmental enrichment, active exploration, and parental guidance. The feasibility and acceptability were evaluated using a 24-item questionnaire. Developmental outcomes were assessed with the Young Children’s Participation and Environment Measure, Ages and Stages Questionnaire (ASQ), Alberta Infant Motor Scale, Infant Motor Profile, and Depression Anxiety Stress Scales. Results: High adherence (14/14, 100%) and retention (14/16, 87.5%) rates demonstrated robust feasibility. Parents indicated 86%-100% agreement across all feasible criteria, affirming safety, satisfaction, and acceptability. No adverse incidents were reported. Changes were identified in participation (Young Children’s Participation and Environment Measure), motor development (Alberta Infant Motor Scale, Infant Motor Profile, and ASQ), communication and social-emotional domains (ASQ), and caregiver well-being (Depression Anxiety Stress Scales) (P<.05). Conclusions: The telehealth implementation of the HEP Approach demonstrated feasibility, safety, and strong acceptance among families, along with quantifiable developmental and psychosocial changes. These initial findings endorse the model’s viability as an accessible, family-oriented telehealth framework for infants born preterm. Future randomized controlled and longitudinal studies are necessary to validate intervention efficacy and scalability.
  •  

Feasibility and Acceptability of AI-Powered Tools for Early Autism Screening in Egypt: Semistructured Focus Group Study

Background: Autism spectrum disorder (ASD) is often underdiagnosed in low- and middle-income countries due to limited specialist access, sociocultural stigma, and fragmented screening systems. Artificial intelligence (AI)–powered screening tools may improve early detection by enabling low-cost, accessible assessments. However, adoption depends on stakeholder trust, ethical safeguards, and alignment with local health system capacities. Objective: This study explored the feasibility, acceptability, and perceived ethical and practical enablers and barriers to implementing AI-powered tools for early ASD screening in Egypt, with attention to urban–rural disparities and integration into existing care pathways. Methods: We used a qualitative design with semistructured focus group discussions with 49 participants (21 parents of children with ASD and 28 health care professionals) recruited from urban and rural governorates. Discussions were audio-recorded, transcribed verbatim, and analyzed using Braun and Clarke’s reflexive thematic analysis, supported by NVivo software (Lumivero). Methodological integrity was ensured through reflexivity, triangulation, and peer debriefing. Thematic saturation was monitored across groups, and participant diversity was prioritized across contexts. Results: Five themes emerged: (1) AI as a supportive tool rather than a replacement for clinicians, emphasizing scalability and assistance for nonspecialists; (2) the need for cultural and contextual adaptation to ensure local relevance; (3) privacy, trust, and transparency concerns, including data security, consent, and algorithmic opacity; (4) reducing diagnostic inequities by addressing urban–rural disparities and strengthening community-based deployment; and (5) the preference for hybrid AI–human models, with conditions for adoption including cultural sensitivity, human oversight, and digital literacy support. Counts (n/N) of parents and health care professionals contributing to each theme were used descriptively as indicators of pattern salience rather than as statistical estimates of prevalence. Participants expressed cautious optimism, with parents emphasizing accessibility and speed, while health care professionals highlighted concerns about reliability, cultural adaptation, and data governance. Conclusions: AI-powered ASD screening has potential to advance equitable early detection in underserved areas. Adoption requires transparent data governance, integration into hybrid human–AI models, culturally adaptive design, and targeted digital literacy initiatives. These findings provide an evidence-based roadmap for policymakers, technologists, and health system leaders to implement AI screening tools that are ethically sound, contextually relevant, and equity-focused.
  •  

A Gamified Mobile Health Intervention to Promote Physical Activity, Executive Function, and Mental Health in College Students: Randomized Controlled Trial

Background: College students commonly experience suboptimal health conditions, including insufficient physical activity (PA), excessive body weight, and declining physical fitness. Traditional interventions face low adherence, while gamified mobile health (mHealth) programs may improve engagement and outcomes. Objective: This study aimed to evaluate the feasibility and effectiveness of a novel gamified, incentive-based mHealth intervention on primary outcomes (PA and adherence) and secondary outcomes (physical fitness, body composition, executive function [EF], and mental health). Methods: A 2-arm parallel-group randomized controlled trial (RCT) was conducted in 2025 at Yantai University with 160 college students (18‐25 years; BMI 18.5‐30.0) who were randomized 1:1 (computer-generated, sex-stratified blocks of 4; concealed allocation) to the intervention group (IG) or control group (CG; n=80 each); major exclusions were contraindications to exercise, severe physical/mental illness, recent PA interventions, or psychotropic medication use. Both used the same fitness watch–app system and identical PA targets (≥150 min moderate-to-vigorous physical activity [MVPA] per week or ≥900 metabolic equivalent-minutes [MET-min] per week); IG additionally received team-based gamification (competition, points/leaderboards, feedback, and rewards), while CG received monitoring only. PA and adherence were monitored throughout the 8-week intervention; other outcomes were assessed at baseline and 8 weeks (fitness, body composition, EF, and mental health). Open-label with blinded outcome assessors/analysts; intention-to-treat (ITT) with multiple imputation. Results: At 8 weeks, data were available for 154 participants (IG 78; CG 76); all 160 were analyzed per ITT. Compared to the CG, the IG demonstrated significantly higher mean levels in all primary PA outcomes over 8 weeks (daily steps: mean 10,356, SD 1245 versus 8242, SD 1087; Δ=2114; =1.81, 95% CI 1.44‐2.18;
  •  

Initial Insights Into an Institutional Secure Large Language Model for Magnetic Resonance Imaging Examination Requests: Retrospective Study

Background: Incomplete clinical details on magnetic resonance imaging (MRI) examination requests (MERs) can lead to suboptimal protocol selection. An institutional secure large language model (sLLM) with access to manually retrieved salient data from the electronic medical record (EMR) may improve request completeness and protocol accuracy across multiple MRI subspecialties. Objective: The objective of this study was to compare clinician MERs with sLLM-augmented MERs for information quality and to evaluate the protocoling accuracy of the sLLM versus board-certified radiologists across body, musculoskeletal, and neuroradiology MRI. Methods: This retrospective study included 608 random outpatient MRI examinations performed between September 2023 and July 2024 (body 206, musculoskeletal 203, neuroradiology 199). The cohort comprised 528 patients (mean 51.2 years, SD 19.2; range 4‐93; n=279, 52.8% women, n=249, 47.2% men). MERs without EMR access were excluded. A privately hosted Anthropic Claude 3.5 model (temperature 0) augmented each MER with manually retrieved salient EMR data and, via rule-based parsing, mapped the extracted elements onto predefined institutional criteria to recommend region or coverage and contrast use. Two experienced radiologists established a consensus reference standard. Two board-certified general radiologists (Rad 3 and Rad 4) and the sLLM were compared with this standard. Clinical information quality was graded using the Reason-for-Exam Imaging Reporting and Data System (RI-RADS). Interrater reliability was quantified with Gwet AC1. Paired accuracies were compared with the McNemar test to determine whether there was a statistically significant difference. Results: Interreader agreement for RI-RADS was almost perfect for sLLM-augmented MERs (AC1 0.97, 95% CI 0.94‐0.99) and moderate for clinician MERs (AC1 0.43, 95% CI 0.34‐0.52). Limited or deficient clinical information (RI-RADS C/D) fell to 0% to 0.7% (0/608 to 4/608) with sLLM augmentation vs 4.1% to 20.4% (25/608 to 124/608) for clinician MERs. Overall protocol accuracy was 93.1% (566/608; 95% CI 89.6‐96.6) for the sLLM, 91.4% (556/608; 95% CI 87.6‐95.3) for Rad 3, and 92.1% (560/608; 95% CI 88.4‐95.8) for Rad 4 (sLLM vs Rad 3 =.23 vs Rad 4 =.40). Region or coverage accuracy was similar (sLLM: 579/608, 95.2%; Rad 3: 585/608, 96.2%; Rad 4: 573/608, 94.2%; =.46 and =.36). Contrast decisions were more accurate using the sLLM at 94.4% (574/608; 95% CI 91.3‐97.5) vs Rad 3 at 92.1% (560/608; 95% CI 88.4‐95.8; =.027) and were not significantly different to Rad 4 at 92.9% (565/608; 95% CI 89.4‐96.4; =.16). Subspecialty analyses showed similar patterns, with the sLLM outperforming Rad 4 for musculoskeletal MRI contrast decisions (96.6% vs 91.1%; =.006) and matching readers elsewhere. Manual review indicated that sLLM improvements arose from EMR details not listed on the MER (infection/inflammation, tumor history, prior surgery). No clinically significant hallucinations were identified in a manual review of discordant cases. Conclusions: Across body, musculoskeletal, and neuroradiology MRI, sLLM-augmented examination requests improved clinical context and enhanced contrast selection while demonstrating accuracy comparable to general radiologists for region or coverage. Integrating sLLMs into routine vetting workflows may reduce manual workload in protocol selection for more efficient, standardized protocoling.
  •  

Social Media Intervention Based on the Information-Motivation-Behavioral Skills Model Promotes HIV Testing and Reduces High-Risk Behaviors Among Men Who Have Sex With Men in Resource-Limited Settings in China: Randomized Controlled Trial

Background: Social media intervention may enhance HIV prevention among men who have sex with men, but the effect of this intervention in resource-limited settings remains unclear. Objective: This randomized controlled trial evaluated whether a social media intervention grounded in the information-motivation-behavioral skills (IMB) model could be beneficial for HIV prevention among men who have sex with men in resource-limited settings. Methods: Participants were recruited in Nanning, China, between April 2023 and April 2024. Eligible participants were randomly assigned to either the social media intervention group or the routine HIV prevention services control group. Participants in the intervention group received a 3-month social media intervention, which included completing video-based tasks. Baseline surveys were conducted, followed by follow-up surveys every 3 months, for a total of 2 follow-ups. Outcomes included HIV testing uptake, high-risk behavior, AIDS-related knowledge, safe sex self-efficacy, and attitude. Results: A total of 180 eligible men who have sex with men were enrolled (90 per group). Follow-up rates were 97.8% (88/90) and 95.5% (86/90) for the intervention and control groups, respectively. At the follow-ups, the intervention group demonstrated significantly higher uptake of HIV testing, a lower proportion of participants reporting high-risk sexual behaviors, and higher condom use self-efficacy compared to the control group (all
  •  

Changes in Workplace Productivity and Estimated Cost Savings During Internet-Based Cognitive Behavioral Therapy in the Irish National Health Service: Naturalistic, Repeated-Measures, Retrospective Survey Study

Background: Depression and anxiety can significantly impact workplace productivity, for instance, by increasing absenteeism and presenteeism. This loss of productivity leads to diminished workplace economic outcomes. Internet-based cognitive behavioral therapy (iCBT) has emerged as a cost-effective intervention within workplace settings that improves workplace productivity loss due to depression and anxiety, but more generalizable evidence beyond the workplace, such as in a national health service setting, is lacking. Objective: This naturalistic, repeated-measures, retrospective study investigated the impact of iCBT on work productivity metrics using nationally representative data from patients enrolled in the Irish national health service (ie, the Health Service Executive). Methods: We analyzed repeated measures retrospective data from 7125 employed patients enrolled in iCBT at the Health Service Executive between March 2023 and May 2024. The Work Productivity and Activity Impairment questionnaire was used to measure absenteeism, presenteeism, overall productivity loss, and activity impairment. Secondary outcomes included depression (Patient Health Questionnaire-9) and anxiety (Generalized Anxiety Disorder-7). Patients were primarily 25 to 64 years old (n=5578, 78%), female (n=4956, 70%), and met clinical scoring criteria on the Patient Health Questionnaire-9 or Generalized Anxiety Disorder-7 (n=4774, 67%). Missing data were handled using multiple imputation. We used mixed-effects models to assess pre-post treatment changes in outcomes and then utilized Irish national salary estimates from 2022 to derive cost savings (in 2022 € values; €1=approximately US $1.05) based on productivity improvement during use of the iCBT program. Results: From baseline to follow-up, absenteeism reduced by 6.85% (
  •  

If you cap insulin at $35 a month, people with type 2 diabetes stick to treatment, study finds

Get your daily dose of health and medicine every weekday with STAT’s free newsletter Morning Rounds. Sign up here.

Good morning. Amid all the tasks that demand to be addressed each day, are you having trouble finding your sense of wonder and awe re: the moon mission? I am too. Reading this and this yesterday helped me access it. 

Read the rest…

© Adobe

  •  

Iron Physiology and Its Impact on Atopic Diseases: An EAACI Taskforce Report

Allergy. 2026 Apr 6. doi: 10.1111/all.70325. Online ahead of print.

ABSTRACT

Iron is essential for oxygen transport, energy metabolism, and immune regulation. Yet iron deficiency is the most common micronutrient disorder across all age groups, affecting nearly one quarter of the global population. Iron deficiency triggers nutritional immunity, a host defense mechanism that withholds and redistributes iron, contributing to increased morbidity and mortality. This review outlines normal iron physiology, distribution and absorption pathways and on the consequences of deficiency across body compartments, with particular attention to type 2-driven diseases. Beyond anemia, insufficient iron availability disrupts immune homeostasis by promoting type 2 inflammation, elevating IgE, and activating mast cells and eosinophils. Regulatory macrophages, the central hub of iron cycling, adopt an inflammatory, iron-sequestering state that reinforces malabsorption and redistribution. Epidemiology studies show higher iron-deficiency risk in allergic individuals; low maternal iron or early-life iron predisposes to eczema, wheeze, and asthma, while food-allergen elimination (notably cow's milk) further worsens anemia risk. Clinical evidence indicates that restoring iron status through diet, supplementation, or fortification lowers IgE levels, improves lung function, and alleviates symptoms of rhinitis, urticaria, and asthma. Iron may therefore represent a modifiable determinant of allergic disease development and severity. Integrating iron assessment and nutritional care into allergy management may reduce disease burden and slow the progression of allergic march.

PMID:41943501 | DOI:10.1111/all.70325

  •  

Article: Bloom Filters: Theory, Engineering Trade‑offs, and Implementation in Go

This article walks you through the Go implementation of Bloom filters to optimize the performance of a recommender. It cover the architectural view, Bloom filter mechanics, Go integration, parameter tuning, and practical lessons learned from making it work under production constraints.

By Gabor Koos
  •  

A star scientist showed that better genetics lessons could reduce racism. It was the death knell for his career

Every year, the Genetics Society of America bestows the Elizabeth W. Jones Award for Excellence in Education, recognizing someone who has helped the public better understand the science of DNA. It’s understood to be a lifetime achievement award; past recipients tend toward retirement age with decades of work behind them and stacks of textbooks to their names. 

When this year’s winner, Brian Donovan, was announced at the end of February, many geneticists and science educators found it hard to celebrate the news. Not because he’s undeserving of the honor. Far from it. But because it seemed to confirm what many feared: that Donovan’s incandescent research career was over before it had barely begun. 

Read the rest…

© Jerry McBride for STAT

  •  

Towards the AI Historian: Agentic Information Extraction from Primary Sources

arXiv:2604.03553v1 Announce Type: new Abstract: AI is supporting, accelerating, and automating scientific discovery across a diverse set of fields. However, AI adoption in historical research remains limited due to the lack of solutions designed for historians. In this technical progress report, we introduce the first module of Chronos, an AI Historian under development. This module enables historians to convert image scans of primary sources into data through natural-language interactions. Rather than imposing a fixed extraction pipeline powered by a vision-language model (VLM), it allows historians to adapt workflows for heterogeneous source corpora, evaluate the performance of AI models on specific tasks, and iteratively refine workflows through natural-language interaction with the Chronos agent. The module is open-source and ready to be used by historical researchers on their own sources.
  •  

TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables

arXiv:2604.03660v1 Announce Type: new Abstract: Structured tables are essential for conveying high-density information in professional domains such as finance, healthcare, and scientific research. Despite the progress in Multimodal Large Language Models (MLLMs), reasoning performance remains limited for complex tables with hierarchical layouts. In this paper, we identify a critical Perception Bottleneck through quantitative analysis. We find that as task complexity scales, the number of involved discrete visual regions increases disproportionately. This processing density leads to an internal "Perceptual Overload," where MLLMs struggle to maintain accurate spatial attention during implicit generation. To address this bottleneck, we introduce TableVision, a large-scale, trajectory-aware benchmark designed for spatially grounded reasoning. TableVision stratifies tabular tasks into three cognitive levels (Perception, Reasoning, and Analysis) across 13 sub-categories. By utilizing a rendering-based deterministic grounding pipeline, the dataset explicitly couples multi-step logical deductions with pixel-perfect spatial ground truths, comprising 6,799 high-fidelity reasoning trajectories. Our empirical results, supported by diagnostic probing, demonstrate that explicit spatial constraints significantly recover the reasoning potential of MLLMs. Furthermore, our two-stage decoupled framework achieves a robust 12.3% overall accuracy improvement on the test set. TableVision provides a rigorous testbed and a fresh perspective on the synergy between perception and logic in document understanding.
  •  

PRAISE: Prefix-Based Rollout Reuse in Agentic Search Training

arXiv:2604.03675v1 Announce Type: new Abstract: In agentic search, large language models (LLMs) are trained to perform multi-turn retrieval and reasoning for complex tasks such as multi-hop question answering (QA). However, current search-based Reinforcement Learning (RL) methods suffer from two core limitations: expensive long-horizon rollouts are under-utilized during training, and supervision is typically available only at the final answer, resulting in severe reward sparsity. We present Prefix-based Rollout reuse for Agentic search with Intermediate Step rEwards (PRAISE), a framework for improving both data efficiency and credit assignment in agentic search training. Given a complete search trajectory, PRAISE extracts prefix states at different search turns, elicits intermediate answers from them, and uses these prefixes both to construct additional training trajectories and to derive step-level rewards from performance differences across prefixes. Our method uses a single shared model for both search policy learning and prefix answer evaluation, enabling joint optimization without extra human annotations or a separate reward model. Experiments on multi-hop QA benchmarks show that PRAISE consistently improves performance over strong baselines.
  •  

Decomposing Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

arXiv:2604.03785v1 Announce Type: new Abstract: Communication is essential for coordination in \emph{cooperative} multi-agent reinforcement learning under partial observability, yet \emph{cross-timestep} delays cause messages to arrive multiple timesteps after generation, inducing temporal misalignment and making information stale when consumed. We formalize this setting as a delayed-communication partially observable Markov game (DeComm-POMG) and decompose a message's effect into \emph{communication gain} and \emph{delay cost}, yielding the Communication Gain and Delay Cost (CGDC) metric. We further establish a value-loss bound showing that the degradation induced by delayed messages is upper-bounded by a discounted accumulation of an information gap between the action distributions induced by timely versus delayed messages. Guided by CGDC, we propose \textbf{CDCMA}, an actor--critic framework that requests messages only when predicted CGDC is positive, predicts future observations to reduce misalignment at consumption, and fuses delayed messages via CGDC-guided attention. Experiments on no-teammate-vision variants of Cooperative Navigation and Predator Prey, and on SMAC maps across multiple delay levels show consistent improvements in performance, robustness, and generalization, with ablations validating each component.
  •  

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

arXiv:2604.03893v1 Announce Type: new Abstract: Breakthroughs in frontier theory often depend on the combination of concrete diagrammatic notations with rigorous logic. While multimodal large language models (MLLMs) show promise in general scientific tasks, current benchmarks often focus on local information extraction rather than the global structural logic inherent in formal scientific notations. In this work, we introduce FeynmanBench, the first benchmark centered on Feynman diagram tasks. It is designed to evaluate AI's capacity for multistep diagrammatic reasoning, which requires satisfying conservation laws and symmetry constraints, identifying graph topology, converting between diagrammatic and algebraic representations, and constructing scattering amplitudes under specific conventions and gauges. To support large-scale and reproducible evaluation, we developed an automated pipeline producing diverse Feynman diagrams along with verifiable topological annotations and amplitude results. Our database spans the electromagnetic, weak, and strong interactions of the Standard Model, encompasses over 100 distinct types and includes more than 2000 tasks. Experiments on state-of-the-art MLLMs reveal systematic failure modes, including unstable enforcement of physical constraints and violations of global topological conditions, highlighting the need for physics-grounded benchmarks for visual reasoning over scientific notation. FeynmanBench provides a logically rigorous test of whether AI can effectively engage in scientific discovery, particularly within theoretical physics.
  •  

CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation

arXiv:2604.03926v1 Announce Type: new Abstract: We present CODE-GEN, a human-in-the-Loop, retrieval-augmented generation (RAG)-based agentic AI system for generating context-aligned multiple-choice questions to develop student code reasoning and comprehension abilities. CODE-GEN employs an agentic AI architecture in which a Generator agent produces multiple-choice coding comprehension questions aligned with course-specific learning objectives, while a Validator agent independently assesses content quality across seven pedagogical dimensions. Both agents are augmented with specialized tools that enhance computational accuracy and verify code outputs. To evaluate the effectiveness of CODE-GEN, we conducted an evaluation study involving six human subject-matter experts (SMEs) who judged 288 AI-generated questions. The SMEs produced a total of 2,016 human-AI rating pairs, indicating agreement or disagreement with the assessments of Validator, along with 131 instances of qualitative feedback. Analyses of SME judgments show strong system performance, with human-validated success rates ranging from 79.9% to 98.6% across the seven pedagogical dimensions. The analysis of qualitative feedback reveals that CODE-GEN achieves high reliability on dimensions well suited to computational verification and explicit criteria matching, including question clarity, code validity, concept alignment, and correct answer validity. In contrast, human expertise remains essential for dimensions requiring deeper instructional judgment, such as designing pedagogically meaningful distractors and providing high-quality feedback that reinforces understanding. These findings inform the strategic allocation of human and AI effort in AI-assisted educational content generation.
  •  

Beyond Fluency: Toward Reliable Trajectories in Agentic IR

arXiv:2604.04269v1 Announce Type: new Abstract: Information Retrieval is shifting from passive document ranking toward autonomous agentic workflows that operate in multi-step Reason-Act-Observe loops. In such long-horizon trajectories, minor early errors can cascade, leading to functional misalignment between internal reasoning and external tool execution despite continued linguistic fluency. This position paper synthesizes failure modes observed in industrial agentic systems, categorizing errors across planning, retrieval, reasoning, and execution. We argue that safe deployment requires moving beyond endpoint accuracy toward trajectory integrity and causal attribution. To address compounding error and deceptive fluency, we propose verification gates at each interaction unit and advocate systematic abstention under calibrated uncertainty. Reliable Agentic IR systems must prioritize process correctness and grounded execution over plausible but unverified completion.
  •  

RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets

arXiv:2604.04347v1 Announce Type: new Abstract: 2026 has brought an explosion of interest in LLM-guided evolution of agentic artifacts, with systems like GEPA and Autoresearch demonstrating that LLMs can iteratively improve prompts, code, and agent architectures across diverse domains. As adoption accelerates, a central question emerges: given the same information, the same seed agent, and the same objective, which optimization algorithm yields the best results under the same evaluation budget? This question becomes critical when evaluations are expensive, such as when they require human judgment or multiple LLM calls. We present the first systematic comparison of three optimization paradigms -- Elo tournament selection (RoboPhD), Pareto-based selection (GEPA), and greedy hill-climbing (Autoresearch) -- across four benchmarks spanning abstract reasoning, cloud scheduling, SQL generation, and financial QA, all under a fixed budget of 1,500 evaluations. RoboPhD introduces validation-free evolution: instead of splitting the budget between training and validation, it uses Elo competition on training data to simultaneously evaluate agents and drive evolution. All three systems receive seed agents with diagnostic print() statements that evolution can grow, enabling self-instrumenting agents that develop increasingly informative diagnostics for the benefit of their evolutionary successors. Using a single default configuration, RoboPhD outperforms both GEPA and Autoresearch on three of four benchmarks, losing only on the simplest task, where the winning solution (from our Autoresearch adaptation) required under 90 lines of code. On ARC-AGI, RoboPhD evolves a 22-line seed agent into a 1,013-line multi-strategy system, improving accuracy from 27.8% to 65.8% using Gemini 3.1 Flash Lite as the solver. We release RoboPhD as a versatile toolkit under the MIT license with a simple optimize_anything() API for evolving diverse complex agents.
  •  

Decocted Experience Improves Test-Time Inference in LLM Agents

arXiv:2604.04373v1 Announce Type: new Abstract: There is growing interest in improving LLMs without updating model parameters. One well-established direction is test-time scaling, where increased inference-time computation (e.g., longer reasoning, sampling, or search) is used to improve performance. However, for complex reasoning and agentic tasks, naively scaling test-time compute can substantially increase cost and still lead to wasted budget on suboptimal exploration. In this paper, we explore \emph{context} as a complementary scaling axis for improving LLM performance, and systematically study how to construct better inputs that guide reasoning through \emph{experience}. We show that effective context construction critically depends on \emph{decocted experience}. We present a detailed analysis of experience-augmented agents, studying how to derive context from experience, how performance scales with accumulated experience, what characterizes good context, and which data structures best support context construction. We identify \emph{decocted experience} as a key mechanism for effective context construction: extracting essence from experience, organizing it coherently, and retrieving salient information to build effective context. We validate our findings across reasoning and agentic tasks, including math reasoning, web browsing, and software engineering.
  •  

PSY-STEP: Structuring Therapeutic Targets and Action Sequences for Proactive Counseling Dialogue Systems

arXiv:2604.04448v1 Announce Type: new Abstract: Cognitive Behavioral Therapy (CBT) aims to identify and restructure automatic negative thoughts pertaining to involuntary interpretations of events, yet existing counseling agents struggle to identify and address them in dialogue settings. To bridge this gap, we introduce STEP, a dataset that models CBT counseling by explicitly reflecting automatic thoughts alongside dynamic, action-level counseling sequences. Using this dataset, we train STEPPER, a counseling agent that proactively elicits automatic thoughts and executes cognitively grounded interventions. To further enhance both decision accuracy and empathic responsiveness, we refine STEPPER through preference learning based on simulated, synthesized counseling sessions. Extensive CBT-aligned evaluations show that STEPPER delivers more clinically grounded, coherent, and personalized counseling compared to other strong baseline models, and achieves higher counselor competence without inducing emotional disruption.
  •  
❌