❌

Normal view

Artificial Intelligence-Driven Multiomics and Clinical Investigation Identify Macrophage Migration Inhibitory Factor as a Pan-Cancer Biomarker

Phenomics. 2026 May 20;6(3):213-229. doi: 10.1007/s43657-026-00322-4. eCollection 2026 Jun.

ABSTRACT

Early cancer detection remains challenging due to the lack of reliable pan-cancer screening methods, particularly blood-based biomarkers. Using a novel three-tiered validation framework combining artificial intelligence (AI)-powered literature mining of 180,000 PubMed articles (1950-2024), multiomics integration across major databases, and extensive clinical validation, we identified macrophage migration inhibitory factor (MIF) as a promising blood-based biomarker for pan-cancer detection. Multiomics analysis revealed consistent MIF upregulation across 21 cancer types at the transcriptional level and across 12 cancer types at the protein level. Clinical validation in independent cohorts (n = 4,269) showed that serum MIF protein levels discriminated effectively between cancer patients and healthy controls (median AUC = 0.994) and between cancer and benign conditions (median AUC = 0.881). Notably, comparative analyses showed that MIF demonstrated superior or comparable performance to established cancer-specific markers, including AFP for hepatocellular carcinoma (MIF AUC = 0.885 vs. AFP AUC: 0.744-0.887) and CA125 for ovarian cancer (MIF AUC = 0.831 vs. CA125 AUC: 0.58-0.71). Meta-analysis of 28 cohorts (n = 5,347) confirmed the diagnostic efficacy of MIF (pooled AUC: 0.782). This cost-effective, blood-based ELISA approach establishes MIF as a valuable tool for broad applications in cancer screening.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s43657-026-00322-4.

PMID:42750739 | PMC:PMC13578188 | DOI:10.1007/s43657-026-00322-4

Integrated multi-omic profiling enables recurrence risk stratification beyond pathological stage in resected EGFR-mutant lung adenocarcinoma

J Thorac Oncol. 2026 Sep 16:104204. doi: 10.1016/j.jtho.2026.104204. Online ahead of print.

ABSTRACT

BACKGROUND: Early-stage EGFR-mutant lung adenocarcinoma (LUAD) demonstrates heterogeneous outcomes after curative surgery, yet adjuvant treatment decisions are guided by pathological stage alone. Following the ADAURA trial, adjuvant osimertinib is the standard of care for resected stage IB-IIIA EGFR-mutant LUAD; however, real-world data demonstrate that up to 40% of patients remain disease-free at five years without adjuvant osimertinib, underscoring the need for improved risk stratification.

PATIENTS AND METHODS: We performed integrated clinical, genomic and transcriptomic profiling of 400 patients with resected stage IA-IIIA EGFR-mutant LUAD. EGFR-mutant recurrence risk models integrating clinical, genomic and transcriptomic data were developed and validated across one internal and three external cohorts.

RESULTS: Genomic instability, including TP53 co-mutations, copy number alterations and APOBEC-associated mutational signatures, increased with pathological stage. RBM10 co-mutations were enriched in tumours with L858R mutations and correlated with upregulation of WNT signalling and epithelial-mesenchymal transition. Transcriptomic features outperformed clinical or genomic variables alone in predicting recurrence risk, and a multi-omic model demonstrated superior and reproducible performance, achieving a median concordance index of 75.4% across four independent validation cohorts. The multi-omic model stratified recurrence risk within individual pathological stages, including stage I disease, and identified patients most likely to benefit from adjuvant EGFR TKI.

CONCLUSIONS: These findings define the molecular heterogeneity of early-stage EGFR-mutant LUAD and support multi-omic risk stratification to inform adjuvant EGFR TKI decisions beyond pathological stage. Prospective validation in larger cohorts will be required to confirm these findings.

PMID:42749051 | DOI:10.1016/j.jtho.2026.104204

Advances in single-cell and spatial multi-omics for deciphering the mechanisms of pan-organ metastasis in breast cancer

Biochim Biophys Acta Rev Cancer. 2026 Sep 15:189717. doi: 10.1016/j.bbcan.2026.189717. Online ahead of print.

ABSTRACT

Breast cancer deaths are mainly caused by metastasis to distant organs, not by the primary tumor. Bone, lung, liver, and brain are the most common metastatic sites, each showing different clinical behaviors and treatment responses-a pattern often called metastatic organotropism. Bulk omics can provide tissue-level information, but they fall short in identifying rare metastasis-initiating clones or capturing how tumor cells adapt to distinct organ microenvironments. With recent progress in single-cell sequencing, multi-omics integration, and spatial profiling, it is now possible to study metastasis at much finer cellular and spatial resolution. In this review, we synthesize current evidence from two complementary perspectives. First, we summarize pan-organ programs associated with metastatic competence, including partial epithelial-mesenchymal transition, lineage plasticity, stem-like states, stress tolerance, metabolic flexibility, immune evasion, and stromal-vascular remodeling. Second, we discuss how these programs are reshaped by organ-specific microenvironments: osteolytic and mixed bone remodeling and marrow dormancy in bone, inflammatory vascular niches in lung, tolerogenic antigen presentation and hepatic metabolism in liver, and blood-brain/blood-tumor barrier constraints, glial crosstalk, neuronal interactions, and lipid-metabolic adaptation in brain. We also highlight how CTC/CTM profiling, spatial mapping, and longitudinal integration refine the understanding of dissemination, dormancy, colonization, outgrowth, and treatment resistance. Although these approaches hold translational promise, most remain at the discovery or early validation stage and require assay simplification, prospective testing, and cross-center standardization. Overall, single-cell and spatial multi-omics are reframing breast cancer metastasis as a dynamic, multi-stage, and tissue-shaped process, providing a foundation for future biomarker development and mechanism-guided therapeutic strategies.

PMID:42744123 | DOI:10.1016/j.bbcan.2026.189717

Mechanism of Action of Hedyotis diffusa Extract in a Rat Model of Acute Lung Injury Based on Transcriptomic Analysis

Biology (Basel). 2026 Sep 4;15(17):1549. doi: 10.3390/biology15171549.

ABSTRACT

OBJECTIVE: This study established a rat model of lipopolysaccharide (LPS)-induced acute lung injury (ALI) to evaluate pathological damage, collagen deposition, inflammatory cytokine levels, and key gene/protein expression following Hedyotis diffusa water extract (HDWE) intervention. Combined with ultra-high-performance liquid chromatography-quadrupole Orbitrap high-resolution mass spectrometry (UHPLC-Q-Orbitrap HRMS), transcriptomic analysis, and molecular simulation, this study identified the bioactive components of HDWE, evaluated their potential interactions with ALI-related targets, and explored the multi-omics-based protective mechanisms of HDWE.

METHODS: Thirty-six Sprague-Dawley (SD) rats were randomly divided into six groups: Control group, ALI group, DXMS group, HDWE-L group (100 mg/kg), HDWE-M group (200 mg/kg), and HDWE-H group (300 mg/kg). Hematoxylin and eosin (H&E) and Masson's trichrome staining were used to evaluate lung pathological changes and collagen deposition. Enzyme-linked immunosorbent assay (ELISA) was used to measure serum tumor necrosis factor-α TNF-α interleukin-1β IL-1β, erleukin-6 (IL-6), and interleukin-10 (IL-10) levels. Transcriptomic analysis identified differentially expressed genes (DEGs), followed by Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), receiver operating characteristic (ROC), and immune infiltration analyses. Quantitative real-time polymerase chain reaction (qRT-PCR) detected the mRNA expression levels of SPHK1, RELA, and NFKBIA. Immunohistochemistry evaluated the expression of eight hub targets, including endothelin-1 (EDN1), sphingosine kinase 1 (SPHK1), intercellular adhesion molecule 1 (ICAM1), interleukin-17 (IL-17), prostaglandin-endoperoxide synthase 2 (PTGS2/COX-2), NF-κB p65 (encoded by RELA), WT1-associated protein (WTAP), and myeloperoxidase (MPO). UHPLC-Q-Orbitrap HRMS characterized HDWE constituents. Molecular docking analysis was performed between 22 compounds and eight hub targets, followed by 100 ns molecular dynamics simulations and molecular mechanics-Poisson-Boltzmann surface area (MM/PBSA) binding free energy calculations for five core targets. Compared with the control group, the ALI group showed increased levels of TNF-α (86%), IL-1β (107%), and IL-6 (66%), accompanied by a 43% reduction in IL-10 and a 300% increase in lung collagen deposition. All HDWE doses alleviated inflammatory responses, with medium-dose HDWE showing the most pronounced effects. Specifically, medium-dose HDWE increased IL-10 levels by 52% and reduced IL-6, TNF-α, and IL-1β levels by 18%, 22%, and 11%, respectively. Transcriptomic analysis identified 2512 DEGs between the control group and ALI groups, 832 exclusive DEGs between the ALI group and HDWE-M groups, and 876 overlapping DEGs enriched in TNF, IL-17, and NF-κB signaling pathways. The eight-hub-gene diagnostic model achieved an area under the curve (AUC) of 0.969. RELA, SPHK1, and four other hub genes showed positive correlations with Th1, Th17, and neutrophil infiltration. In the ALI group, SPHK1, RELA, and NFKBIA mRNA expression levels were 1.30-, 0.96-, and 0.71-fold of those in the control group, respectively. Compared with the ALI group, high-dose HDWE treatment and low-dose HDWE treatment reduced SPHK1 expression to 0.62- and 0.57-fold, respectively, and increased NFKBIA expression to 1.68- and 1.58-fold, respectively. High-dose HDWE treatment reduced RELA expression to 0.43-fold. The expression levels of inflammation-related proteins were increased in the ALI group and were reduced after HDWE treatment. Twenty-two HDWE components were identified, 16 of which met the docking criteria. Asperulosidic acid exhibited favorable predicted binding affinities with all eight targets, with calculated binding free energies of -14.74, -14.92, -17.58, -23.04, and -16.10 kcal/mol for MPO, IL-17, NF-κB p65, PTGS2/COX-2, and SPHK1, respectively.

CONCLUSIONS: This study provides systematic in vivo pharmacodynamic and in silico component-target evidence regarding the protective effects of HDWE against LPS-induced ALI. HDWE treatment increased NFKBIA expression and reduced SPHK1, RELA, and multiple inflammatory protein levels, suggesting that HDWE may regulate the IL-17/NF-κB-associated inflammatory network, although direct causal relationships require further validation. Asperulosidic acid may represent a key bioactive component with broad target-binding potential. This study was limited by the use of an LPS-induced rat ALI model without gene knockout or target inhibitor validation; therefore, further functional experiments are required to confirm the proposed regulatory mechanisms.

PMID:42737981 | PMC:PMC13564518 | DOI:10.3390/biology15171549

IFN-I induced-LAP3 promotes embryo resorption by inhibiting trophoblast mitophagy via targeting HSD17B10/PE pathway

Cell Death Discovery, Published online: 15 September 2026; doi:10.1038/s41420-026-03342-1

IFN-I induced-LAP3 promotes embryo resorption by inhibiting trophoblast mitophagy via targeting HSD17B10/PE pathway

Synthetic transcription factors designed by domain recombination enhance CAR T cell antitumor function

Recombining domains across an entire protein family, rather than relying on natural sequences shaped by evolution, generates synthetic “DESynR” transcription factors with enhanced function. DESynR AP-1 TFs reprogram CAR T cells into non-natural, therapeutically optimized states and outperform natural AP-1 factors in antitumor immunity.

Advancing cancer detection and treatment using longitudinal routine clinical data

Liu et al. develop Oncoformer, a multimodal transformer that reads routine laboratory tests and chest X-rays already collected in everyday care. Across more than 3.6 million individuals, it detects cancer, infers tumor stage, and stratifies treatment response and recurrence risk, pointing toward risk-adapted cancer care built on data already in hand.

Spatial proximity sequencing maps developmental dynamics in the germinal center

Sprox-seq enables spatial profiling of protein complexes, surface proteins, and mRNAs in intact tissues by combining proximity ligation with spatial transcriptomics. In human tonsils, Sprox-seq maps germinal center interaction networks, links CD21-CD35 complexes to proliferative programs, reveals interaction-based B cell state transitions, and directly captures B cell-follicular dendritic cell communication.

Digitally Adapting LGBTQ-Affirmative Cognitive Behavioral Therapy for Chinese Men Who Have Sex With Men Living With HIV: User-Centered Design Approach

Background: Chinese men who have sex with men living with HIV (MSMLWH) experience substantial psychological distress driven by minority stress and HIV-related challenges. However, culturally tailored digital mental health interventions that address HIV-specific maladaptive cognitive schemas and culturally specific psychosocial stressors remain scarce in China. Objective: This study aimed to systematically adapt an evidence-based cognitive behavioral therapy (CBT) intervention Effective Skills to Empower Effective Men (ESTEEM) into a WeChat (Tencent) Mini-Program–based intervention (iESTEEM) specifically for Chinese MSMLWH and to evaluate its preliminary feasibility and usability. Methods: We used a three-phase user-centered design approach guided by the Assessment, Decision, Adaptation, Production, Topical Experts, Integration, Training, and Testing (ADAPT-ITT) framework. The study proceeded in three phases: (1) a qualitative needs assessment using semistructured interviews with 20 MSMLWH (mean age 23.25, SD 3.08 years); (2) systematic intervention adaptation and platform development, including theater testing (n=5); and (3) a 2-week pilot study involving 10 MSMLWH and five counselors to evaluate feasibility, usability, and acceptability through focus groups and objective platform analytics. Results: Phase 1 identified 3 major themes of psychological distress: persistent health anxiety fueled by catastrophizing, intersectional stigma internalization, the disclosure dilemma, and intimacy barriers rooted in defectiveness and shame schemas. Participants also prioritized anonymity and bite-sized learning. Guided by these findings, iESTEEM was developed as a counselor-assisted, privacy-preserving WeChat Mini-Program incorporating HIV-specific scenarios, multimodal learning modules, and a back-end risk-alert system. During the 2-week pilot, participants logged into the platform 14.1 (SD 6.7) times per person and completed 134.3 (SD 103.1) minutes of learning activities; all participants accessed module 1, and 90% (9/10) accessed modules 2‐5. Anxiety scores decreased from 8.9 (SD 2.3) to 7.2 (SD 3.0), whereas depression scores remained stable. All participants expressed a willingness to continue using the program and to recommend it to peers. Participants and counselors endorsed its contextual relevance, privacy protections, and clinical utility. Conclusions: This study provides a theory- and evidence-informed model for culturally adapting digital mental health interventions for highly stigmatized populations. By integrating lesbian, gay, bisexual, transgender, and queer (LGBTQ)-affirmative CBT principles, HIV-specific adaptations, and a privacy-preserving, counselor-assisted WeChat Mini-Program, iESTEEM demonstrated promising preliminary feasibility, acceptability, and engagement among Chinese MSMLWH. These findings support the potential of culturally tailored digital interventions to expand access to psychological support for this stigmatized population in resource-constrained settings. Ongoing randomized controlled trials will further evaluate its efficacy, implementation outcomes, and mechanism of action. Trial Registration: Chinese Clinical Trial Registry ChiCTR2400080263; https://www.chictr.org.cn/showproj.html?proj=216926

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.

When Successful Knowledge Graph Edits Displace Correct Answers: Rank-Level Locality beyond Parameter Support

arXiv:2609.12116v1 Announce Type: new Abstract: Editing a knowledge graph embedding (KGE) model to promote a desired answer can displace correct answers from the returned list. Locality tests based only on facts that reuse the edited parameter can miss this ranking effect. We introduce a common rank-displacement audit at three scopes: facts supported by the edited parameter, other correct answers to the target query, and correct answers across queries with the same relation. We also derive dimensional and geometric conditions for an update to improve the target while exactly preserving selected scores. On FB15k-237 with DistMult and ComplEx, direct promotion always moves the target into the top ten, but does so without damage in only 23.0--23.2\% of edits. Strict preservation causes no measured damage, yet succeeds in only 1.3--1.4\%. Support-regularized entity editing gives the highest joint success, 36.3--37.7\%, while rank-truncated preservation reaches 32.8--34.7\% and reduces the mean number of displaced answers from about 14 to 1.2. Experiments across dimensions, scorers, ranking conventions, and a learned editor show that locality depends on both the protected scope and the editing mechanism. KGE editing should therefore report correction success together with the incidence and severity of rank displacement.

GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs

arXiv:2609.12265v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly asked to reason over structured data such as graphs, yet how reliably they can carry out multi-step graph algorithms in language remains unclear. Existing evaluations tend to use simple tasks on small graphs, to score code generation rather than reasoning over the graph itself, or to fix a single input format. We introduce Graph Theory Bench (GT Bench), a benchmark covering 24 classical graph problems in 44 task-structure settings, with over 100,000 examples across four representations: natural language, structured language, adjacency list, and adjacency matrix. Evaluating eight LLMs on GT Bench shows that accuracy is strongly tied to the input representation, that the best representation shifts with graph density, size, and topology as well as with the model, and that this sensitivity persists, attenuated, in the strongest reasoning models. Building on these observations, we propose the Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM. GTA lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split, outperforming eight prompting and agent baselines, and transfers without retraining to GraCoRe and NLGraph. Code for benchmark generation and evaluation: https://github.com/xzx34/GTA. The project homepage is available at https://xzx34.github.io/gta/.

Hybrid Physics-AI Framework of Body Center of Mass Dynamics from Wrist-Worn Sensors

arXiv:2609.12304v1 Announce Type: new Abstract: Wrist-worn IMU has been widely used for daily-life health monitoring. Yet, it does not fully represent whole-body dynamics, for which the body center of mass (COM) is considered the physiological reference standard. Therefore, this work proposes a simplified kinematic model (KM), which is designed to map the wrist IMU to the COM acceleration. It is built upon several reductive assumptions that enable the solvability of the dynamic equations based on wrist IMU measurements alone. This work further proposes three types of hybrid AI modeling methods, namely human kinematic model-based neural network (HKM-NN) models, to leverage the power of both grey-box and black-box modeling. The HKM-NN methods include serial learning (ser-) and two approaches of simultaneous learning (sim1- and sim2-). The proposed models are trained and tested using our dataset, which includes wrist IMU measurements and ground-truth COM measurements from 10 healthy volunteers during six gait activities and sit-to-stand (SS) transitional movement. The results demonstrate the feasibility of estimating COM acceleration from wrist IMU measurements. Our KM model yields satisfactory results, with an error ranging from 6.7% to 12.5% for gait activities and 5.6% for the SS. In comparison with the KM model, our HKM-NN models significantly enhance the performance, achieving 5.3% to 9.3% errors for gait activities, and the best error of 3.9% for the SS. In addition, the HKM-NN models demonstrate distinct robustness characteristics under noisy test conditions, with sim1-/sim2- generally maintaining greater robustness under Gaussian perturbations, while the KM model exhibits comparatively strong robustness under salt-and-pepper noise. These findings highlight the importance of combining biomechanical structure with data-driven learning for wearable sensing applications operating under imperfect and noisy measurement conditions.

AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

arXiv:2609.12320v1 Announce Type: new Abstract: Traditional large language models (LLMs) are scoped to individual user sessions, limiting their knowledge to a single conversation and preventing them from learning user preferences that evolve over time. Existing agentic memory systems address this limitation but generally operate at the individual-user level, restricting the public knowledge that could be shared across users to improve downstream responses. We introduce AIM (Agentic Interoperable Memory), a unified, privacy-aware memory framework that enables multi-agent, multi-user LLM systems to persistently manage private and shared memory. AIM dynamically classifies information as private, scoped to one user and inaccessible to others, or public, accessible to all users. It enforces index-level access controls so that private memories are retrievable only by their owner, protecting sensitive data while allowing beneficial shared knowledge to improve coordination and consistency. We also introduce MUMBench (Multi-User Memory Benchmark), a dataset of multi-user interactions containing private and shareable information across four domains. To our knowledge, MUMBench is the first public dataset designed to evaluate multiple memory operations, including retrieval, creation, update, and deletion, in a multi-user environment. Across three independent runs on MUMBench, AIM achieves 96.0% visibility classification accuracy, 58.8% strict operation accuracy, and 70.5% state-aware operation accuracy.

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.

Is Gaussian Splatting Becoming Neural Again? A Taxonomy and Controlled Study of Learned Parameterization

arXiv:2609.12395v1 Announce Type: new Abstract: Three-dimensional Gaussian Splatting (3DGS) combines explicit primitives with efficient rasterization, yet recent systems increasingly use neural networks to generate or share Gaussian parameters. We characterize this trend along five axes: attribute decoding, spatial sharing, view-conditioned decoding, topology generation, and amortized inference. An analysis of 19 representative methods shows that these choices address different limitations and cannot be reduced to a binary neural label. We also isolate three forms of neural parameterization in a controlled mip-NeRF 360 study. Sharing appearance and opacity improves reconstruction quality, while decoding geometric structure offers no further gain. The evidence favors selective neuralization: shared functions help when they capture reusable correlations without sacrificing the local geometric freedom of explicit splats.

OneLA: Scaling Linear-Attention Decoding to Large Beams in Generative Recommendation

arXiv:2609.12399v1 Announce Type: new Abstract: Generative recommendation (GR) relies on large-beam decoding to generate hundreds of candidate items, creating a new scaling challenge for recurrent linear attention. Existing linear attention serving systems either materialize a full recurrent state for every beam or repeatedly replay shared history, incurring substantial memory and traffic overhead. To address this, we present OneLA, a linear-attention decoding framework that exploits the shared prompt and short divergent suffixes of GR workloads. Specifically, OneLA represents all beam states using a single shared prompt-derived state and compact, append-only records of their divergent transitions. Using this representation, OneLA computes only the state information required at each decoding step, without reconstructing a full recurrent state for every beam. Furthermore, OneLA uses a lightweight ancestry index to track the transition records that make up each beam's history, allowing beams to be updated without moving or copying existing records. A fused GPU kernel further reuses the shared state across beams. Our analysis shows that OneLA achieves 1.54-2.46x end-to-end decode speedups while substantially reducing recurrent-state memory use and data movement.

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal composition. We introduce EvoRS, a self-evolving RL framework that evolves the reward system from on-policy experience, representing it as an executable Reward-DAG. Specifically, an agentic designer updates this system from on-policy rollouts and reward traces to maintain train-time reliability. Across writing and roleplay, EvoRS achieves the best quality under all three judges, outperforming the policy by \(2.107\) and \(4.767\) points, respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations confirm that a comprehensive fixed reward system cannot remain reliable in open-ended tasks and must evolve throughout training.
❌