❌

Normal view

NBR1-Mediated Autophagic Degradation of YTHDF1 Curtails <em>FDX1</em> Translation to Drive Concurrent Multikinase Inhibitor Resistance and Cuproptosis Tolerance

Cancer Commun (Lond). 2026 Sep 11;46:0048. doi: 10.34133/cancomm.0048. eCollection 2026.

ABSTRACT

Background: Cancer cells frequently acquire adaptive resistance to targeted therapies; however, strategies capable of concurrently overcoming treatment tolerance and reactivating cell death pathways are currently lacking. Here, we investigated the dual role of ferredoxin 1 (FDX1) in modulating both multikinase inhibitor (MKI) sensitivity and cuproptosis susceptibility in hepatocellular carcinoma (HCC), and sought to develop a therapeutic approach for reversing resistance. Methods: HCC models, both in vitro and in vivo, were employed to investigate the role of FDX1 in MKI resistance and cuproptosis evasion. Polysome profiling, SunTag translation reporters, CRISPR-Cas9 mutagenesis, and mass spectrometry were employed to delineate the underlying mechanisms. A codelivery nanoliposome system was engineered and tested in orthotopic HCC models. Results: Prolonged exposure to MKIs led to the down-regulation of FDX1 protein levels, resulting in MKI resistance and cuproptosis tolerance in HCC both in vitro and in vivo. Mechanistically, we found that MKIs inactivated protein kinase B (PKB, also known as AKT)-mechanistic target of rapamycin (mTOR) signaling, thereby suppressing the SET and MYND domain-containing protein 2 (SMYD2)-mediated methylation of YTH domain family protein 1 (YTHDF1) at lysine 515 (K515). Hypomethylated YTHDF1 was degraded via next to BRCA1 gene 1 protein (NBR1)-dependent autophagy, leading to the repression of N6-methyladenosine modification-dependent translation of FDX1 mRNA. FDX1 deficiency drove MKI resistance by reactivating AKT survival signaling while impairing cuproptosis through reduced divalent copper ions (Cu2+) to monovalent copper ions (Cu+) conversion and the loss of protein lipoylation. Additionally, restoring FDX1 expression through NBR1 knockdown or YTHDF1 overexpression overcame MKI resistance and resensitized HCC cells to cuproptosis. Finally, a nanoliposomal system, super cuproptosis detonator liposome, designed for the codelivery of NBR1 small interfering RNA, a copper ionophore, and sorafenib restored FDX1-dependent cuproptosis and exhibited marked anti-HCC efficacy, suppressing HCC growth in vivo. Conclusions: MKIs suppressed SMYD2-mediated YTHDF1 methylation at K515 via the inactivation of AKT-mTOR signaling. This led to the inhibition of FDX1 translation, resulting in AKT signaling reactivation and protein lipoylation impairment, effects that contributed to both MKI resistance and cuproptosis tolerance in HCC. Overcoming MKI resistance and resensitizing cells to cuproptosis by targeting NBR1-mediated YTHDF1 degradation using a nanoliposomal codelivery system represents a promising strategy for HCC treatment.

PMID:42729649 | PMC:PMC13562797 | DOI:10.34133/cancomm.0048

Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G

arXiv:2609.09591v1 Announce Type: cross Abstract: Sixth-generation (6G) wireless networks are expected to provide a key infrastructure for large-scale embodied intelligence, where heterogeneous robots collaborate through low-latency connectivity, edge intelligence, and distributed sensing. Vision-language-action (VLA) models offer a foundation by integrating visual perception, language understanding, and action generation into a unified closed-loop policy. However, training and adapting VLA models to distributed robotic agents introduce challenges in privacy protection, communication efficiency, and model heterogeneity. Existing federated learning (FL) methods overlook the intrinsic differences among vision, language, and action pathways in parameter scale, privacy exposure, update dynamics, and tolerance to compression or perturbation. To address this issue, this article proposes FedMVLA, a modality-decoupled FL framework for privacy-preserving embodied intelligence in 6G networks. FedMVLA incorporates three mechanisms: modality-aware federated aggregation (MAFA), modality-aware privacy allocation (MAPA), and modality-aware communication compression (MACO), together with a modality-sliced transport design that routes the precision-critical action stream through a protected ultra-reliable low-latency slice. A case study on federated robotic manipulation over the Third Generation Partnership Project (3GPP)-based wireless substrate, covering fading, co-channel interference, and malicious jamming, shows that FedMVLA achieves an 84.8% task success rate, exceeds FedAvg by 22.2 percentage points, sustains a widening margin when scaling to 128 clients across eight cells, and reduces the schedule-averaged per-client uplink model-update payload by 95.6% (approximately 96%), while keeping the 95th percentile (p95) of the round-critical uplink completion time near 1.5s.

MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation

arXiv:2510.05124v3 Announce Type: replace-cross Abstract: We propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS employs three coordinated agents: User Agents designed to simulate diverse persona-driven behaviors by leveraging personality signifiers such as Zodiac Signs and MBTI types, a Dialog Agent executing task-oriented persuasion strategies and an Optimization Agent evaluating and refining dialogue outcomes. We further validate its effectiveness through users' Chain-of-Attitude (CoA) modeling and dedicated LLMs' persuasion assessment. This approach enables low-cost generation of training data without human annotation, addressing key industry challenges such as lack of user data, cold-start evaluation difficulties, and prompt inefficiency. Applied to a real-world marketing scenario, MADS significantly improved the persuasion capacity of small LLMs, increasing the organic traffic conversion rate by 22.4% (from 1.83% to 2.24%) , demonstrating clear business value.

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | PMC:PMC13555834 | DOI:10.3322/caac.70100

Targeting KRAS reprograms a Treg-dominant immunosuppressive microenvironment and sensitizes KRAS-mutant gastric adenocarcinoma to CTLA-4 immunotherapy

Sci China Life Sci. 2026 Sep 3. doi: 10.1007/s11427-026-3438-4. Online ahead of print.

ABSTRACT

Oncogenic KRAS mutations define a distinct molecular subset of gastric adenocarcinoma (GA), yet their impact on the tumor immune microenvironment remains incompletely understood. In this study, we established a genetically faithful and immunocompetent KRASG12D-driven mouse model of GA, together with matched organoids and cell lines, to investigate how oncogenic KRAS shapes tumor-immune interactions. KRAS-mutant tumors consistently developed an immunosuppressive microenvironment characterized by enrichment of regulatory T cells (Tregs), accompanied by reduced cytotoxic lymphocyte infiltration and intrinsic resistance to PD-1 blockade. Although pharmacologic targeting of KRAS effectively suppressed tumor growth and increased immune cell infiltration, functional immune analyses revealed persistent Treg-mediated immunosuppression that limited effective antitumor immunity. Mechanistically, TGF-β signaling was required to maintain Treg dominance and suppress effector T cell function in KRAS-driven tumors. Importantly, disruption of this suppressive axis through combined KRAS inhibition and CTLA-4 blockade attenuated TGF-β activity, impaired Treg function, and enhanced antitumor immune responses in vivo. Collectively, these findings identify oncogenic KRAS as a key regulator of TGF-β-dependent immune suppression in GA and provide mechanistic insight into immune evasion within this molecular subtype.

PMID:42714795 | DOI:10.1007/s11427-026-3438-4

Early stage nonsmall cell lung cancer: Toward a risk-adaptive paradigm in the era of biologic precision

CA Cancer J Clin. 2026 Sep-Oct;76(5):e70100. doi: 10.3322/caac.70100.

ABSTRACT

The clinical landscape of early stage nonsmall cell lung cancer is at transformative crossroads. Driven by the widespread adoption of low-dose computed tomography screening, the frequent detection of ground-glass opacities, and a rising incidence among never-smokers, the diagnostic center of gravity has shifted toward earlier, potentially curable disease. This shift has been accompanied by equally important therapeutic advances, including parenchyma-sparing surgical techniques, minimally invasive platforms enhanced by digital navigation, and the transformative integration of perioperative immunotherapy and targeted agents. Concurrently, noninvasive monitoring approaches, such as liquid biopsy, have emerged as powerful tools to guide precision management. Despite this progress, substantial barriers to achieving a universal cure persist. Clinicians continue to face uncertainty in the management of ground-glass opacities, the anatomy-based TNM staging system fails to capture the biologic heterogeneity of early tumors, and global disparities in access to innovation remain unresolved. To address these challenges, the authors propose a shift toward a risk-adaptive management paradigm that harnesses artificial intelligence-driven analytics and multi-omics profiling to tailor treatment intensity according to each patient's biologic risk. Such an approach would enable appropriate escalation for high-risk individuals while permitting safe de-escalation for those at low risk. This holistic, lifespan-oriented strategy must be embraced to deliver equitable and durable cures for patients with early stage nonsmall cell lung cancer.

PMID:42713910 | DOI:10.3322/caac.70100

Multi-Omics Biomarker Signatures for Precision Diagnosis and Prognosis in Primary Liver Cancer: A Literature Review

4 September 2026 at 18:00

Biofactors. 2026 Sep-Oct;52(5):e70136. doi: 10.1002/biof.70136.

ABSTRACT

Primary liver cancer (PLC) is a biologically heterogeneous group of malignancies dominated by hepatocellular carcinoma (HCC), intrahepatic cholangiocarcinoma (iCCA), and a smaller subset of combined hepatocellular-cholangiocarcinoma (cHCC-CCA), and its clinical burden remains high because current diagnostic and prognostic tools do not adequately capture molecular diversity. Conventional imaging, serum markers, and histopathological assessment remain insufficient for precise early diagnosis, subtype-resolved classification, and outcome stratification, while tissue and liquid biopsy approaches have expanded the range of analytes available for clinical assessment. Recent studies have identified candidate biomarker signatures across genomic, epigenomic, transcriptomic, proteomic, metabolomic, and circulating layers, suggesting that integrated multi-omics profiling may better represent tumor lineage, clonal evolution, immune context, and therapeutic vulnerability than isolated molecular readouts. However, these layers are not equally mature for clinical use: genomic testing is closest to routine therapeutic application in iCCA, plasma methylation assays are advancing for HCC surveillance augmentation, and many proteomic or metabolomic panels remain validation-stage tools. Their clinical value remains constrained by sampling bias, biospecimen-dependent signal loss, assay standardization, cost, and the need for prospective validation across clinically diverse populations. This narrative review critically synthesizes current evidence on multi-omics biomarker signatures for precision diagnosis and prognosis in primary liver cancer and argues that clinically useful signatures should be question-specific, stage-aware, and specimen-aware rather than universal multi-analyte panels.

PMID:42697859 | PMC:PMC13545153 | DOI:10.1002/biof.70136

CAFs shape the immunosuppressive microenvironment of pancreatic cancer through the Lin28b-STING Axis

Nat Commun. 2026 Aug 7;17(1):9491. doi: 10.1038/s41467-026-76495-3.

ABSTRACT

Cancer-associated fibroblasts comprise diverse functionally distinct cellular subsets, with certain subpopulations exerting pivotal influence in shaping the pancreatic cancer immune microenvironment. Here we show that Lin28b+ cancer-associated fibroblasts contribute to establishing an immunologically cold tumor microenvironment in pancreatic ductal adenocarcinoma. Mechanistically, Lin28b directly binds to STING mRNA and promotes its degradation, thereby suppressing STING expression and downstream type I interferon signaling. Loss of Lin28b in cancer-associated fibroblasts activates the cGAS-STING-interferon signaling cascade, enhancing dendritic cell antigen presentation and CD8+ T cell cytotoxic function. Importantly, genetic inhibition of Lin28b in cancer-associated fibroblasts enhances sensitivity to anti-PD-L1 immune checkpoint blockade therapy. These findings reveal that targeting the Lin28b-STING axis represents a promising therapeutic strategy for overcoming the intrinsic resistance of pancreatic ductal adenocarcinoma to immunotherapy.

PMID:42693143 | PMC:PMC13542369 | DOI:10.1038/s41467-026-76495-3

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, mining valid supervision faces three challenges: (1) Label Noise via Mimetic Bias, where rewards prioritize statistical likelihood over logical truth, creating a "correctness illusion" that masks compounding errors; (2) Coarse-Grained Supervision, where sparse global outcomes (e.g., in GRPO) fail to provide granular guidance, treating reasoning chains as monolithic; and (3) Distributional Collapse, where signals fail to generalize without amplifying pre-training biases. To address these, we introduce LC-ERD (Logic-Consistent Endogenous Reward Decomposition), a framework framing self-alignment as latent structure mining. We derive a Variational Logic Potential by aggregating consensus from the model's Latent Logic Expertise (LLE) to denoise the reasoning manifold, and introduce a Multi-Agent Value Decomposition protocol based on the IGM principle to quantify individual step utility. Experiments show LC-ERD delivers a robust self-evolution path, uncovering trade-offs between logic consistency and accuracy while identifying high-value reasoning patterns missed by standard rewards. Our code is available at https://github.com/Reinhardmannn/LC-ERD.

HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models

arXiv:2605.24140v1 Announce Type: new Abstract: Multi-step reasoning remains a central challenge for large language models: single-pass generation is efficient but lacks accuracy; tree-search methods explore multiple paths but are computation-heavy. We address this gap by distilling reasoning progress into a hyperbolic geometric signal that guides step-by-step generation. Our approach is motivated by a structural observation: in combinatorial reasoning trees, solution-bearing states are few while dead ends are exponentially numerous. The hyperbolic space matches this asymmetry, with compact volume near the origin and exponentially expanding capacity toward the boundary, so that distance-to-origin naturally encodes solution proximity while angular separation distinguishes branches requiring different next operations. We train a lightweight head to project LLM hidden states into this space, then fine-tune a low-rank adapter interactively on its own reasoning attempts to act on the injected signal. Across multiple benchmarks, the geometric signal yields consistent gains, with larger improvements on deeper reasoning chains. Our code is publicly available at https://github.com/yuyuliu11037/HyperGuide.

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term this phenomenon Structural Cognitive Overload (SCO), a byproduct of the contention between deep reasoning and safety alignment. However, prior work has predominantly targeted typographic and pixel-level perturbations, leaving the study of SCO largely unexplored. To this end, we propose StructBreak, an automated end-to-end framework designed to quantify SCO. By leveraging StructBreak, we uncover a novel higher-order cognitive overload attack paradigm; notably, this attack operates under a practical black-box setting, requiring no internal model access. Consequently, we utilize this framework to establish a comprehensive benchmark spanning ten diverse threat scenarios. Empirical evaluations on six leading MLLMs reveal that SCO readily triggers toxic generation, yielding a 92% average ASR (up to 97% on Gemini 2.5). To elucidate the mechanism of SCO, we further conduct model-level interpretations spanning attention dynamics, latent space topology, and geometric analysis. Our findings reveal that StructBreak acts as a novel structural channel to circumvent safety filters. Furthermore, the limited efficacy of inherent safety mechanisms underscores that current alignment paradigms are insufficient for the era of complex multimodal reasoning.

PHGNet: Prototype-Guided Hypergraph Construction for Heterogeneous Spatiotemporal Forecasting

arXiv:2605.25554v1 Announce Type: new Abstract: As a core task in intelligent transportation systems, traffic forecasting plays a critical role in urban traffic management. Accurate traffic forecasting relies on modeling complex spatiotemporal dependencies, which is inherently challenging due to spatial heterogeneity in traffic systems.Despite significant progress, most existing methods are still limited to pairwise spatial dependency modeling, making it difficult to capture dynamic high-order interactions among nodes with similar traffic patterns. To address this issue, we propose PHGNet, a novel spatiotemporal forecasting framework based on prototype-guided hypergraph construction. At the core of PHGNet, a prototype learning mechanism is designed to adaptively assign pattern-similar nodes to hyperedges, thereby capturing high-order interactions with time-varying structures. To improve the reliability of dynamic hypergraph construction, we further develop a global-local node representation module to extract time-consistent features. For forecasting, iterative residual refinement and Temporal Query Attention are introduced to improve forecasting accuracy while supporting efficient parallel decoding. Extensive experiments on multiple real-world datasets demonstrate that PHGNet achieves superior predictive performance compared with state-of-the-art methods.

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

arXiv:2605.24860v1 Announce Type: cross Abstract: Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stability. The performance of ADAS critically relies on accurate and reliable vehicle state estimation, particularly from vehicle dynamic sensors. Among these signals, wheel load is a key variable for chassis control and safety-critical functions, yet it remains difficult to estimate robustly due to complex suspension geometry, nonlinear dynamics, and measurement noise. To address this issue, we propose DBPnet, a Bayesian physics-informed neural network (PINN) with a physics-aware embedding module inspired by damper characteristics. First, this paper presents a suspension linkage-level modeling (SLLM) approach that constructs a nonlinear instantaneous dynamic model by explicitly considering the complex geometric structure of the suspension. Building upon SLLM, Bayesian inference is integrated into the PINN to effectively cope with noise and uncertainty in the vehicle chassis system, thereby improving the model's robustness. Then, a physics-informed loss function is employed to ensure consistency with fundamental physical principles, while the damper characteristics-inspired embedding module extracts temporal variation features of input signals and incorporates them into each layer of the PINN, ensuring that physical observations guide the neural network without being constrained by fixed physical models. Extensive evaluations on high-fidelity simulations and real-world experiments demonstrate that our DBPnet consistently achieves lower RMSE and MaxError than baseline methods. These results highlight the potential of our DBPnet to advance wheel load estimation and contribute to the development of more reliable ADAS actuator functions.

STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media

arXiv:2605.25162v1 Announce Type: cross Abstract: Large language models for vertical domains are bottlenecked by the scarcity of complex, domain-specific task-oriented dialogues. Existing data acquisition pipelines face a persistent trilemma: expert annotation is expensive, real-world service conversations are constrained by privacy and commercial restrictions, and static corpora quickly become temporally stale. We propose Stream, a data-centric framework that leverages publicly available streaming media (live streams and short videos) to synthesize high-value service dialogues at scale. Stream mines authentic interaction signals from noisy streams and synthesizes conversations by integrating role-grounded persona construction with Conversational Blueprint construction; it further adopts retrieval-augmented generation (RAG) to support knowledge-aware responses. Based on Stream, we release StreamDial, a large-scale multi-domain dataset covering Automotive, Restaurant, and Hotel. StreamDial contains 87,498 dialogue sessions and 1,497,320 turns in total, with an average of 17.11 turns per session and a comparable scale across domains. Each session is organized as a structured quadruplet $\langle P_u, P_a, B, H \rangle$ that pairs dialogue history with explicit user/agent personas and a Conversational Blueprint, capturing realistic service behaviors such as requirement mining, constraint conflicts, negotiation, and recovery. Evaluations with automatic judges and downstream tasks show that StreamDial improves intrinsic dialogue quality over strong baselines, and models trained with StreamDial improve Dialogue State Tracking across backbones; we further report a completed human-evaluation set and encouraging multilingual transfer on Qwen3-8B under a controlled training budget. The data is released in https://github.com/hitxueliang/DialogDataSetBySTREAM.

Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective

arXiv:2605.02010v2 Announce Type: replace Abstract: This position paper argues that reliable AI requires infrastructure for human validation of implicit knowledge. AI learns from both explicit knowledge (papers, documentation, structured databases) and implicit knowledge (reasoning patterns, debugging processes, intermediate steps). Implicit knowledge remains unexternalized because documentation cost exceeds perceived value -- yet AI learns from it indiscriminately, acquiring both beneficial patterns and harmful biases. Current reliability methods can only verify explicit knowledge against sources, creating a fundamental gap: the most valuable AI capabilities (reasoning, judgment, intuition) are precisely those we cannot verify. We propose Knowledge Objects (KOs) -- structured artifacts that externalize implicit knowledge into forms humans can inspect, verify, and endorse. KOs transform verification economics: what was previously too costly to verify becomes feasible, enabling accumulated human validation to improve reliability over time.

Coupled Variational Reinforcement Learning for Language Model General Reasoning

arXiv:2512.12576v3 Announce Type: replace-cross Abstract: While reinforcement learning has achieved impressive progress in language model reasoning, it is constrained by the requirement for verifiable rewards. Recent verifier-free RL methods address this limitation by utilizing the probabilities that LLMs generate reference answers as reward signals. However, these approaches typically sample reasoning traces conditioned only on the question. This design decouples reasoning-trace sampling from answer information, leading to inefficient exploration and incoherence between traces and final answers. In this paper, we propose \textit{\b{Co}upled \b{V}ariational \b{R}einforcement \b{L}earning} (CoVRL), which bridges variational inference and reinforcement learning by coupling prior and posterior distributions through a hybrid sampling strategy. By constructing and optimizing a composite distribution that integrates these two distributions, CoVRL enables efficient exploration while preserving strong thought-answer coherence. Extensive experiments on mathematical and general reasoning benchmarks show that CoVRL improves performance by 12.4\% over the base model and achieves an additional 2.3\% improvement over state-of-the-art verifier-free RL baselines, providing a principled framework for enhancing the general reasoning capabilities of language models.

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

arXiv:2605.02900v2 Announce Type: replace-cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act under uncertain sensing, incomplete knowledge, and dynamic human-robot interactions, where failures can directly lead to physical harm. This survey provides a comprehensive and structured review of safety research in embodied AI, examining attacks and defenses across the full embodied pipeline, from perception and cognition to planning, action and interaction, and agentic system. We introduce a multi-level taxonomy that unifies fragmented lines of work and connects embodied-specific safety findings with broader advances in vision, language, and multimodal foundation models. Our review synthesizes insights from over 500 papers spanning adversarial, backdoor, jailbreak, and hardware-level attacks; attack detection, safe training and robust inference; and risk-aware human-agent interaction. This analysis reveals several overlooked challenges, including the fragility of multimodal perception fusion, the instability of planning under jailbreak attacks, and the trustworthiness of human-agent interaction in open-ended scenarios. By organizing the field into a coherent framework and identifying critical research gaps, this survey provides a roadmap for building embodied agents that are not only capable and autonomous but also safe, robust, and reliable in real-world deployment.

SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

arXiv:2605.10989v3 Announce Type: replace-cross Abstract: The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.

Simply Stabilizing the Loop via Fully Looped Transformer

arXiv:2605.18797v2 Announce Type: replace-cross Abstract: Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading additional computation for improved performance without increasing parameter count or context length. Because the number of loop iterations can be adjusted at inference, it also provides a natural mechanism for balancing performance and test-time compute. However, Looped Transformer still suffers from training instability when the number of loop iterations increases. Our analysis reveals that this instability stems from two sources: gradient oscillation and residual explosion. To address these two problems, we propose the Fully Looped Transformer, which introduces two parameter-free modifications: (1) Fully Looped Architecture, which distributes inter-loop signals across all layers to mitigate residual explosion; (2) Attention Injection, which reuses the existing attention block to suppress gradient oscillation. These modifications stabilize training dynamics, enabling the Fully Looped Transformer to be trained stably up to 12 loop iterations, whereas other baseline looped models collapse in this regime. In milder settings where Looped Transformer does not collapse, Fully Looped Transformer still improves average downstream-task performance by up to 13.2\%. Overall, our experiments demonstrate that Fully Looped Transformer improves training stability, enhances downstream performance, and provides preliminary adaptability under different test-time compute budgets by varying loop iterations at inference.
❌