❌

Normal view

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-world deployment faces a fundamental device--cloud dilemma: on-device models are efficient but often brittle, while cloud models are stronger but costly in computation. State-of-the-art LLM device--cloud routers usually make coarse task-level decisions, which cannot adapt to the changing difficulty of multi-step agent interactions. To address this issue, we present Hera, a step-level device--cloud LLM agent coordinator for long-horizon tasks achieving a strong performance--cost Pareto frontier. Hera adopts a novel two-stage training paradigm: (1) imitation learning for cold-start, followed by (2) reinforcement learning that jointly optimizes task success and cloud usage efficiency. The first stage casts step-level routing as a supervised classification problem: the device agent is replayed on cloud trajectories, with each state labeled by the agreement between device and cloud actions. In the second stage, we perform cost-aware reinforcement learning by grouping identical states across trajectories and updating Hera with labels favoring higher expected return and fewer future cloud calls. We evaluate Hera on ALFWorld, WebShop, and AppWorld, where it consistently outperforms prior methods, achieving 92.5% of the cloud-only success rate with cloud use in only 46.3% of steps.

Electric dipole moment drives the dynamics of the TNFR1 complex I signalosome

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10304-1

Long-range interactions mediated by protein electric dipole moments have a role in driving the assembly and disassembly of super-signalling complex I for promoting NF-κB signalling.

Trem1 regulates neutrophil metabolism and recruitment in lung ischemia-reperfusion injury

Redox Biol. 2026 Jan 14;92:104026. doi: 10.1016/j.redox.2026.104026. Online ahead of print.

ABSTRACT

Primary graft dysfunction (PGD) caused by ischemia-reperfusion injury (IRI) is a major complication after lung transplantation, yet its underlying mechanisms remain unclear. Triggering receptor expressed on myeloid cells 1 (Trem1) is an important mediator of inflammation, but its role in neutrophil function and metabolic reprogramming during lung IRI is not well understood. In this study, we used a murine orthotopic lung transplantation model with cold ischemia and reperfusion, and Trem1 knockout (Trem1-/-) and myeloid-specific Trem1 conditional knockout mice (LysmCreTrem1fl) to explore the role of Trem1 in neutrophil recruitment, neutrophil extracellular trap (NET) formation, and metabolism. Our results show that Trem1 expression increases in both mouse and human lungs after reperfusion and correlates with neutrophil infiltration and lung injury. Trem1 deficiency significantly reduced neutrophil and macrophage recruitment, NET formation, and tissue damage. Multi-omics analysis revealed that Trem1 deletion suppressed oxidative phosphorylation (OXPHOS) and induced a metabolic shift in neutrophils toward glycolysis. In clinical samples, the abundance of TREM1+ neutrophils was correlated with PGD severity and OXPHOS activity. These findings identify Trem1 as a key regulator of neutrophil metabolism and recruitment in lung IRI, and suggest that targeting Trem1 may provide a novel therapeutic strategy to mitigate PGD and improve lung transplant outcomes.

PMID:41861599 | DOI:10.1016/j.redox.2026.104026

Top-Down Semantic Refinement for Image Captioning

arXiv:2510.22391v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (VLMs) face an inherent contradiction in image captioning: their powerful single-step generation capabilities often lead to a myopic decision-making process. This makes it difficult to maintain global narrative coherence while capturing rich details, a limitation that is particularly pronounced in tasks that require multi-step and complex scene description. To overcome this fundamental challenge, we redefine image captioning as a goal-oriented hierarchical refinement planning problem, and further propose a novel framework, named Top-Down Semantic Refinement (TDSR), which models the generation process as a Markov Decision Process (MDP). However, planning within the vast state space of a VLM presents a significant computational hurdle. Our core contribution, therefore, is the design of a highly efficient Monte Carlo Tree Search (MCTS) algorithm tailored for VLMs. By incorporating a visual-guided parallel expansion and a lightweight value network, our TDSR reduces the call frequency to the expensive VLM by an order of magnitude without sacrificing planning quality. Furthermore, an adaptive early stopping mechanism dynamically matches computational overhead to the image's complexity. Extensive experiments on multiple benchmarks, including DetailCaps, COMPOSITIONCAP, and POPE, demonstrate that our TDSR, as a plug-and-play module, can significantly enhance the performance of existing VLMs (e.g., LLaVA-1.5, Qwen2.5-VL) by achieving state-of-the-art or highly competitive results in fine-grained description, compositional generalization, and hallucination suppression.
❌