❌

Reading view

SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning

arXiv:2609.12599v1 Announce Type: cross Abstract: Multi-task learning (MTL) requires navigating unavoidable trade-offs among competing objectives. This paradigm is frequently formulated as multi-objective optimization (MOO), where the scalarization is favored to reduce an MOO problem to a single objective. We empirically find that existing merit-function-based scalarization approaches are sensitive to the relative scales of different objectives in practical MTL, where task losses commonly differ by orders of magnitude. The optimization process often favors objectives with larger scales even though the underlying Pareto optimal solutions remains invariant to rescaling (i.e., multiplying an objective by a positive constant). To address this issue, we propose Scale-Invariant Merit-function-based Scalarization (SIMS) for MTL. Specifically, SIMS adopts a transformation-induced merit function to convert the MOO problem of MTL to a single objective that renders optimization invariant to the magnitudes of losses. Theoretically, we prove that the requirement for scale invariance uniquely determines this transformation to be logarithmic. We further show that this general transformation-induced merit function preserves weak Pareto optimality and admits a smooth surrogate with controllable approximation error. Extensive experiments on representative multi-task benchmarks demonstrate that SIMS consistently outperforms existing scalarization methods and achieves state-of-the-art performance.
  •  

From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

arXiv:2505.08548v3 Announce Type: replace-cross Abstract: Achieving generalization in robotic manipulation remains a critical challenge, particularly for unseen scenarios and novel tasks. Current Vision-Language-Action (VLA) models, while building on top of general Vision-Language Models (VLMs), still fall short of achieving robust zero-shot performance due to the scarcity and heterogeneity prevalent in embodied datasets. To address these limitations, we propose FSD (From Seeing to Doing), a novel vision-language model that generates intermediate representations through spatial relationship reasoning, providing fine-grained guidance for robotic manipulation. Our approach combines a hierarchical data pipeline for training with a self-consistency mechanism that aligns spatial coordinates with visual signals. Through extensive experiments, we comprehensively validated FSD's capabilities in both "seeing" and "doing," achieving outstanding performance across 8 benchmarks for general spatial reasoning and embodied reference abilities, as well as on our proposed more challenging benchmark VABench. We also verified zero-shot capabilities in robot manipulation, demonstrating significant performance improvements over baseline methods in both SimplerEnv and real robot settings. Experimental results show that FSD achieves 40.6% success rate in SimplerEnv and 72% success rate across 8 real-world tasks, outperforming the strongest baseline by 30%.
  •  

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

arXiv:2508.13998v2 Announce Type: replace-cross Abstract: Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation, defining four core embodied pointing abilities that bridge high-level vision-language comprehension with low-level action primitives. We introduce Embodied-R1, a 3B Vision-Language Model (VLM) specifically designed for embodied reasoning and pointing. We use a wide range of embodied and general visual reasoning datasets as sources to construct a large-scale dataset, Embodied-Points-200K, which supports key embodied pointing capabilities. We then train Embodied-R1 using a two-stage Reinforced Fine-tuning (RFT) curriculum with a specialized multi-task reward design. Embodied-R1 achieves state-of-the-art performance on 11 embodied spatial and pointing benchmarks. Critically, it demonstrates robust zero-shot generalization by achieving a 56.2% success rate in the SIMPLEREnv and 87.5% across 8 real-world XArm tasks without any task-specific fine-tuning, representing a 62% improvement over strong baselines. Furthermore, the model exhibits high robustness against diverse visual disturbances. Our work shows that a pointing-centric representation, combined with an RFT training paradigm, offers an effective and generalizable pathway to closing the perception-action gap in robotics.
  •  

Integrative Multi-Omics and Single-Cell Analysis Reveal THOC3 and THOC7 as Oncogenic RNA Processing Regulators in Lung Adenocarcinoma

Int J Med Sci. 2026 Mar 9;23(4):1408-1430. doi: 10.7150/ijms.128975. eCollection 2026.

ABSTRACT

Lung adenocarcinoma (LUAD) remains a leading cause of cancer-related mortality worldwide. Although the transcription-export (TREX) complex plays a central role in RNA maturation and nuclear export, the clinical and biological relevance of individual THO Complex Subunit (including THOC1, THOC2, THOC3, THOC5, THOC6, and THOC7) in LUAD is not well defined. We performed integrative analyses combining bulk transcriptomics from TCGA/GTEx and independent GEO cohorts, survival modeling, DNA methylation profiling, protein-level annotation from public resources, protein-protein interaction network analysis, immune infiltration estimation (TIMER), and single-cell RNA sequencing (scRNA-seq) to evaluate the relevance of THOC3 and THOC7 in LUAD. Across TCGA and external GEO validation datasets, THOC3 and THOC7 were consistently upregulated in LUAD and associated with poorer overall and disease-free survival, whereas other THO complex members showed weaker or inconsistent associations. Given these comparatively consistent and reproducible signals, we therefore prioritized THOC3 and THOC7 for downstream multi-layer analyses. Epigenetic profiling and interaction network analyses placed both genes within conserved RNA processing and export programs linked to genome maintenance pathways. Single-cell transcriptomic analysis provided additional resolution, demonstrating predominant enrichment of THOC3 and THOC7 in malignant epithelial clusters, with THOC3 aligning with transcriptional programs associated with DNA replication and repair, and THOC7 with proliferative and checkpoint-related states. Notably, expression of both genes was also detectable in myeloid and neutrophil subsets, and THOC7 expression remained elevated in recurrent LUAD samples, indicating association with aggressive and treatment-resistant disease states. Collectively, by integrating bulk, single-cell, epigenetic, and immune profiling across multiple independent cohorts, this study identifies THOC3 and THOC7 as reproducible molecular correlates of aggressive LUAD phenotypes. These highlight dysregulated RNA export programs as potential biomarkers of poor prognosis and motivate future functional studies to assess RNA export dependencies in LUAD.

PMID:41938520 | PMC:PMC13048885 | DOI:10.7150/ijms.128975

  •  

Integrative Multi-Omics and Single-Cell Analysis Reveal THOC3 and THOC7 as Oncogenic RNA Processing Regulators in Lung Adenocarcinoma

Int J Med Sci. 2026 Mar 9;23(4):1408-1430. doi: 10.7150/ijms.128975. eCollection 2026.

ABSTRACT

Lung adenocarcinoma (LUAD) remains a leading cause of cancer-related mortality worldwide. Although the transcription-export (TREX) complex plays a central role in RNA maturation and nuclear export, the clinical and biological relevance of individual THO Complex Subunit (including THOC1, THOC2, THOC3, THOC5, THOC6, and THOC7) in LUAD is not well defined. We performed integrative analyses combining bulk transcriptomics from TCGA/GTEx and independent GEO cohorts, survival modeling, DNA methylation profiling, protein-level annotation from public resources, protein-protein interaction network analysis, immune infiltration estimation (TIMER), and single-cell RNA sequencing (scRNA-seq) to evaluate the relevance of THOC3 and THOC7 in LUAD. Across TCGA and external GEO validation datasets, THOC3 and THOC7 were consistently upregulated in LUAD and associated with poorer overall and disease-free survival, whereas other THO complex members showed weaker or inconsistent associations. Given these comparatively consistent and reproducible signals, we therefore prioritized THOC3 and THOC7 for downstream multi-layer analyses. Epigenetic profiling and interaction network analyses placed both genes within conserved RNA processing and export programs linked to genome maintenance pathways. Single-cell transcriptomic analysis provided additional resolution, demonstrating predominant enrichment of THOC3 and THOC7 in malignant epithelial clusters, with THOC3 aligning with transcriptional programs associated with DNA replication and repair, and THOC7 with proliferative and checkpoint-related states. Notably, expression of both genes was also detectable in myeloid and neutrophil subsets, and THOC7 expression remained elevated in recurrent LUAD samples, indicating association with aggressive and treatment-resistant disease states. Collectively, by integrating bulk, single-cell, epigenetic, and immune profiling across multiple independent cohorts, this study identifies THOC3 and THOC7 as reproducible molecular correlates of aggressive LUAD phenotypes. These highlight dysregulated RNA export programs as potential biomarkers of poor prognosis and motivate future functional studies to assess RNA export dependencies in LUAD.

PMID:41938520 | PMC:PMC13048885 | DOI:10.7150/ijms.128975

  •  

Predicting LLM Output Length via Entropy-Guided Representations

arXiv:2602.11812v2 Announce Type: replace Abstract: The long-tailed distribution of sequence lengths in LLM serving and reinforcement learning (RL) sampling causes significant computational waste due to excessive padding in batched inference. Existing methods rely on auxiliary models for static length prediction, but they incur high overhead, generalize poorly, and fail in stochastic "one-to-many" sampling scenarios. We introduce a lightweight framework that reuses the main model's internal hidden states for efficient length prediction. Our framework features two core components: 1) Entropy-Guided Token Pooling (EGTP), which uses on-the-fly activations and token entropy for highly accurate static prediction with negligible cost, and 2) Progressive Length Prediction (PLP), which dynamically estimates the remaining length at each decoding step to handle stochastic generation. To validate our approach, we build and release ForeLen, a comprehensive benchmark with long-sequence, Chain-of-Thought, and RL data. On ForeLen, EGTP achieves state-of-the-art accuracy, reducing MAE by 29.16\% over the best baseline. Integrating our methods with a length-aware scheduler yields significant end-to-end throughput gains. Our work provides a new technical and evaluation baseline for efficient LLM inference.
  •  

Mechanisms Underlying Drought Adaptability in Duolang Sheep Based on Metabolomic and Transcriptomic Analyses

Biology (Basel). 2026 Mar 12;15(6):461. doi: 10.3390/biology15060461.

ABSTRACT

This study investigates the mechanisms underlying drought adaptability in Duolang sheep, a local breed from two distinct habitats in Xinjiang-an arid southern region and a grassland northern region-aiming to identify key factors driving differential environmental adaptation. Integrated multi-omics analyses were performed, including serum biochemical assays, untargeted metabolomics of perirenal and tail fat tissues, and transcriptomic profiling of lung, liver, and kidney samples. Our results revealed notable differences: (1) serum levels of GSH-Px, IL-2, and IgG were significantly higher in the southern group (p < 0.01); (2) metabolomic analysis identified key differential metabolites, including EPA (involved in unsaturated fatty acid biosynthesis), choline (glycerophospholipid metabolism), L-serine and glutathione (cofactor biosynthesis), and taurine (sulfur metabolism); and (3) transcriptomic analysis revealed significant differential expression of genes such as FGF21 (thermogenesis), CD14 and DUSP2 (MAPK signaling pathway), GOT1 (arginine biosynthesis), and AVPR2 (vasopressin-regulated water reabsorption). Integrative correlation analysis further indicated that glutathione, EPA, GOT1, and CD14 are involved in energy and lipid metabolism, while taurine, AVPR2, and DUSP2 contribute to oxidative stress resistance and immune regulation. These molecular and metabolic adjustments collectively enhance drought adaptability in southern Xinjiang Duolang sheep. In conclusion, adaptation to arid environments requires enhanced antioxidant capacity and immune function, with metabolites such as EPA supporting lipid metabolism and genes such as FGF21 regulating fatty acid oxidation to limit triglyceride accumulation.

PMID:41892221 | PMC:PMC13023672 | DOI:10.3390/biology15060461

  •  

From Noisy Labels to Intrinsic Structure: A Geometric-Structural Dual-Guided Framework for Noise-Robust Medical Image Segmentation

arXiv:2509.02419v2 Announce Type: replace-cross Abstract: The effectiveness of convolutional neural networks in medical image segmentation relies on large-scale, high-quality annotations, which are costly and time-consuming to obtain. Even expert-labeled datasets inevitably contain noise arising from subjectivity and coarse delineations, which disrupt feature learning and adversely impact model performance. To address these challenges, this study propose a Geometric-Structural Dual-Guided Network (GSD-Net), which integrates geometric and structural cues to improve robustness against noisy annotations. It incorporates a Geometric Distance-Aware module that dynamically adjusts pixel-level weights using geometric features, thereby strengthening supervision in reliable regions while suppressing noise. A Structure-Guided Label Refinement module further refines labels with structural priors, and a Knowledge Transfer module enriches supervision and improves sensitivity to local details. To comprehensively assess its effectiveness, we evaluated GSD-Net on six publicly available datasets: four containing three types of simulated label noise, and two with multi-expert annotations that reflect real-world subjectivity and labeling inconsistencies. Experimental results demonstrate that GSD-Net achieves state-of-the-art performance under noisy annotations, achieving improvements of 1.58% on Kvasir, 22.76% on Shenzhen, 8.87% on BU-SUC, and 1.77% on BraTS2020 under SR simulated noise. The codes of this study are available at https://github.com/ortonwang/GSD-Net.
  •  

Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation

arXiv:2603.07048v1 Announce Type: cross Abstract: Although large vision-language models (LVLMs) have demonstrated remarkable capabilities, they are prone to hallucinations in multi-image tasks. We attribute this issue to limitations in existing attention mechanisms and insufficient cross-image modeling. Inspired by this, we propose a structured hallucination mitigation framework involving Cross-Image Attention calibration and Preference Learning (CAPL). CAPL explicitly enhances inter-image interactions at the architectural level while reinforcing reliance on genuine cross-image evidence during training, thereby improving the model's perception and modeling of cross-image associations. Specifically, we (i) introduce a selectable image token interaction attention mechanism to establish fine-grained cross-image entity alignment and information flow; (ii) design a cross-image modeling-based preference optimization strategy that contrasts reasoning outcomes under full inter-image interaction and those obtained when images are mutually invisible, encouraging the model to ground its predictions in authentic visual evidence and mitigating erroneous inferences driven by textual priors. Experimental results demonstrate that CAPL consistently improves performance across multiple model architectures, achieving stable gains on both multi-image hallucination and general benchmarks. Notably, performance on single-image visual tasks remains stable or slightly improves, indicating strong generalization capability.
  •  

SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training

arXiv:2603.07927v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed the software engineering landscape. Recently, numerous LLM-based agents have been developed to address real-world software issue fixing tasks. Despite their state-of-the-art performance, Despite achieving state-of-the-art performance, these agents face a significant challenge: \textbf{Insufficient high-quality issue descriptions.} Real-world datasets often exhibit misalignments between issue descriptions and their corresponding solutions, introducing noise and ambiguity that mislead automated agents and limit their problem-solving effectiveness. We propose \textbf{\textit{SWE-Fuse}}, an issue-description-aware training framework that fuses issue-description-guided and issue-free samples for training SWE agents. It consists of two key modules: (1) An issue-free-driven trajectory learning module for mitigating potentially misleading issue descriptions while enabling the model to learn step-by-step debugging processes; and (2) An entropy-aware RLVR training module, which adaptively adjusts training dynamics through entropy-driven clipping. It applies relaxed clipping under high entropy to encourage exploration, and stricter clipping under low entropy to ensure training stability. We evaluate SWE-Fuse on the widely studied SWE-bench Verified benchmark shows to demonstrate its effectiveness in solving real-world software problems. Specifically, SWE-Fuse outperforms the best 8B and 32B baselines by 43.0\% and 60.2\% in solve rate, respectively. Furthermore, integrating SWE-Fuse with test-time scaling (TTS) enables further performance improvements, achieving solve rates of 49.8\% and 65.2\% under TTS@8 for the 8B and 32B models, respectively.
  •  

MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems

arXiv:2602.19843v1 Announce Type: cross Abstract: As LLM-based Multi-Agent Systems (MAS) are increasingly deployed for complex tasks, ensuring their reliability has become a pressing challenge. Since MAS coordinate through unstructured natural language rather than rigid protocols, they are prone to semantic failures (e.g., hallucinations, misinterpreted instructions, and reasoning drift) that propagate silently without raising runtime exceptions. Prevailing evaluation approaches, which measure only end-to-end task success, offer limited insight into how these failures arise or how effectively agents recover from them. To bridge this gap, we propose MAS-FIRE, a systematic framework for fault injection and reliability evaluation of MAS. We define a taxonomy of 15 fault types covering intra-agent cognitive errors and inter-agent coordination failures, and inject them via three non-invasive mechanisms: prompt modification, response rewriting, and message routing manipulation. Applying MAS-FIRE to three representative MAS architectures, we uncover a rich set of fault-tolerant behaviors that we organize into four tiers: mechanism, rule, prompt, and reasoning. This tiered view enables fine-grained diagnosis of where and why systems succeed or fail. Our findings reveal that stronger foundation models do not uniformly improve robustness. We further show that architectural topology plays an equally decisive role, with iterative, closed-loop designs neutralizing over 40% of faults that cause catastrophic collapse in linear workflows. MAS-FIRE provides the process-level observability and actionable guidance needed to systematically improve multi-agent systems.
  •  

Detecting Brick Kiln Infrastructure at Scale: Graph, Foundation, and Remote Sensing Models for Satellite Imagery Data

arXiv:2602.13350v1 Announce Type: cross Abstract: Brick kilns are a major source of air pollution and forced labor in South Asia, yet large-scale monitoring remains limited by sparse and outdated ground data. We study brick kiln detection at scale using high-resolution satellite imagery and curate a multi city zoom-20 (0.149 meters per pixel) resolution dataset comprising over 1.3 million image tiles across five regions in South and Central Asia. We propose ClimateGraph, a region-adaptive graph-based model that captures spatial and directional structure in kiln layouts, and evaluate it against established graph learning baselines. In parallel, we assess a remote sensing based detection pipeline and benchmark it against recent foundation models for satellite imagery. Our results highlight complementary strengths across graph, foundation, and remote sensing approaches, providing practical guidance for scalable brick kiln monitoring from satellite imagery.
  •  

GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization

arXiv:2602.13921v1 Announce Type: cross Abstract: Repository-level bug localization-the task of identifying where code must be modified to fix a bug-is a critical software engineering challenge. Standard Large Language Modles (LLMs) are often unsuitable for this task due to context window limitations that prevent them from processing entire code repositories. As a result, various retrieval methods are commonly used, including keyword matching, text similarity, and simple graph-based heuristics such as Breadth-First Search. Graph Neural Networks (GNNs) offer a promising alternative due to their ability to model complex, repository-wide dependencies; however, their application has been hindered by the lack of a dedicated benchmark. To address this gap, we introduce GREPO, the first GNN benchmark for repository-scale bug localization tasks. GREPO comprises 86 Python repositories and 47294 bug-fixing tasks, providing graph-based data structures ready for direct GNN processing. Our evaluation of various GNN architectures shows outstanding performance compared to established information retrieval baselines. This work highlights the potential of GNNs for bug localization and established GREPO as a foundation resource for future research, The code is available at https://github.com/qingpingmo/GREPO.
  •  
❌