❌

Reading view

NED-Tree: Bridging the Semantic Gap with Nonlinear Element Decomposition Tree for LLM Nonlinear Optimization Modeling

arXiv:2604.01588v1 Announce Type: new Abstract: Automating the translation of Operations Research (OR) problems from natural language to executable models is a critical challenge. While Large Language Models (LLMs) have shown promise in linear tasks, they suffer from severe performance degradation in real-world nonlinear scenarios due to semantic misalignment between mathematical formulations and solver codes, as well as unstable information extraction. In this study, we introduce NED-Tree, a systematic framework designed to bridge the semantic gap. NED-Tree employs (a) a sentence-by-sentence extraction strategy to ensure robust parameter mapping and traceability; and (b) a recursive tree-based structure that adaptively decomposes complex nonlinear terms into solver-compatible sub-elements. Additionally, we present NEXTOR, a novel benchmark specifically designed for complex nonlinear, extensive-constraint OR problems. Experiments across 10 benchmarks demonstrate that NED-Tree establishes a new state-of-the-art with 72.51% average accuracy, NED-Tree is the first framework that drives LLMs to resolve nonlinear modeling difficulties through element decomposition, achieving alignment between modeling semantics and code semantics. The NED-Tree framework and benchmark are accessible in the anonymous repository https://anonymous.4open.science/r/NORA-NEXTOR.
  •  

Integrated spatial transcriptomics and pan-cancer XGBoost modeling uncover spatial drivers of immune exclusion and predict immunotherapy response

Cancer Immunol Immunother. 2026 Apr 2;75(4):131. doi: 10.1007/s00262-026-04374-3.

ABSTRACT

Immunotherapy has revolutionized cancer treatment, yet characterizing the spatial complexity of the tumor immune microenvironment remains a challenge. In this study, we established a comprehensive computational framework integrating multi-omics profiling across 27 cancer types to decode immune-related non-coding RNA regulatory networks. Moving beyond traditional bulk analysis, we utilized spatial transcriptomics to dissect the spatial localization of these regulators. We identified the SNHG6-BIRC5 axis as a critical driver of the "immune-cold" phenotype in lung adenocarcinoma. We provide visual evidence that this axis localizes to tumor nests and negatively correlates with T- cell infiltration, elucidating a mechanism of spatial immune exclusion. Validating the clinical relevance of these findings, genome-scale CRISPR-Cas9 screening data confirmed the functional essentiality of these targets for cancer cell survival. Furthermore, pharmacogenomic analysis revealed that high expression of this axis correlates with sensitivity to chemotherapy agents like Vinblastine, suggesting a potential stratification strategy for patients with immune-excluded tumors. To expand the clinical utility to immunotherapy prediction, we developed a pan-cancer XGBoost machine learning model incorporating 14 high-performance regulatory features. This model achieved robust performance in distinguishing immunotherapy responders from non-responders with an AUC of 0.771, outperforming traditional markers such as PD-L1. Collectively, this study highlights spatial determinants of immune exclusion and chemotherapy sensitivity- and presents a generalized machine- learning tool for precision immunotherapy stratification. The developed online resource is freely available to facilitate community-wide biomarker discovery.

PMID:41925746 | DOI:10.1007/s00262-026-04374-3

  •  

Integrated spatial transcriptomics and pan-cancer XGBoost modeling uncover spatial drivers of immune exclusion and predict immunotherapy response

Cancer Immunol Immunother. 2026 Apr 2;75(4):131. doi: 10.1007/s00262-026-04374-3.

ABSTRACT

Immunotherapy has revolutionized cancer treatment, yet characterizing the spatial complexity of the tumor immune microenvironment remains a challenge. In this study, we established a comprehensive computational framework integrating multi-omics profiling across 27 cancer types to decode immune-related non-coding RNA regulatory networks. Moving beyond traditional bulk analysis, we utilized spatial transcriptomics to dissect the spatial localization of these regulators. We identified the SNHG6-BIRC5 axis as a critical driver of the "immune-cold" phenotype in lung adenocarcinoma. We provide visual evidence that this axis localizes to tumor nests and negatively correlates with T- cell infiltration, elucidating a mechanism of spatial immune exclusion. Validating the clinical relevance of these findings, genome-scale CRISPR-Cas9 screening data confirmed the functional essentiality of these targets for cancer cell survival. Furthermore, pharmacogenomic analysis revealed that high expression of this axis correlates with sensitivity to chemotherapy agents like Vinblastine, suggesting a potential stratification strategy for patients with immune-excluded tumors. To expand the clinical utility to immunotherapy prediction, we developed a pan-cancer XGBoost machine learning model incorporating 14 high-performance regulatory features. This model achieved robust performance in distinguishing immunotherapy responders from non-responders with an AUC of 0.771, outperforming traditional markers such as PD-L1. Collectively, this study highlights spatial determinants of immune exclusion and chemotherapy sensitivity- and presents a generalized machine- learning tool for precision immunotherapy stratification. The developed online resource is freely available to facilitate community-wide biomarker discovery.

PMID:41925746 | PMC:PMC13046951 | DOI:10.1007/s00262-026-04374-3

  •  

Excessive pyroptosis mediates the exacerbation of pneumonia caused by low-lethality influenza virus and secondary MRSA co-infection

Cell Death Discovery, Published online: 02 April 2026; doi:10.1038/s41420-026-03031-z

Excessive pyroptosis mediates the exacerbation of pneumonia caused by low-lethality influenza virus and secondary MRSA co-infection
  •  

PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent

arXiv:2603.29318v1 Announce Type: new Abstract: Smartphone GUI agents execute tasks by operating directly on app interfaces, offering a path to broad capability without deep system integration. However, real-world smartphone use is highly personalized: users adopt diverse workflows and preferences, challenging agents to deliver customized assistance rather than generic solutions. Existing GUI agent benchmarks cannot adequately capture this personalization dimension due to sparse user-specific data and the lack of fine-grained evaluation metrics. To address this gap, we present PSPA-Bench, the benchmark dedicated to evaluating personalization in smartphone GUI agents. PSPA-Bench comprises over 12,855 personalized instructions aligned with real-world user behaviors across 10 representative daily-use scenarios and 22 mobile apps, and introduces a structure-aware process evaluation method that measures agents' personalized capabilities at a fine-grained level. Through PSPA-Bench, we benchmark 11 state-of-the-art GUI agents. Results reveal that current methods perform poorly under personalized settings, with even the strongest agent achieving limited success. Our analysis further highlights three directions for advancing personalized GUI agents: (1) reasoning-oriented models consistently outperform general LLMs, (2) perception remains a simple yet critical capability, and (3) reflection and long-term memory mechanisms are key to improving adaptation. Together, these findings establish PSPA-Bench as a foundation for systematic study and future progress in personalized GUI agents.
  •  

Owl-AuraID 1.0: An Intelligent System for Autonomous Scientific Instrumentation and Scientific Data Analysis

arXiv:2603.29828v1 Announce Type: new Abstract: Scientific discovery increasingly depends on high-throughput characterization, yet automation is hindered by proprietary GUIs and the limited generalizability of existing API-based systems. We present Owl-AuraID, a software-hardware collaborative embodied agent system that adopts a GUI-native paradigm to operate instruments through the same interfaces as human experts. Its skill-centric framework integrates Type-1 (GUI operation) and Type-2 (data analysis) skills into end-to-end workflows, connecting physical sample handling with scientific interpretation. Owl-AuraID demonstrates broad coverage across ten categories of precision instruments and diverse workflows, including multimodal spectral analysis, microscopic imaging, and crystallographic analysis, supporting modalities such as FTIR, NMR, AFM, and TGA. Overall, Owl-AuraID provides a practical, extensible foundation for autonomous laboratories and illustrates a path toward evolving laboratory intelligence through reusable operational and analytical skills. The code are available at https://github.com/OpenOwlab/AuraID.
  •  

A Semi-amortized Lifted Learning-to-Optimize Masked (SALLO-M) Transformer Model for Scalable and Generalizable Beamforming

arXiv:2510.13077v3 Announce Type: replace-cross Abstract: We develop an unsupervised deep learning framework for real-time scalable and generalizable downlink beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed semi-amortized lifted learning-to-optimize (SALLO) framework employs a multi-layer Transformer to iteratively refine an auxiliary variable and the beamformer solution, with a few projected gradient ascent steps at each layer. A key feature of our SALLO Transformer model is that it can handle varying numbers of users and antennas, enabled by a user-antenna dual tokenization and a structured sample/attention masking scheme, leading to generalization across different configurations without retraining. To improve convergence and robustness, we introduce three training strategies: (a) sliding-window training to stabilize gradient propagation, (b) curriculum learning with random masking to enable user-antenna configuration generalization and prevent poor early-stage convergence, and (c) sample replay to mitigate catastrophic forgetting during multi-stage training. Ablation studies validate several key architecture designs and show that the enhanced training scheme improves both generalizability and solution quality. Simulation results over both Gaussian and sparse channels show that the proposed scheme consistently outperforms existing deep learning baselines across diverse system configurations and channel conditions. The performance gain becomes more pronounced in overloaded regimes, highlighting improved robustness under challenging scenarios. Furthermore, our scheme surpasses the WMMSE benchmark in underloaded systems and even in overloaded systems when the overloading factor is below certain threshold. These gains are achieved with fast inference and a substantially more lightweight model than wireless foundation models.
  •  

Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma

Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.

ABSTRACT

Intratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8% of tumors by existing subtyping systems. To overcome this, we identify a low-intratumor-heterogeneity/high-intertumor-variability (LIHV) gene set and develop an ITH-insensitive classification system defining five subgroups: inflammatory (SI), metabolic (SII), atypical (SIII-1), immune-silent (SIII-2), and neurodegenerative (SIII-3). These subgroups exhibit distinct clinical outcomes, molecular features, immune landscapes, and therapeutic vulnerabilities. GPRC5A and VTCN1 serve as robust immunohistochemical biomarkers for SI and SIII tumors, while serum CEA and CA19-9 identify inflammatory iCCA. Therapeutically, HSP90 inhibition synergizes with anti-PD1 in inflammatory iCCA, whereas combined anti-PD1 and anti-TIM3 suppresses neurodegenerative iCCA. Collectively, our study provides a robust molecular framework and actionable therapeutic strategies for iCCA.

PMID:41916296 | DOI:10.1016/j.xcrm.2026.102708

  •  

Liquid Biopsies in HNSCC: Current Landscape and Emerging Opportunities in the Era of HPV Stratification

Int J Mol Sci. 2026 Mar 20;27(6):2847. doi: 10.3390/ijms27062847.

ABSTRACT

Head and neck squamous cell carcinoma (HNSCC) is biologically and clinically dichotomous according to HPV status, a distinction that fundamentally dictates the design, implementation, and interpretation of liquid biopsy strategies. Conventional anatomical imaging lacks sufficient sensitivity for minimal residual disease (MRD) detection, contributing significantly to treatment failure and suboptimal clinical outcomes. This review provides a critical, evidence-based synthesis of the three principal circulating analytes, circulating tumor DNA (ctDNA), exosomes, and circulating tumor cells (CTCs), and their evolving roles in real-time, non-invasive molecular monitoring. Critically, the clinical readiness of these analytes differs substantially: while ctDNA, particularly HPV-related ctDNA, is approaching clinical validation for MRD detection and recurrence surveillance in HPV-positive HNSCC, exosomes and CTCs remain investigational tools hindered by ongoing technical challenges including lack of standardized assays, limited reproducibility across platforms, and insufficient prospective validation. We review how the presence of a clonal, virally derived DNA target in HPV-positive HNSCC contrasts with the heterogeneous somatic mutational landscape of HPV-negative tumors, necessitating divergent analytical platforms and yielding distinct clinical utility profiles for MRD detection and recurrence surveillance. We further outline a pragmatic translational pathway focused on assay standardization, particularly for exosomes and CTCs where this foundational work is most urgently needed, integration of complementary multimodal liquid biopsy approaches, and rigorously designed prospective interventional clinical trials to establish clinical utility. Collectively, these efforts aim to transition HNSCC management from reactive, anatomy-based surveillance to proactive, molecularly guided precision oncology, with the potential to improve therapeutic decision-making and patient outcomes.

PMID:41898706 | PMC:PMC13027142 | DOI:10.3390/ijms27062847

  •  

Advances in Spatial Multi-Omics in Gastric Cancer

Cells. 2026 Mar 17;15(6):535. doi: 10.3390/cells15060535.

ABSTRACT

Gastric cancer (GC) remains a major global health burden, with its unfavorable prognosis primarily driven by extensive tumor heterogeneity. Traditional bulk omics, while informative, are inherently limited by the averaging effect of diverse cell populations and fail to capture the critical spatial molecular disparities within the tumor and its microenvironment (TME). Single-cell omics can capture cellular heterogeneity but lack spatial context. Therefore, there is an urgent clinical need for spatial multi-omics to provide a high-definition dissection of GC heterogeneity and to optimize therapeutic efficacy. This review first outlines briefly the evolution of spatial technologies, including transcriptomics, proteomics, metabolomics, genomics and epigenomics, and their transformative applications in GC research. We further explore how these platforms refine molecular classification beyond traditional models, identify next-generation biomarkers, and decode the intricate cellular interactions governing immune evasion and metastasis. Next, we highlight the pivotal role of spatial profiling in unravelling the multidimensional mechanisms of resistance to chemotherapy, targeted therapy and immunotherapy. Finally, we address current technical bottlenecks and discuss prospects for clinical translation.

PMID:41892326 | PMC:PMC13025482 | DOI:10.3390/cells15060535

  •  

Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

arXiv:2602.22983v3 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying across language contexts. This paper investigates the role of classical Chinese in jailbreak attacks. Owing to its conciseness and obscurity, classical Chinese can partially bypass existing safety constraints, exposing notable vulnerabilities in LLMs. Based on this observation, this paper proposes a framework, CC-BOS, for the automatic generation of classical Chinese adversarial prompts based on multi-dimensional fruit fly optimization, facilitating efficient and automated jailbreak attacks in black-box settings. Prompts are encoded into eight policy dimensions-covering role, behavior, mechanism, metaphor, expression, knowledge, trigger pattern and context; and iteratively refined via smell search, visual search, and cauchy mutation. This design enables efficient exploration of the search space, thereby enhancing the effectiveness of black-box jailbreak attacks. To enhance readability and evaluation accuracy, we further design a classical Chinese to English translation module. Extensive experiments demonstrate that effectiveness of the proposed CC-BOS, consistently outperforming state-of-the-art jailbreak attack methods.
  •  

From Context to Intent: Reasoning-Guided Function-Level Code Completion

arXiv:2508.09537v2 Announce Type: replace-cross Abstract: The growing capabilities of Large Language Models (LLMs) have led to their widespread adoption for function completion within code repositories. Recent studies on such tasks show promising results when explicit instructions, often in the form of docstrings, are available to guide the completion. However, in real-world scenarios, clear docstrings are frequently absent. Under such conditions, LLMs typically fail to produce accurate completions. To enable more automated and accurate function completion in such settings, we aim to enable LLMs to accurately infer the developer's intent prior to code completion. Our key insight is that the preceding code, namely the code context before the function to be completed, often contains valuable cues that help the model understand the intended functionality. However, inferring intent from such implicit context is non-trivial and constitutes a core challenge in function-level code completion. To tackle this challenge, inspired by how humans interpret context, we propose a reasoning-based prompting framework that guides LLMs to utilize these contextual cues to infer intent step by step. To incentivize LLMs to reason through the preceding code and infer intent, we further curate a dataset of 40k examples, each annotated with intermediate reasoning traces and corresponding docstrings. Extensive experiments on DevEval and ComplexCodeEval demonstrate consistent performance improvements across multiple models, achieving over 25% relative gains in pass@1 for both DeepSeekCoder and CodeLLaMA families. Building upon our framework, we further develop an intent-interactive platform that supports lightweight human feedback. This platform allows developers to select from a set of candidate intentions or edit the intent to better guide the model. Our experiments show that this interactive approach leads to further performance improvements.
  •  

Schr\"odinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation

arXiv:2512.21201v2 Announce Type: replace-cross Abstract: Zero-shot object navigation (ZSON) requires robots to locate target objects in unseen environments without task-specific fine-tuning or pre-built maps, a capability crucial for service and household robotics. Existing methods perform well in simulation but struggle in realistic, cluttered environments where heavy occlusions and latent hazards make large portions of the scene unobserved. These approaches typically act on a single inferred scene, making them prone to overcommitment and unsafe behavior under uncertainty. To address these challenges, we propose Schr\"odinger's Navigator, a belief-aware framework that explicitly reasons over multiple trajectory-conditioned imagined 3D futures at inference time. A trajectory-conditioned 3D world model generates hypothetical observations along candidate paths, maintaining a superposition of plausible scene realizations. An adaptive, occluder-aware trajectory sampling strategy focuses imagination on uncertain regions, while a Future-Aware Value Map (FAVM) aggregates imagined futures to guide robust, proactive action selection. Evaluations in simulation and on a physical Go2 quadruped robot demonstrate that Schr\"odinger's Navigator outperforms strong ZSON baselines, achieving more robust self-localization, object localization, and safe navigation under severe occlusions and latent hazards. These results highlight the effectiveness of reasoning over imagined 3D futures as a scalable and generalizable strategy for zero-shot navigation in uncertain real-world environments.
  •  

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

arXiv:2603.07980v1 Announce Type: cross Abstract: As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench \$OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, \$OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.
  •  

Process-Centric Analysis of Agentic Software Systems

arXiv:2512.02393v2 Announce Type: replace-cross Abstract: Agentic systems are modern software systems: they consist of orchestrated modules, expose interfaces, and are deployed in software pipelines. Unlike conventional programs, their execution, i.e., trajectories, is inherently stochastic and adaptive to the problems they solve. Evaluation of such systems is often outcome-centric. This narrow focus overlooks detailed insights, failing to explain how agents reason, plan, act, or change their strategies. Inspired by the structured representation of conventional software systems as graphs, we introduce Graphectory to systematically encode the temporal and semantic relations in such systems. Using Graphectory, we automatically analyze 4000 trajectories of two dominant agentic programming workflows, SWE-agent and OpenHands, with four backbone Large Language Models (LLMs), attempting to resolve SWE-bench issues. Our automated analyses (completed within four minutes) reveal that: (1) agents using richer prompts or stronger LLMs exhibit more complex Graphectory, reflecting deeper exploration, broader context gathering, and more thorough validation; (2) agents' strategies vary with problem difficulty and the underlying LLM - for resolved issues, strategies often follow coherent localization-patching-validation steps, while unresolved ones exhibit chaotic or backtracking behaviors; and (3) even successful agentic systems often display inefficient processes. We also implement a novel technique for real-time construction and analysis of Graphectory and Langutory during agent execution to flag trajectory issues. Upon detecting such issues, the technique notifies the agent with a diagnostic message and, when applicable, rolls back the trajectory. Experiments show that online monitoring and interventions improve resolution rates by 6.9%-23.5% across models for problematic instances, while significantly shortening trajectories with near-zero overhead.
  •  

GIPO: Gaussian Importance Sampling Policy Optimization

arXiv:2603.03955v1 Announce Type: cross Abstract: Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation. However, RL remains limited by poor data efficiency, particularly in settings where interaction data are scarce and quickly become outdated. To address this challenge, GIPO (Gaussian Importance sampling Policy Optimization) is proposed as a policy optimization objective based on truncated importance sampling, replacing hard clipping with a log-ratio-based Gaussian trust weight to softly damp extreme importance ratios while maintaining non-zero gradients. Theoretical analysis shows that GIPO introduces an implicit, tunable constraint on the update magnitude, while concentration bounds guarantee robustness and stability under finite-sample estimation. Experimental results show that GIPO achieves state-of-the-art performance among clipping-based baselines across a wide range of replay buffer sizes, from near on-policy to highly stale data, while exhibiting superior bias--variance trade-off, high training stability and improved sample efficiency.
  •  

ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools

arXiv:2508.03284v2 Announce Type: replace Abstract: Integrating external tools into Large Foundation Models (LFMs) has emerged as a promising approach to enhance their problem-solving capabilities. While existing studies have demonstrated strong performance in tool-augmented Visual Question Answering (VQA), recent benchmarks reveal significant gaps in real-world tool-use proficiency, particularly in functionally diverse multimodal settings requiring multi-step reasoning. In this work, we introduce ToolVQA, a large-scale multimodal dataset comprising 23K instances, designed to bridge this gap. Unlike previous datasets that rely on synthetic scenarios and simplified queries, ToolVQA features real-world visual contexts and challenging implicit multi-step reasoning tasks, better aligning with real user interactions. To construct this dataset, we propose ToolEngine, a novel data generation pipeline that employs Depth-First Search (DFS) with a dynamic in-context example matching mechanism to simulate human-like tool-use reasoning. ToolVQA encompasses 10 multimodal tools across 7 diverse task domains, with an average inference length of 2.78 reasoning steps per instance. The fine-tuned 7B LFMs on ToolVQA not only achieve impressive performance on our test set but also surpass the large close-sourced model GPT-3.5-turbo on various out-of-distribution (OOD) datasets, demonstrating strong generalizability to real-world tool-use scenarios.
  •  

Rigidity-Aware Geometric Pretraining for Protein Design and Conformational Ensembles

arXiv:2603.02406v1 Announce Type: cross Abstract: Generative models have recently advanced $\textit{de novo}$ protein design by learning the statistical regularities of natural structures. However, current approaches face three key limitations: (1) Existing methods cannot jointly learn protein geometry and design tasks, where pretraining can be a solution; (2) Current pretraining methods mostly rely on local, non-rigid atomic representations for property prediction downstream tasks, limiting global geometric understanding for protein generation tasks; and (3) Existing approaches have yet to effectively model the rich dynamic and conformational information of protein structures. To overcome these issues, we introduce $\textbf{RigidSSL}$ ($\textit{Rigidity-Aware Self-Supervised Learning}$), a geometric pretraining framework that front-loads geometry learning prior to generative finetuning. Phase I (RigidSSL-Perturb) learns geometric priors from 432K structures from the AlphaFold Protein Structure Database with simulated perturbations. Phase II (RigidSSL-MD) refines these representations on 1.3K molecular dynamics trajectories to capture physically realistic transitions. Underpinning both phases is a bi-directional, rigidity-aware flow matching objective that jointly optimizes translational and rotational dynamics to maximize mutual information between conformations. Empirically, RigidSSL variants improve designability by up to 43\% while enhancing novelty and diversity in unconditional generation. Furthermore, RigidSSL-Perturb improves the success rate by 5.8\% in zero-shot motif scaffolding and RigidSSL-MD captures more biophysically realistic conformational ensembles in G protein-coupled receptor modeling. The code is available at: https://github.com/ZhanghanNi/RigidSSL.git.
  •  

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

arXiv:2507.01352v3 Announce Type: replace-cross Abstract: Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced human preferences. We hypothesize that this brittleness stems primarily from limitations in preference datasets, which are often narrowly scoped, synthetically labeled, or lack rigorous quality control. To address these challenges, we present SynPref-40M, a large-scale preference dataset comprising 40 million preference pairs. To enable data curation at scale, we design a human-AI synergistic two-stage pipeline that leverages the complementary strengths of human annotation quality and AI scalability. In this pipeline, humans provide verified annotations, while LLMs perform automatic curation based on human guidance. Training on this preference mixture, we introduce Skywork-Reward-V2, a suite of eight reward models ranging from 0.6B to 8B parameters, trained on a carefully curated subset of 26 million preference pairs from SynPref-40M. We demonstrate that Skywork-Reward-V2 is versatile across a wide range of capabilities, including alignment with human preferences, objective correctness, safety, resistance to stylistic biases, and best-of-N scaling. These reward models achieve state-of-the-art performance across seven major reward model benchmarks, outperform generative reward models, and demonstrate strong downstream performance. Ablation studies confirm that effectiveness stems not only from data scale but also from high-quality curation. The Skywork-Reward-V2 series represents substantial progress in open reward models, demonstrating how human-AI curation synergy can unlock significantly higher data quality.
  •  

xLLM Technical Report

arXiv:2510.14686v2 Announce Type: replace-cross Abstract: We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locates online and offline tasks through unified elastic scheduling to maximize cluster utilization. This module also relies on a workload-adaptive dynamic Prefill-Decode (PD) disaggregation policy and a novel Encode-Prefill-Decode (EPD) disaggregation policy designed for multimodal inputs. Furthermore, it incorporates a distributed architecture to provide global KV Cache management and robust fault-tolerant capabilities for high availability. At the engine layer, xLLM-Engine co-optimizes system and algorithm designs to fully saturate computing resources. This is achieved through comprehensive multi-layer execution pipeline optimizations, an adaptive graph mode and an xTensor memory management. xLLM-Engine also further integrates algorithmic enhancements such as optimized speculative decoding and dynamic EPLB, collectively serving to substantially boost throughput and inference efficiency. Extensive evaluations demonstrate that xLLM delivers significantly superior performance and resource efficiency. Under identical TPOT constraints, xLLM achieves throughput up to 1.7x that of MindIE and 2.2x that of vLLM-Ascend with Qwen-series models, while maintaining an average throughput of 1.7x that of MindIE with Deepseek-series models. xLLM framework is publicly available at https://github.com/jd-opensource/xllm and https://github.com/jd-opensource/xllm-service.
  •  
❌