❌

Normal view

GraphAHA: Graph-Based Adaptive Search with Heterogeneous Actions for Test-Time Code Generation

arXiv:2609.12757v1 Announce Type: cross Abstract: Test-time scaling improves code generation by spending additional inference budget (e.g., calls or tokens) on direct sampling, feedback-conditioned repair, and reasoning-guided implementation. Search-based methods can allocate this budget adaptively, but two challenges remain. First, tree-structured search treats each generation history as a separate state even when trajectories converge to the same program, duplicating evaluation and preventing statistics from being shared. Second, sampling, repair, and reasoning have complementary and state-dependent payoffs, making online allocation among them difficult under a finite budget. To address these challenges, we propose an adaptive graph search method with heterogeneous actions (GraphAHA). GraphAHA organizes the test-time code generation in a typed directed acyclic graph. Equivalent programs are merged into a single code node, allowing their downstream search statistics to be reused across all discovery paths. Hierarchical Thompson sampling then selects whether to generate a new state or follow an existing successor and, for generation, chooses among the type-valid sampling, reasoning, implementation, and repair operations. Evaluated on LiveCodeBench and CodeContests with Qwen2.5-Coder and DeepSeek-Coder, GraphAHA achieves the best score in 18 of 20 cases. For Pass@1 measured using visible tests, it outperforms the strongest baseline for both models on both benchmarks by 4.1 percentage points on average, demonstrating more effective use of a fixed inference budget.

NBR1-Mediated Autophagic Degradation of YTHDF1 Curtails <em>FDX1</em> Translation to Drive Concurrent Multikinase Inhibitor Resistance and Cuproptosis Tolerance

Cancer Commun (Lond). 2026 Sep 11;46:0048. doi: 10.34133/cancomm.0048. eCollection 2026.

ABSTRACT

Background: Cancer cells frequently acquire adaptive resistance to targeted therapies; however, strategies capable of concurrently overcoming treatment tolerance and reactivating cell death pathways are currently lacking. Here, we investigated the dual role of ferredoxin 1 (FDX1) in modulating both multikinase inhibitor (MKI) sensitivity and cuproptosis susceptibility in hepatocellular carcinoma (HCC), and sought to develop a therapeutic approach for reversing resistance. Methods: HCC models, both in vitro and in vivo, were employed to investigate the role of FDX1 in MKI resistance and cuproptosis evasion. Polysome profiling, SunTag translation reporters, CRISPR-Cas9 mutagenesis, and mass spectrometry were employed to delineate the underlying mechanisms. A codelivery nanoliposome system was engineered and tested in orthotopic HCC models. Results: Prolonged exposure to MKIs led to the down-regulation of FDX1 protein levels, resulting in MKI resistance and cuproptosis tolerance in HCC both in vitro and in vivo. Mechanistically, we found that MKIs inactivated protein kinase B (PKB, also known as AKT)-mechanistic target of rapamycin (mTOR) signaling, thereby suppressing the SET and MYND domain-containing protein 2 (SMYD2)-mediated methylation of YTH domain family protein 1 (YTHDF1) at lysine 515 (K515). Hypomethylated YTHDF1 was degraded via next to BRCA1 gene 1 protein (NBR1)-dependent autophagy, leading to the repression of N6-methyladenosine modification-dependent translation of FDX1 mRNA. FDX1 deficiency drove MKI resistance by reactivating AKT survival signaling while impairing cuproptosis through reduced divalent copper ions (Cu2+) to monovalent copper ions (Cu+) conversion and the loss of protein lipoylation. Additionally, restoring FDX1 expression through NBR1 knockdown or YTHDF1 overexpression overcame MKI resistance and resensitized HCC cells to cuproptosis. Finally, a nanoliposomal system, super cuproptosis detonator liposome, designed for the codelivery of NBR1 small interfering RNA, a copper ionophore, and sorafenib restored FDX1-dependent cuproptosis and exhibited marked anti-HCC efficacy, suppressing HCC growth in vivo. Conclusions: MKIs suppressed SMYD2-mediated YTHDF1 methylation at K515 via the inactivation of AKT-mTOR signaling. This led to the inhibition of FDX1 translation, resulting in AKT signaling reactivation and protein lipoylation impairment, effects that contributed to both MKI resistance and cuproptosis tolerance in HCC. Overcoming MKI resistance and resensitizing cells to cuproptosis by targeting NBR1-mediated YTHDF1 degradation using a nanoliposomal codelivery system represents a promising strategy for HCC treatment.

PMID:42729649 | PMC:PMC13562797 | DOI:10.34133/cancomm.0048

Gut dysbiosis, metabolic signals, and pulmonary immune reprogramming: decoding the gut microbiota -immune axis in stroke-associated pneumonia

Front Immunol. 2026 Aug 27;17:1812306. doi: 10.3389/fimmu.2026.1812306. eCollection 2026.

ABSTRACT

Stroke-associated pneumonia (SAP) is the most common infectious complication following acute stroke. The limited efficacy of conventional antimicrobial therapy suggests that SAP may be fundamentally a syndrome driven by dysregulated cross-system interactions. This review proposes the "gut microbiota-immune axis" (GMIA) as a comprehensive framework for the development of SAP and systematically discusses the potential mechanisms by which post-stroke microbial-derived metabolic signals-including short-chain fatty acids (SCFAs), bile acids, tryptophan metabolites, and endotoxins-drive systemic immune reprogramming, predisposing patients to SAP. Based on the GMIA, we highlight several promising intervention strategies, including dietary modulation, precision antibiotic use, probiotics, fecal microbiota transplantation (FMT), supplementation with microbial metabolites, and receptor-targeted therapies, and summarize the current clinical translation related to the GMIA. Future research directions require high-quality clinical trials that integrate multi-omics data from the microbiome with immune biomarkers and clinical parameters. Such an approach is essential for constructing validated risk stratification models and advancing the management of SAP from empirical anti-infective treatment toward a precision medicine model centered on GMIA-based immune modulation.

PMID:42724580 | PMC:PMC13560329 | DOI:10.3389/fimmu.2026.1812306

Engineering inflammation-responsive proteins through nitric oxide-caged amino acids

Nature Biomedical Engineering, Published online: 31 August 2026; doi:10.1038/s41551-026-01782-9

A protein engineering strategy enables nitric oxide-triggered reactivation of proteins using genetically encoded caged amino acids, allowing inflammation-localized control of protein activity, viral gene delivery and biosensing in vivo.

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

arXiv:2609.09905v1 Announce Type: cross Abstract: Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on fixed preference pairs rely primarily on positive-only fine-tuning or DPO-style likelihood-ratio surrogates. We organize these approaches through a divergence-based framework and introduce FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts. For linear interpolation, we show under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data. We further show that this loss is nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In the in-domain setting, FlowCPO achieves higher mean GenEval and OCR scores than the evaluated baselines, reaching 0.84 and 0.87 versus 0.81 and 0.74 for FlowDPO at CFG 3.0. In the out-of-domain setting, the results are mixed, with the best GenEval result but lower reward scores than RFT on several metrics.

Ancient proteins identify various Denisovan remains from Southwest China

Nature, Published online: 09 September 2026; doi:10.1038/s41586-026-10976-9

Identification and proteomic analysis of bone fragments and teeth from an excavation in Southwest China provide insight into the evolution and phenotype of Denisovans and fill a geographical gap in their documented distribution.

Integrated single-cell multi-omics characterization reveals lipid-associated macrophage-mediated immunosuppression in neoadjuvant immunotherapy of hepatocellular carcinoma

Nat Commun. 2026 Jul 31;17(1):9381. doi: 10.1038/s41467-026-75949-y.

ABSTRACT

Hepatocellular carcinoma (HCC) is a cancer with high incidence and mortality rate. Although immune checkpoint inhibitors (ICIs) improved survival outcomes for HCC patients, limited objective response rate highlights the urgency of investigating determinants of immunotherapy. Here, we explore HCC resistance mechanisms following neoadjuvant αPD-1 immunotherapy by constructing a comprehensive multi-modal single-cell transcriptomic atlas consisting of 14 HCC patients treated with αPD-1 from our cohort (ClinicalTrials.gov ID: NCT06571396) and 60 external HCC cases with heterogeneous treatment backgrounds. Supervised by clinical outcomes of our cohort, we identify positive and negative regulators of immunotherapy within the tumor immune microenvironment (TIME), especially lipid-associated macrophages (LAM) with increased lipid metabolic state in non-responders and characterized by C1QA, FABP1, and APOA1 expression. We further show the presence, exogenous inducements and immunosuppressive functions of LAM, along with regulation strategies of its lipid-associated condition, including lycopene and chiglitazar. Furthermore, we construct interaction networks of immune regulators across responders and non-responders, showing distinct ligand-receptor landscapes with intervention targets. We reveal the TIME components including immunosuppressive LAMs that influence immunotherapy outcomes, thus providing evidence and insights for exploring immune landscape and therapeutic strategies for HCC immunotherapy. ClinicalTrials.gov ID: NCT06571396.

PMID:42680737 | PMC:PMC13534469 | DOI:10.1038/s41467-026-75949-y

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented autonomous execution capabilities for LLM Agents by integrating external open-domain knowledge and tools. However, this interoperability introduces a covert attack surface targeting the agent's cognitive planning layer. This paper systematically investigates Tool Description Poisoning (TDP), a novel semantic attack. In TDP, malicious instructions are not embedded in a tool's executable code, but rather covertly injected into its descriptive metadata, the very "manual" an agent relies on for secure planning and decision-making. To rigorously and systematically evaluate this emerging threat, we introduce the MCP-TDP Security Benchmark. This high-fidelity sandbox environment comprises 32 realistic, real-world test cases spanning 6 distinct risk categories. Our evaluation of 8 mainstream LLMs reveals severe vulnerabilities, with leading models like GPT-4o exhibiting a nearly 100% Attack Success Rate (ASR) in six high-risk scenarios. Furthermore, our findings demonstrate that common prompt-guardrail defenses are largely ineffective and can, counterintuitively, even be counterproductive (a phenomenon which we term the "Firewall Fallacy"). Crucially, we also propose a defense mechanism: "Reactive Self-Correction," where an agent autonomously detects and reverts its own malicious actions post-execution. This work provides the first specialized security benchmark tailored for TDP, offering essential insights for securing the cognitive and planning layers of advanced agentic systems.

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parameter updates, and reusable components opens entry points for attackers. Threats have evolved from data poisoning and weight tampering to agent manipulation and interface exploitation, yet existing reviews lack a unified framework spanning the full fine-tuning lifecycle. Objective: This paper presents a systematic survey of LLM fine-tuning security and establishes a lifecycle-based framework for comparing attacks and defenses, complemented by unified empirical evaluation. Methods: We divide attack and defense mechanisms into three phases by intervention timing: pre-tuning, during-tuning, and post-tuning. Within each phase, strategies are reviewed and contrasted to expose their evolution and limitations. Representative methods are then evaluated under a unified model, hardware, and protocol setup, with cross-phase experiments pairing attacks and defenses from different phases. Results: Attack effectiveness is highly model-dependent and non-monotonic with scale: weight-editing attacks effective on earlier models lose impact on modern open-source LLMs; cross-lingual backdoor transfer, reported as near-perfect at larger scales, fails entirely on tested 1B-4B models; and purely benign samples can compromise safety alignment in instruction-tuned models. Single-phase defenses rarely generalize across phases, and defense effectiveness depends jointly on model architecture and alignment state. Conclusion: We identify key open problems (configuration-robust defense, cross-phase defense composition, and embedding-space attacks beyond behavioral assumptions) and propose concrete future research directions.

Hide to Guide: Learning via Semantic Masking

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limited by exploration. For example, models often fail on hard problems, leaving little useful reward signal. External expert traces offer a natural source of guidance, yet they may also expose reward-relevant content along the critical path to the verifier target, such as final answers, intermediate values, executable implementations, or answer-related entities. This content can create an unintended reward hacking channel, allowing the policy to obtain reward by copying the trace rather than learning the underlying reasoning or agentic behavior. Existing guided-RL methods reduce this risk by using partial trajectories, but they mainly control how much expert information is shown heuristically rather than which parts should be hidden. To this end, we propose Semantic Masked Expert Policy Optimization (SMEPO), a fine-grained semantic masking strategy for expert-guided RLVR. Instead of truncating traces coarsely or revealing them unchanged, SMEPO masks reward-relevant semantic spans along the critical path while preserving the expert's decomposition, plan, and procedural structure. This turns hard problems from reasoning from scratch into a fill-in-the-blank process: the policy can follow the expert's problem-solving route, but must still reconstruct the missing values, code, or entities by itself. SMEPO is simple to apply and requires no changes to the reward function or RL objective. Across diverse domains, including math, code, and agentic search, SMEPO improves accuracy by up to 3.2 points over GRPO and reduces training time by up to 4.2x. The code is available at https://github.com/mit-han-lab/SMEPO.

NSR-Boost: A Neuro-Symbolic Residual Boosting Framework for Industrial Legacy Models

arXiv:2601.10457v3 Announce Type: replace Abstract: Although the Gradient Boosted Decision Trees (GBDTs) dominate industrial tabular applications, upgrading legacy models in high-concurrency production environments still faces prohibitive retraining costs and systemic risks. To address this problem, we present NSR-Boost, a neuro-symbolic residual boosting framework designed specifically for industrial scenarios. Its core advantage lies in being ``non-intrusive''. It treats the legacy model as a frozen model and performs targeted repairs on "hard regions" where predictions fail. The framework comprises three key stages: First, finding hard regions through residuals, then generating interpretable experts by generating symbolic code structures using Large Language Model (LLM) and fine-tuning parameters using Bayesian optimization, and finally dynamically integrating experts with legacy model output through a lightweight aggregator. Experimental results demonstrate that the framework significantly outperforms state-of-the-art (SOTA) baselines across six public datasets and one private dataset. More importantly, we report the successful deployment of NSR-Boost within the core financial risk control system of Qfin Holdings, where empirical results on real-world online traffic exhibit superior performance improvements and a significant reduction in the bad rate. In conclusion, it effectively captures long-tail risks missed by traditional models and offers a safe, low-cost evolutionary paradigm for industry.

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

arXiv:2511.16449v5 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployment, these models must process continuous visual streams, incurring substantial computational overhead. Visual token pruning -- a mainstream technique for accelerating Vision-Language Models (VLMs) by retaining salient tokens while discarding redundant ones -- offers a natural candidate solution to this challenge. However, directly applying VLM-oriented pruning methods to VLA inference can cause severe degradation in manipulation performance. Our analysis attributes this degradation to a key mismatch: VLA inference exhibits distinct attention patterns between the vision-language prefill stage and the action-decode stage, so pruning based only on context-prefill semantic salience is biased toward semantic cues and may remove action-critical visual tokens. Motivated by this observation, we propose VLA-Pruner, an effective plug-and-play token pruning method grounded in the visual requirements of VLA inference, further exploiting the temporal continuity of robot manipulation. Specifically, VLA-Pruner estimates visual-token importance from both semantic prefilling and temporally smoothed action relevance, and then applies a Combine-then-Filter strategy to retain compact, non-redundant tokens under the compute budget. Experiments show that VLA-Pruner outperforms state-of-the-art approaches across multiple VLA architectures, achieving up to 1.99x speedup with comparable manipulation quality.

A pathogen lncRNA secreted into rice sequesters a host miRNA for virulence

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10572-x

A fungal long non-coding RNA from Magnaporthe oryzae translocates into rice cells to sequester a host microRNA that normally represses PKR1, a negative immunity regulator, thereby facilitating infection and revealing a widespread RNA-based pathogen–host interaction mechanism.

Effect of the Maxing Huoqiao granule on nonsevere community-acquired pneumonia: A multicenter, double-blind, placebo-controlled randomized trial

Pharmacol Res. 2026 Apr 9:108186. doi: 10.1016/j.phrs.2026.108186. Online ahead of print.

ABSTRACT

Community-acquired pneumonia (CAP) remains a major global public health challenge with substantial morbidity and mortality. Although preclinical studies suggest that Maxing Huoqiao (MXHQ) granule may have therapeutic potential for pneumonia, high-quality clinical evidence is still limited. We conducted a multicenter, double-blind, randomized, placebo-controlled trial at two tertiary hospitals in China to evaluate the clinical efficacy of MXHQ as adjunctive therapy and to explore its potential mechanisms in adults with nonsevere CAP receiving standard moxifloxacin treatment. A total of 96 patients were enrolled and randomized (1:1:1) to receive standard-dose MXHQ, low-dose MXHQ, or placebo in addition to moxifloxacin for 7 days, with a 14-day follow-up. The primary endpoint was clinical cure, defined as composite recovery of major respiratory symptoms, lung rales, and fever; secondary endpoints included symptom relief, radiographic improvement, and safety. Compared with placebo, standard-dose MXHQ was associated with a higher day-14 clinical cure rate (30.78% vs. 68.97%; RR = 0.45, 95% CI = 0.24-0.83; P < 0.01). Furthermore, the standard-dose intervention was correlated with a shorter time to relief and recovery of cough and sputum (P < 0.05), as well as improvements in symptom scores (P < 0.05) and promoting lesion absorption on chest CT (P < 0.05). Low-dose MXHQ showed no significant clinical benefit, whereas safety profiles were comparable across all groups. Transcriptomic analyses of peripheral blood mononuclear cells, complemented by a Streptococcus pneumonia animal model, indicated that the clinical benefits of MXHQ are linked to the modulation of inflammation and innate immunity. These omics and in vivo observations suggest a potential mechanism underlying the protective effects of MXHQ against inflammatory injury and promotion of tissue repair, involving the regulation of anti-inflammatory mediators and tissue repair-related factors. (Chictr.org.cn, ID Number: ChiCTR2400082095).

PMID:41966499 | DOI:10.1016/j.phrs.2026.108186

Talk to Right Specialists: Iterative Routing in Multi-agent Systems for Question Answering

arXiv:2501.07813v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) agents are increasingly deployed to answer questions over local knowledge bases that cannot be centralized due to knowledge-sovereignty constraints. This results in two recurring failures in production: users do not know which agent to consult, and complex questions require evidence distributed across multiple agents. To overcome these challenges, we propose RIRS, a training-free orchestration framework to enable a multi-agent system for question answering. In detail, RIRS summarizes each agent's local corpus in an embedding space, enabling a user-facing server to route queries only to the most relevant agents, reducing latency and avoiding noisy "broadcast-to-all" contexts. For complicated questions, the server can iteratively aggregate responses to derive intermediate results and refine the question to bridge the gap toward a comprehensive answer. Extensive experiments demonstrate the effectiveness of RIRS, including its ability to precisely select agents and provide accurate responses to single-hop queries, and its use of an iterative strategy to achieve accurate, multi-step resolutions for complex queries.

From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?

arXiv:2604.01608v1 Announce Type: new Abstract: Multi-agent systems (MAS) tackle complex tasks by distributing expertise, though this often comes at the cost of heavy coordination overhead, context fragmentation, and brittle phase ordering. Distilling a MAS into a single-agent skill can bypass these costs, but this conversion lacks a principled answer for when and what to distill. Instead, the empirical outcome is surprisingly inconsistent: skill lift ranges from a 28% improvement to a 2% degradation across metrics of the exact same task. In this work, we reveal that skill utility is governed not by the task, but by the evaluation metric. We introduce Metric Freedom ($F$), the first a priori predictor of skill utility. $F$ measures the topological rigidity of a metric's scoring landscape by quantifying how output diversity couples with score variance via a Mantel test. Guided by $F$, we propose a two-stage adaptive distillation framework. Stage 1 acts as a selective extraction mechanism, extracting tools and knowledge while discarding restrictive structures on "free" metrics to preserve exploration. Stage 2 targets computationally intensive iterative refinement exclusively toward "rigid" metrics ($F \lesssim 0.6$) to eliminate trajectory-local overfitting. Evaluating across 4 tasks, 11 datasets, and 6 metrics, $F$ strongly predicts skill utility ($\rho = -0.62$, $p

SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model

arXiv:2603.26738v2 Announce Type: replace-cross Abstract: While automated sleep staging has achieved expert-level accuracy, its clinical adoption is hindered by a lack of auditable reasoning. We introduce SleepVLM, a rule-grounded vision-language model (VLM) designed to stage sleep from multi-channel polysomnography (PSG) waveform images while generating clinician-readable rationales based on American Academy of Sleep Medicine (AASM) scoring criteria. Utilizing waveform-perceptual pre-training and rule-grounded supervised fine-tuning, SleepVLM achieved Cohen's kappa scores of 0.767 on an held out test set (MASS-SS1) and 0.743 on an external cohort (ZUAMHCS), matching state-of-the-art performance. Expert evaluations further validated the quality of the model's reasoning, with mean scores exceeding 4.0/5.0 for factual accuracy, evidence comprehensiveness, and logical coherence. By coupling competitive performance with transparent, rule-based explanations, SleepVLM may improve the trustworthiness and auditability of automated sleep staging in clinical workflows. To facilitate further research in interpretable sleep medicine, we release MASS-EX, a novel expert-annotated dataset.

Ran Score: a LLM-based Evaluation Score for Radiology Report Generation

arXiv:2603.22935v1 Announce Type: new Abstract: Chest X-ray report generation and automated evaluation are limited by poor recognition of low-prevalence abnormalities and inadequate handling of clinically important language, including negation and ambiguity. We develop a clinician-guided framework combining human expertise and large language models for multi-label finding extraction from free-text chest X-ray reports and use it to define Ran Score, a finding-level metric for report evaluation. Using three non-overlapping MIMIC-CXR-EN cohorts from a public chest X-ray dataset and an independent ChestX-CN validation cohort, we optimize prompts, establish radiologist-derived reference labels and evaluate report generation models. The optimized framework improves the macro-averaged score from 0.753 to 0.956 on the MIMIC-CXR-EN development cohort, exceeds the CheXbert benchmark by 15.7 percentage points on directly comparable labels, and shows robust generalization on the ChestX-CN validation cohort. Here we show that clinician-guided prompt optimization improves agreement with a radiologist-derived reference standard and that Ran Score enables finding-level evaluation of report fidelity, particularly for low-prevalence abnormalities.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

arXiv:2603.06652v1 Announce Type: cross Abstract: Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.

Not All Candidates are Created Equal: A Heterogeneity-Aware Approach to Pre-ranking in Recommender Systems

arXiv:2603.03770v1 Announce Type: cross Abstract: Most large-scale recommender systems follow a multi-stage cascade of retrieval, pre-ranking, ranking, and re-ranking. A key challenge at the pre-ranking stage arises from the heterogeneity of training instances sampled from coarse-grained retrieval results, fine-grained ranking signals, and exposure feedback. Our analysis reveals that prevailing pre-ranking methods, which indiscriminately mix heterogeneous samples, suffer from gradient conflicts: hard samples dominate training while easy ones remain underutilized, leading to suboptimal performance. We further show that the common practice of uniformly scaling model complexity across all samples is inefficient, as it overspends computation on easy cases and slows training without proportional gains. To address these limitations, this paper presents Heterogeneity-Aware Adaptive Pre-ranking (HAP), a unified framework that mitigates gradient conflicts through conflict-sensitive sampling coupled with tailored loss design, while adaptively allocating computational budgets across candidates. Specifically, HAP disentangles easy and hard samples, directing each subset along dedicated optimization paths. Building on this separation, it first applies lightweight models to all candidates for efficient coverage, and further engages stronger models on the hard ones, maintaining accuracy while reducing cost. This approach not only improves pre-ranking effectiveness but also provides a practical perspective on scaling strategies in industrial recommender systems. HAP has been deployed in the Toutiao production system for 9 months, yielding up to 0.4% improvement in user app usage duration and 0.05% in active days, without additional computational cost. We also release a large-scale industrial hybrid-sample dataset to enable the systematic study of source-driven candidate heterogeneity in pre-ranking.
❌