❌

Normal view

KIFC1 engages RUNX2/TGF-β signaling to promote lung cancer bone metastasis via disrupting bone homeostasis

Oncogene, Published online: 25 September 2026; doi:10.1038/s41388-026-03998-0

KIFC1 engages RUNX2/TGF-β signaling to promote lung cancer bone metastasis via disrupting bone homeostasis

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

arXiv:2607.19363v3 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.

When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs

arXiv:2511.07318v3 Announce Type: replace-cross Abstract: Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and proprietary LLMs (including GPT-5), we show that existing hallucination detection methods, such as confidence-based filtering and inner-state probing, fundamentally fail in the presence of spurious correlations. Our theoretical analysis further elucidates why these statistical biases intrinsically undermine confidence-based detection techniques. Our findings thus emphasize the urgent need for new approaches explicitly designed to address hallucinations caused by spurious correlations.

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

npj Digital Medicine, Published online: 12 September 2026; doi:10.1038/s41746-026-03238-5

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

Distilling Image Prototypes for Guided Test-Time Adaptation

arXiv:2609.09737v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) enhances the robustness of models against distribution shifts but faces two critical challenges: error accumulation from noisy pseudo-labels and catastrophic forgetting of source knowledge. Uncertainty-based approaches designed to mitigate error accumulation often yield overconfident or computationally expensive estimates, while strategies intended to prevent forgetting via prototype replay rely on static representations that easily become misaligned as the model adapts. To address these issues, this paper proposes a novel framework, Distilling Image Prototype for Guided Test-Time Adaptation (DIPTTA). The core of the proposed approach is the introduction of a Distill Image Prototype (DIP), a compact set of synthetic images that serves as a dynamic and regenerative anchor of source knowledge. This prototype enables a dynamic feature replay mechanism that continuously generates feature prototypes aligned with the current state of the model, thus effectively preventing catastrophic forgetting. Furthermore, the DIP anchors a source-calibrated uncertainty estimation method, which provides a less biased measure of sample reliability by leveraging stable source knowledge, thereby robustly suppressing error accumulation. Extensive experiments on multiple benchmarks demonstrate that DIPTTA significantly outperforms state-of-the-art methods, particularly under severe domain shifts. The source code is available at https://github.com/LiwenWang919/DIPTTA.

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for CUAs requires consistent task instruction, executable environment, and verifiable reward. However, hand-curated benchmarks achieve high reward fidelity but cover few applications and LLM-as-judge-based datasets scale broadly but lack reliable verification. We present CUA-Gym, a scalable pipeline that co-generates task instructions, environment states, and reward functions. Concretely, a Generator agent constructs the initial and golden environment states, and a separate Discriminator agent writes the reward function from the task specification. An orchestrator agent drives the two through iterative rounds upon execution. Generated tuples then pass a final filter combining LLM majority voting and agent rollouts, ensuring quality beyond the per-task adversarial loop. To address the scarcity of training environments, we further synthesize CUA-Gym-Hub, a broad suite of high-fidelity mock web applications grounded in real-world software-use distributions, expanding the scale of CUA RLVR data by magnitude. Using this pipeline, we construct CUA-Gym, a dataset of 32,112 verified RLVR training tuples grounded in 110 environments. Trained with GSPO on CUA-Gym, our CUA-Gym-A3B and CUA-Gym-A17B achieve 62.1% and 72.6% on OSWorld-Verified, outperforming prior open-source CUAs at comparable scales, with performance scaling smoothly in both data volume and environment diversity. The same checkpoints also improve on the held-out WebArena benchmark, indicating transfer beyond the training environments. We will open-source the full synthesis pipeline, dataset, CUA-Gym-Hub environments, and models.

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers

arXiv:2605.25949v1 Announce Type: cross Abstract: Neural PDE solvers have followed the scaling trajectory of vision and language, with recent foundation models reaching billions of parameters. We argue that scale is a poor substitute for architectural inductive bias in this domain: structured priors deliver outsized parameter efficiency, and the pattern of where they succeed and fail is itself informative about what they capture. We instantiate this argument in WaveLiT, an architecture combining a discrete wavelet transform for lossless multi-resolution tokenization, an augmented linear attention block, a shared-weight multiscale feature pyramid, and a wavelet-domain auxiliary loss. Bespoke 1-10M-parameter WaveLiT models compete with foundation models of 100-1000$\times$ their size across eight TheWell benchmarks, with the largest gains on wave and acoustic-dominated benchmarks where the wavelet-multiscale prior fits the dominant dynamical structure and small per-step errors do not compound geometrically under rollout. Trained jointly across all eight benchmarks, a 10M-parameter foundation variant exhibits a structured, physically interpretable transfer pattern -- strongest where the wavelet-multiscale prior matches the dynamics, weakest on chaotic advection-dominated flows. The entire pipeline trains on a single GPU. The results suggest that small-model PDE performance is shaped by architectural inductive bias rather than scale, and that the structure of a prior's failures is a useful empirical signal about its content.

SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering

arXiv:2601.03014v3 Announce Type: replace-cross Abstract: Traditional Retrieval-Augmented Generation (RAG) effectively supports single-hop question answering with large language models but faces significant limitations in multi-hop question answering tasks, which require combining evidence from multiple documents. Existing chunk-based retrieval often provides irrelevant and logically incoherent context, leading to incomplete evidence chains and incorrect reasoning during answer generation. To address these challenges, we propose SentGraph, a sentence-level graph-based RAG framework that explicitly models fine-grained logical relationships between sentences for multi-hop question answering. Specifically, we construct a hierarchical sentence graph offline by first adapting Rhetorical Structure Theory to distinguish nucleus and satellite sentences, and then organizing them into topic-level subgraphs with cross-document entity bridges. During online retrieval, SentGraph performs graph-guided evidence selection and path expansion to retrieve fine-grained sentence-level evidence. Extensive experiments on four multi-hop question answering benchmarks demonstrate the effectiveness of SentGraph, validating the importance of explicitly modeling sentence-level logical dependencies for multi-hop reasoning.

Nonlinear atomic tunnelling boosted by bright squeezed vacuum

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10485-9

Bright squeezed vacuum light boosts nonlinear atomic tunnelling ionization more than 20-fold compared with coherent light, enabling quantum control of strong-field processes without increasing classical intensity.

Pulmonary-Intestinal Axis: Shared Genetic Basis and Mediating Factors Identified Through Multi-Omics Analysis

Int J Chron Obstruct Pulmon Dis. 2026 Apr 7;21:561645. doi: 10.2147/COPD.S561645. eCollection 2026.

ABSTRACT

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a systemic condition with comorbidities beyond the lung (eg, cardiovascular and metabolic disorders), and gastrointestinal (GI) disorders are also common. The shared genetic basis of COPD-GI comorbidity and its mediating factors remain unclear. We hypothesized that COPD and GI diseases share pleiotropic genetic architecture implicating lipid-metabolic pathways, with smoking mediating part of the association.

METHODS: We analyzed publicly available European-ancestry GWAS summary statistics for COPD (Global Biobank Meta-analysis Initiative), 15 GI diseases (FinnGen), and smoking phenotypes (UK Biobank). Genetic correlation was estimated using linkage disequilibrium score regression (LDSC) and high-definition likelihood (HDL). Multi-trait analysis of GWAS (MTAG) boosted COPD discovery by leveraging genetically correlated GI traits. We integrated locus-to-gene mapping with multi-tissue expression quantitative trait loci (eQTL) and plasma protein quantitative trait loci (pQTL) evidence to prioritize shared loci, genes, and proteins. Bidirectional two-sample Mendelian randomization (MR) tested causal directions, and two-step mediation MR evaluated smoking.

RESULTS: COPD showed significant genetic correlation with nine GI diseases. We identified six comorbidity-associated loci (three with CADD > 12.37) and 13 unique candidate pleiotropic genes; APOE was supported by proteomic evidence. Enrichment analyses highlighted lipid-metabolism pathways. MR suggested COPD increases risk of gastroesophageal reflux disease (GERD), irritable bowel syndrome (IBS), acute appendicitis, and gastric ulcer, while diverticular disease showed reverse causality toward COPD. Smoking partially mediated the COPD effect on GERD, acute appendicitis, and gastric ulcer.

CONCLUSION: COPD and multiple GI disorders share a distributed pleiotropic genetic basis within the broader systemic comorbidity spectrum of COPD. Multi-omics evidence supports a genomic pulmonary-intestinal axis in which lipid metabolism and smoking-related mechanisms contribute to COPD and GI comorbidity, providing targets for risk stratification and potential intervention.

PMID:41978582 | PMC:PMC13070119 | DOI:10.2147/COPD.S561645

Integrin β3 deficiency unleashes spontaneous pulmonary inflammation by promoting B cell hyperactivation via the CD40-CD40L axis

Front Immunol. 2026 Mar 24;17:1796926. doi: 10.3389/fimmu.2026.1796926. eCollection 2026.

ABSTRACT

BACKGROUND: Pulmonary immune homeostasis requires tight control of adaptive responses. Integrin β3 is a well-known mediator of cell adhesion and platelet function. However, its role in adaptive immunity, especially in B cell responses, remains unclear.

METHODS: We defined the pulmonary phenotype of constitutive β3-deficient (β3-/-) mice by histopathology. We performed integrated transcriptomic and proteomic profiling of lung tissue to map the molecular signature of spontaneous pulmonary inflammation. We further probed the underlying mechanisms with additional histology and functional assays and tested for biological significance using transcriptomics data from auto-immune disease patients.

RESULTS: β3-/- mice developed spontaneous pulmonary inflammation marked by B cell activation and in situ immune-complex deposition within alveoli. Multi-omics integration implicated the CD40-CD40 Ligand (CD40L) axis as a central driver of this pathology. Mechanistically, loss of β3 enhanced CD40L-CD40 engagement on B cells, resulting in NF-κB pathway hyperactivation. Consistent with our murine data, reduced ITGB3 expression in patients with autoimmune disease correlated with transcriptional signatures of B cell activation and inflammation.

CONCLUSIONS: These results reframe integrin β3 as a threshold regulator of B cell activation. The β3-CD40L-CD40 axis therefore represents a potential therapeutic target for B cell-mediated autoimmune diseases.

PMID:41953039 | PMC:PMC13055533 | DOI:10.3389/fimmu.2026.1796926

Integrin β3 deficiency unleashes spontaneous pulmonary inflammation by promoting B cell hyperactivation via the CD40-CD40L axis

Front Immunol. 2026 Mar 24;17:1796926. doi: 10.3389/fimmu.2026.1796926. eCollection 2026.

ABSTRACT

BACKGROUND: Pulmonary immune homeostasis requires tight control of adaptive responses. Integrin β3 is a well-known mediator of cell adhesion and platelet function. However, its role in adaptive immunity, especially in B cell responses, remains unclear.

METHODS: We defined the pulmonary phenotype of constitutive β3-deficient (β3-/-) mice by histopathology. We performed integrated transcriptomic and proteomic profiling of lung tissue to map the molecular signature of spontaneous pulmonary inflammation. We further probed the underlying mechanisms with additional histology and functional assays and tested for biological significance using transcriptomics data from auto-immune disease patients.

RESULTS: β3-/- mice developed spontaneous pulmonary inflammation marked by B cell activation and in situ immune-complex deposition within alveoli. Multi-omics integration implicated the CD40-CD40 Ligand (CD40L) axis as a central driver of this pathology. Mechanistically, loss of β3 enhanced CD40L-CD40 engagement on B cells, resulting in NF-κB pathway hyperactivation. Consistent with our murine data, reduced ITGB3 expression in patients with autoimmune disease correlated with transcriptional signatures of B cell activation and inflammation.

CONCLUSIONS: These results reframe integrin β3 as a threshold regulator of B cell activation. The β3-CD40L-CD40 axis therefore represents a potential therapeutic target for B cell-mediated autoimmune diseases.

PMID:41953039 | PMC:PMC13055533 | DOI:10.3389/fimmu.2026.1796926

InCoder-32B: Code Foundation Model for Industrial Scenarios

arXiv:2603.16790v3 Announce Type: replace-cross Abstract: Recent code large language models have achieved remarkable progress on general programming tasks. Nevertheless, their performance degrades significantly in industrial scenarios that require reasoning about hardware semantics, specialized language constructs, and strict resource constraints. To address these challenges, we introduce InCoder-32B (Industrial-Coder-32B), the first 32B-parameter code foundation model unifying code intelligence across chip design, GPU kernel optimization, embedded systems, compiler optimization, and 3D modeling. By adopting an efficient architecture, we train InCoder-32B from scratch with general code pre-training, curated industrial code annealing, mid-training that progressively extends context from 8K to 128K tokens with synthetic industrial reasoning data, and post-training with execution-grounded verification. We conduct extensive evaluation on 14 mainstream general code benchmarks and 9 industrial benchmarks spanning 4 specialized domains. Results show InCoder-32B achieves highly competitive performance on general tasks while establishing strong open-source baselines across industrial domains.

How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation

arXiv:2603.07540v1 Announce Type: cross Abstract: Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffer from a critical reliability gap: as sequences grow, generation quality rapidly collapses. In this work, we investigate the mechanism behind this failure and argue that it is distinct from standard long-context challenges. We reveal that in generation, accumulated visual history acts as a source of active pollution, a decay governed specifically by the number of image events rather than raw token count. We identify a structural vulnerability where dense visual tokens overwhelm the attention mechanism, creating noise that distorts future synthesis. Guided by these mechanistic insights, we propose UniLongGen, a training-free inference strategy that prioritizes safe conditioning over total recall. Instead of retaining all history, UniLongGen dynamically curates the model's memory, identifying and discarding interfering visual signals based on the model's own internal relevance rankings. Extensive experiments demonstrate that this active forgetting approach is essential for stability: UniLongGen significantly outperforms baselines in long-horizon fidelity and consistency, while simultaneously reducing memory footprint and inference time.

BoxMind: Closed-loop AI strategy optimization for elite boxing validated in the 2024 Olympics

arXiv:2601.11492v2 Announce Type: replace Abstract: Competitive sports require sophisticated tactical analysis, yet combat disciplines like boxing remain underdeveloped in AI-driven analytics due to the complexity of action dynamics and the lack of structured tactical representations. To address this, we present BoxMind, a closed-loop AI expert system validated in elite boxing competition. By defining atomic punch events with precise temporal boundaries and spatial and technical attributes, we parse match footage into 18 hierarchical technical-tactical indicators. We then propose a graph-based predictive model that fuses these explicit technical-tactical profiles with learnable, time-variant latent embeddings to capture the dynamics of boxer matchups. Modeling match outcome as a differentiable function of technical-tactical indicators, we turn winning probability gradients into executable tactical adjustments. Experiments show that the outcome prediction model achieves state-of-the-art performance, with 69.8% accuracy on BoxerGraph test set and 87.5% on Olympic matches. Using this predictive model as a foundation, the system generates strategic recommendations that demonstrate proficiency comparable to human experts. BoxMind is validated through a closed-loop deployment during the 2024 Paris Olympics, directly contributing to the Chinese National Team's historic achievement of three gold and two silver medals. BoxMind establishes a replicable paradigm for transforming unstructured video data into strategic intelligence, bridging the gap between computer vision and decision support in competitive sports. Code and data is available at https://github.com/gouba2333/BoxingWeb.

Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration

arXiv:2603.02760v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) have recently attracted significant attention for their ability to enhance diversity, controllability, and parallelism. However, their non-sequential, bidirectionally masked generation makes quality assessment difficult, underscoring the need for effective self-evaluation. In this work, we propose DiSE, a simple yet effective self-evaluation confidence quantification method for dLLMs. DiSE quantifies confidence by computing the probability of regenerating the tokens in the entire generated sequence, given the full context. This method enables more efficient and reliable quality assessment by leveraging token regeneration probabilities, facilitating both likelihood estimation and robust uncertainty quantification. Building upon DiSE, we further introduce a flexible-length generation framework, which adaptively controls the sequence length based on the model's self-assessment of its own output. We analyze and validate the feasibility of DiSE from the perspective of dLLM generalization, and empirically demonstrate that DiSE is positively correlated with both semantic coherence and answer accuracy. Extensive experiments on likelihood evaluation, uncertainty quantification, and flexible-length generation further confirm the effectiveness of the proposed DiSE.

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

arXiv:2506.03922v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines, while overlooking the distinct needs and potential of the Humanities and Social Sciences (HSS). Tasks in the HSS domain require more horizontal, interdisciplinary thinking and a deep integration of knowledge across related fields, which presents unique challenges for MLLMs, particularly in linking abstract concepts with corresponding visual representations. Addressing this gap, we present HSSBench, a dedicated benchmark designed to assess the capabilities of MLLMs on HSS tasks in multiple languages, including the six official languages of the United Nations. We also introduce a novel data generation pipeline tailored for HSS scenarios, in which multiple domain experts and automated agents collaborate to generate and iteratively refine each sample. HSSBench contains over 13,000 meticulously designed samples, covering six key categories. We benchmark more than 20 mainstream MLLMs on HSSBench and demonstrate that it poses significant challenges even for state-of-the-art models. We hope that this benchmark will inspire further research into enhancing the cross-disciplinary reasoning abilities of MLLMs, especially their capacity to internalize and connect knowledge across fields.

Generative Reasoning Re-ranker

arXiv:2602.07774v4 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge. However, existing work has three key limitations: (1) most efforts focus on retrieval and ranking, while the reranking phase, critical for refining final recommendations, is largely overlooked; (2) LLMs are typically used in zero-shot or supervised fine-tuning settings, leaving their reasoning abilities, especially those enhanced through reinforcement learning (RL) and high-quality reasoning data, underexploited; (3) items are commonly represented by non-semantic IDs, creating major scalability challenges in industrial systems with billions of identifiers. To address these gaps, we propose the Generative Reasoning Reranker (GR2), an end-to-end framework with a three-stage training pipeline tailored for reranking. First, a pretrained LLM is mid-trained on semantic IDs encoded from non-semantic IDs via a tokenizer achieving $\ge$99% uniqueness. Next, a stronger larger-scale LLM generates high-quality reasoning traces through carefully designed prompting and rejection sampling, which are used for supervised fine-tuning to impart foundational reasoning skills. Finally, we apply Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO), enabling scalable RL supervision with verifiable rewards designed specifically for reranking. Experiments on two real-world datasets demonstrate GR2's effectiveness: it surpasses the state-of-the-art OneRec-Think by 2.4% in Recall@5 and 1.3% in NDCG@5. Ablations confirm that advanced reasoning traces yield substantial gains across metrics. We further find that RL reward design is crucial in reranking: LLMs tend to exploit reward hacking by preserving item order, motivating conditional verifiable rewards to mitigate this behavior and optimize reranking performance.
❌