❌

Normal view

KIFC1 engages RUNX2/TGF-β signaling to promote lung cancer bone metastasis via disrupting bone homeostasis

Oncogene, Published online: 25 September 2026; doi:10.1038/s41388-026-03998-0

KIFC1 engages RUNX2/TGF-β signaling to promote lung cancer bone metastasis via disrupting bone homeostasis

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

arXiv:2609.12446v1 Announce Type: new Abstract: Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.

A complement C5-targeted GalNAc-conjugated siRNA with sustained efficacy in a non-human primate model of IgA nephropathy

This study characterizes a GalNAc-C5 small interfering RNA with potent in vitro and in vivo activity. Single subcutaneous dosing sustains long-term C5 suppression in cynomolgus monkeys with IgA nephropathy, outperforming Nefecon in blocking glomerular complement deposition, supporting its standalone or combinational clinical application.

In vivo-directed evolution identifies AAV-WM04 as a next-generation vector for potent and sustained hearing restoration in DFNB9

AAV-WM04, an AAV vector identified through in-vivo-directed screening in the adult cochlea, enables highly efficient and selective inner hair cell transduction. Dual-AAV delivery of OTOF using AAV-WM04 restores hearing in a DFNA9 deafness mouse model at low doses, highlighting its translational potential for gene therapy.

Integrative Multi-Omics Mendelian Randomization Analysis Identifies NIT2 as a Potential Metabolic Risk Gene in Hepatocellular Carcinoma

J Gene Med. 2026 Sep;28(9):e70111. doi: 10.1002/jgm.70111.

ABSTRACT

BACKGROUND: Metabolic pathways are crucial in hepatocellular carcinoma (HCC) pathogenesis, but causal metabolic genes remain unclear. This study used Summary data-based Mendelian Randomization (SMR) and colocalization to identify metabolism-related genetic loci influencing HCC risk.

METHODS: Differentially expressed genes in hepatic malignancy phenotype versus normal tissues from TCGA and GTEx were analyzed. Metabolism-related candidates were examined via SMR and colocalization using multi-omics data: methylation (mQTL), expression (eQTL), and protein (pQTL) quantitative trait loci.

RESULTS: Multi-omics integration identified NIT2 as a key metabolic regulator for HCC. The cg13016775 locus of NIT2 was associated with elevated HCC risk at gene (OR = 1.618, 95% CI: 1.199-2.182) and protein (OR = 4.432, 95% CI: 1.783-11.018) levels. Colocalization supported a shared causal variant (PPH4 > 0.6), linking NIT2 to hepatocarcinogenesis via metabolic regulation.

CONCLUSIONS: This study provides multi-omics evidence for NIT2 as a potential causal gene in HCC, enhancing understanding of metabolic contributions to HCC pathogenesis and highlighting integrative genomics for uncovering causal relationships.

PMID:42681890 | PMC:PMC13534973 | DOI:10.1002/jgm.70111

Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration

arXiv:2605.24957v1 Announce Type: new Abstract: The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to address this issue - ranging from expensive data-driven fine-tuning and high-latency contrastive decoding to rigid attention head truncation - frequently compromise either computational efficiency or the continuity of the model's feature space. To overcome these limitations, we introduce a novel, training-free inference strategy that operates as a region-aware adaptive weighting mechanism to dynamically correct semantic drift without relying on abrupt heuristic truncations. By computing an outlier-resistant statistical midpoint across various attention heads, we establish a stable anchor for reliable visual representations. We then utilize the inter-head disagreement mapped across regions to dynamically determine intervention budgets, gently suppressing hallucination-inducing attention paths through a continuous penalty modulation. This recalibration process effectively rectifies visual-semantic misalignments while fully preserving generative fluency and language priors. Comprehensive evaluations on standard multimodal benchmarks, including CHAIR, POPE, and MME, reveal that our strategy substantially curtails both instance- and sentence-level hallucinations. The results demonstrate state-of-the-art performance against contemporary baselines, confirming our method's efficiency and algorithmic robustness. Our code will be public.

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

arXiv:2605.24993v1 Announce Type: new Abstract: Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically faithful surface-based models. We argue this is partly driven by inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present NeurIPS, a framework that improves surface-based decoding by reframing anatomical variation from a nuisance to a powerful inductive prior. NeurIPS unites two innovations: a Selective ROI Spherical Tokenizer (SRST) for efficient geometric encoding, and a Structure-Guided Mixture of Experts (SG-MoE) that explicitly models individual anatomy using cortical features. On the Natural Scenes Dataset, NeurIPS establishes a new state-of-the-art for surface decoders and achieves performance comparable to strong 1D baselines. This is achieved with unprecedented efficiency, as the model converges dramatically faster (10 vs. 600 epochs). This efficiency enables rapid adaptation to new subjects using only 20% of data and ensures robust scalability as the training cohort is expanded. Ablations provide causal evidence that these gains are driven by the model's use of cortical features, not by memorizing subject IDs. By leveraging anatomical priors, NeurIPS provides a principled and scalable path toward robust, generalizable brain decoding.

Cascade-KDE: Robust Time-Series Restoration under Out-of-Distribution Impulse Corruptions

arXiv:2605.24055v1 Announce Type: cross Abstract: Real-world time-series data in industrial sensing, healthcare, and energy systems is often corrupted by a mixture of Gaussian noise and occasional large-magnitude impulse outliers. For tasks that depend on local shape, such as ECG morphology analysis and battery degradation monitoring, the main requirement is not only low reconstruction error but also preservation of derivative peaks and task-critical features. We propose Cascade-KDE, a training-free restoration framework for corrupted time series. The method first estimates a two-dimensional temporal-amplitude density, then applies a Density-Truncated Robust Expectation to limit the influence of distant abnormal points, and finally refines the sequence through an exponential cascade with adaptive stopping. This design aims to improve robustness under out-of-distribution impulse corruptions while keeping the restored trajectory close to the original local structure. Across several benchmark datasets, the proposed method shows consistent gains over classical filters and representative learning-based baselines on curve fidelity, derivative preservation, downstream classification, and runtime efficiency. These results suggest that bounded density-based restoration is a practical option for feature-preserving preprocessing in noisy time-series pipelines.

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

arXiv:2605.24810v1 Announce Type: cross Abstract: Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under mismatched transition dynamics. Existing approaches such as reward augmentation and data filtering are constrained to the source dataset and cannot synthesize new target behavior to improve coverage beyond the collected source trajectories. While recent model-based methods attempt to address this by learning target-aware dynamics, the generated experience is constructed only at the transition level, which leads to accumulated errors over long horizons. These limitations necessitate a shift toward trajectory-level generation for off-dynamics offline RL. We propose CEDGE, a Cross-domain Energy-guided Diffusion GEneration framework. CEDGE trains a trajectory diffusion model on source-domain trajectories and adapts the generated samples to the target domain through energy guidance. This guidance is derived by minimizing the distribution mismatch between the source and desired target-domain trajectories and is decomposed into return, domain, and behavior energy components. The resulting energy-guided trajectories are useful both for direct planning and as synthetic data for policy learning. Since target adaptation is achieved via energy guidance rather than retraining the diffusion model, CEDGE can be efficiently adapted to new target dynamics compared to previous methods. Experiments on the ODRL benchmark demonstrate that trajectory-level energy-guided generation improves diffusion planning under dynamics shifts and produces synthetic data that improves downstream target policy learning.

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its weights. Recent Sequential Monte Carlo (SMC)-based steering methods approximate reward-tilted target distributions in a principled way, but their proposals remain largely tied to the base sampler. Since reward information is mainly used after propagation through particle reweighting and resampling, these methods can require large particle budgets and suffer from weight degeneracy and high-variance estimates. One way to reduce variance and improve particle efficiency is to iteratively learn twisting functions that provide look-ahead guidance, as in twisted SMC. However, existing learnable twisting methods are developed mainly for classical sequential inference and can be unstable when applied to diffusion-based alignment with high-dimensional state spaces and terminal, noisy, or black-box rewards. We propose Trust-Region Iterative Twisted Sequential Monte Carlo (TRI-TSMC), a trust-region framework for learning twisting functions in SMC-based inference-time alignment. Each iteration computes an exact KL-constrained update in path space, which admits a closed-form solution by tempered importance reweighting, and projects this target back to the parameterized twisted family by weighted maximum likelihood. Theoretically, we formalize the value-function interpretation of the optimal twisting function and show that it yields a zero-variance sampler. We prove that the trust-region update follows an escort path toward the target distribution, that the weighted maximum-likelihood update is a forward-KL projection, and that the path reduces residual importance-weight variance. Empirically, TRI-TSMC improves primary alignment objectives on discrete diffusion text generation and text-to-image generation under matched inference-time budgets.

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

arXiv:2605.20025v2 Announce Type: replace Abstract: Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail and inform the next attempt, and lessons accumulate across cycles. Existing autonomous research systems often model this process as a linear pipeline: they rely on single-agent reasoning, stop when execution fails, and do not carry experience across runs. We present AutoResearchClaw, a multi-agent autonomous research pipeline built on five mechanisms: structured multi-agent debate for hypothesis generation and result analysis, a self-healing executor with a \textsc{Pivot}/\textsc{Refine} decision loop that transforms failures into information, verifiable result reporting that prevents fabricated numbers and hallucinated citations, human-in-the-loop collaboration with seven intervention modes spanning full autonomy to step-by-step oversight, and cross-run evolution that converts past mistakes into future safeguards. On ARC-Bench, a 25-topic experiment-stage benchmark, AutoResearchClaw outperforms AI Scientist v2 by 54.7%. A human-in-the-loop ablation across seven intervention modes reveals that precise, targeted collaboration at high-leverage decision points consistently outperforms both full autonomy and exhaustive step-by-step oversight. We position AutoResearchClaw as a research amplifier that augments rather than replaces human scientific judgment. Code is available at https://github.com/aiming-lab/AutoResearchClaw.

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

arXiv:2605.09270v2 Announce Type: replace-cross Abstract: Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is not memorization itself, but its target: vanilla SFT drives models to exploit and memorize spurious surface correlations in problem-solution pairs, leaving them brittle to superficial input variations. To address this, we propose Theorem-SFT, which reorients supervision toward explicit theorem application by teaching models how rules are invoked rather than what answers look like. Theorem-SFT yields consistent gains across benchmarks and model families: +8.8% on MATH (LLaMA3.2-3B-Instruct) and +20.27% on GeoQA (Qwen2.5-VL-7B-Instruct) without modality-specific re-training. Fine-tuning MLP layers alone matches full-layers performance, implicating feed-forward components as the primary locus of reasoning rules. Our findings reframe the debate: Generalization failures stem not from memorization as a mechanism, but from memorizing the wrong inductive targets.

Epigenome-wide Mendelian randomization with multi-omics validation identifies epigenetic drivers of idiopathic pulmonary fibrosis

Commun Biol. 2026 Apr 11. doi: 10.1038/s42003-026-10033-1. Online ahead of print.

ABSTRACT

Idiopathic pulmonary fibrosis (IPF) is a complex disease without clear etiology or effective therapy. While DNA methylation has been implicated in IPF pathogenesis, the tissue-specific causal effects of the epigenetic factors on IPF remain undetermined. Here, we perform epigenome-wide Mendelian randomization using blood-based methylation quantitative trait loci of 420,509 CpG sites and genome-wide association study for IPF to elucidate the causal effects of the CpG sites on IPF. Totally, 452 CpG sites has shown putative causal effects on IPF risk after Bonferroni correction. Among them, 13 CpG sites have shown strong colocalization evidence with genetic factors associated with IPF. Specifically, DNA methylation at CpG sites within MAN2A2 and TRIM27 shows significant differences between IPF lungs and controls, correlating with altered mRNA expressions of these genes in lung tissues. The CpG site in MAN2A2 is a binding site of ZNF384 according to transcription factor databases. RNA sequencing in the TGFβ1-induced alveolar epithelia confirms significantly reduced expression of MAN2A2 and ZNF384 comparing to the controls. Collectively, our study suggests a putative causal link between DNA methylation within MAN2A2 and IPF risk, wherein lung-specific DNA methylation in MAN2A2 may perturb the interaction between ZNF384 and MAN2A2, revealing novel roles for these genes in IPF pathogenesis.

PMID:41965819 | DOI:10.1038/s42003-026-10033-1

Epigenome-wide Mendelian randomization with multi-omics validation identifies epigenetic drivers of idiopathic pulmonary fibrosis

Commun Biol. 2026 Apr 11. doi: 10.1038/s42003-026-10033-1. Online ahead of print.

ABSTRACT

Idiopathic pulmonary fibrosis (IPF) is a complex disease without clear etiology or effective therapy. While DNA methylation has been implicated in IPF pathogenesis, the tissue-specific causal effects of the epigenetic factors on IPF remain undetermined. Here, we perform epigenome-wide Mendelian randomization using blood-based methylation quantitative trait loci of 420,509 CpG sites and genome-wide association study for IPF to elucidate the causal effects of the CpG sites on IPF. Totally, 452 CpG sites has shown putative causal effects on IPF risk after Bonferroni correction. Among them, 13 CpG sites have shown strong colocalization evidence with genetic factors associated with IPF. Specifically, DNA methylation at CpG sites within MAN2A2 and TRIM27 shows significant differences between IPF lungs and controls, correlating with altered mRNA expressions of these genes in lung tissues. The CpG site in MAN2A2 is a binding site of ZNF384 according to transcription factor databases. RNA sequencing in the TGFβ1-induced alveolar epithelia confirms significantly reduced expression of MAN2A2 and ZNF384 comparing to the controls. Collectively, our study suggests a putative causal link between DNA methylation within MAN2A2 and IPF risk, wherein lung-specific DNA methylation in MAN2A2 may perturb the interaction between ZNF384 and MAN2A2, revealing novel roles for these genes in IPF pathogenesis.

PMID:41965819 | DOI:10.1038/s42003-026-10033-1

Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing

arXiv:2604.02288v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models. While Group Relative Policy Optimization (GRPO) is widely adopted, its coarse credit assignment uniformly penalizes failed rollouts, lacking the token-level focus needed to efficiently address specific deviations. Self-Distillation Policy Optimization (SDPO) addresses this by providing denser, more targeted logit-level supervision that facilitates rapid early improvement, yet it frequently collapses during prolonged training. We trace this late-stage instability to two intrinsic flaws: self-distillation on already-correct samples introduces optimization ambiguity, and the self-teacher's signal reliability progressively degrades. To resolve these issues, we propose Sample-Routed Policy Optimization (SRPO), a unified on-policy framework that routes correct samples to GRPO's reward-aligned reinforcement and failed samples to SDPO's targeted logit-level correction. SRPO further incorporates an entropy-aware dynamic weighting mechanism to suppress high-entropy, unreliable distillation targets while emphasizing confident ones. Evaluated across five benchmarks and two model scales, SRPO achieves both the rapid early improvement of SDPO and the long-horizon stability of GRPO. It consistently surpasses the peak performance of both baselines, raising the five-benchmark average on Qwen3-8B by 3.4% over GRPO and 6.3% over SDPO, while simultaneously yielding moderate response lengths and lowering per-step compute cost by up to 17.2%.

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

arXiv:2603.21289v2 Announce Type: replace-cross Abstract: Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self-consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self-evolving multimodal models. The code are available at https://github.com/OPPO-Mente-Lab/LLM-Self-Judge.

Lysine attenuates acute lung injury by restoring α-tubulin acetylation and ciliary activity

Cell Death Discovery, Published online: 16 March 2026; doi:10.1038/s41420-026-03025-x

Lysine attenuates acute lung injury by restoring α-tubulin acetylation and ciliary activity

Deconstructing Multimodal Mathematical Reasoning: Towards a Unified Perception-Alignment-Reasoning Paradigm

arXiv:2603.08291v1 Announce Type: new Abstract: Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems that involve both textual and visual modalities. However, current models still face significant challenges in real-world visual math tasks. They often misinterpret diagrams, fail to align mathematical symbols with visual evidence, and produce inconsistent reasoning steps. Moreover, existing evaluations mainly focus on checking final answers rather than verifying the correctness or executability of each intermediate step. To address these limitations, a growing body of recent research addresses these issues by integrating structured perception, explicit alignment, and verifiable reasoning within unified frameworks. To establish a clear roadmap for understanding and comparing different MMR approaches, we systematically study them around four fundamental questions: (1) What to extract from multimodal inputs, (2) How to represent and align textual and visual information, (3) How to perform the reasoning, and (4) How to evaluate the correctness of the overall reasoning process. Finally, we discuss open challenges and offer perspectives on promising directions for future research.

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

arXiv:2603.07980v1 Announce Type: cross Abstract: As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench \$OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, \$OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.

Performance of Conventional EEG Biomarkers Across Different Clinical Phases of Major Depressive Disorder: A Comprehensive Evaluation

arXiv:2603.03864v1 Announce Type: new Abstract: While EEG features differentiate Major Depressive Disorder (MDD) from healthy controls (HC), their clinical utility as biomarkers depends on a monotonic trajectory across the disease spectrum, from the acute (AC) phase to the maintenance (MA) phase and finally to the healthy baseline. However, the progression of the MA phase remains poorly understood in traditional marker analysis. Analyzing EEG data from 74 individuals (24 AC, 23 MA, and 27 HC), this study provides a comprehensive evaluation of classic ERP and resting-state indices across AC, MA, and HC groups. Our results demonstrate that almost no conventional metrics strictly satisfy the criterion of monotonic progression, likely due to profound inter-individual heterogeneity. These findings highlight the inherent limitations of group-level feature extraction and provide critical insights for developing future paradigms and algorithms to identify neurobiological markers with genuine clinical utility.
❌