❌

Normal view

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

npj Digital Medicine, Published online: 12 September 2026; doi:10.1038/s41746-026-03238-5

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

Coupled Variational Reinforcement Learning for Language Model General Reasoning

arXiv:2512.12576v3 Announce Type: replace-cross Abstract: While reinforcement learning has achieved impressive progress in language model reasoning, it is constrained by the requirement for verifiable rewards. Recent verifier-free RL methods address this limitation by utilizing the probabilities that LLMs generate reference answers as reward signals. However, these approaches typically sample reasoning traces conditioned only on the question. This design decouples reasoning-trace sampling from answer information, leading to inefficient exploration and incoherence between traces and final answers. In this paper, we propose \textit{\b{Co}upled \b{V}ariational \b{R}einforcement \b{L}earning} (CoVRL), which bridges variational inference and reinforcement learning by coupling prior and posterior distributions through a hybrid sampling strategy. By constructing and optimizing a composite distribution that integrates these two distributions, CoVRL enables efficient exploration while preserving strong thought-answer coherence. Extensive experiments on mathematical and general reasoning benchmarks show that CoVRL improves performance by 12.4\% over the base model and achieves an additional 2.3\% improvement over state-of-the-art verifier-free RL baselines, providing a principled framework for enhancing the general reasoning capabilities of language models.

Unlocking the Future of Hepatocellular Carcinoma Early Diagnosis: The Promise of Extracellular Vesicle Biomarkers

J Clin Transl Hepatol. 2026 Apr 28;14(4):462-477. doi: 10.14218/JCTH.2025.00589. Epub 2026 Apr 8.

ABSTRACT

Hepatocellular carcinoma (HCC) is one of the most prevalent and aggressive malignant tumors globally, with a notably low five-year survival rate. Its high mortality is largely attributed to challenges in early detection. Extracellular vesicles (EVs) are naturally occurring nanoparticles secreted by nearly all cell types and carry a diverse array of bioactive molecules, including proteins, nucleic acids (particularly non-coding RNAs), and lipids. EVs play pivotal roles in remodeling the tumor microenvironment and driving cancer progression through intercellular communication. Accumulating evidence has established that EVs are critically involved in the pathogenesis of HCC and are emerging as promising biomarkers for its early detection. With advances in EV isolation technologies, these vesicles have garnered considerable attention in the field of liquid biopsy for HCC. This review provides a comprehensive overview of the diagnostic potential of EV-derived biomarkers in HCC, including DNA, RNA, proteins, and lipids. Additionally, it discusses the advantages of integrating multi-omics approaches for HCC diagnosis. Furthermore, the review highlights the technical challenges in EV isolation and characterization, as well as the crucial role of reference genes in the standardization of EV data. These insights underscore the potential of EVs as novel, minimally invasive liquid biopsy biomarkers for the early diagnosis of HCC.

PMID:42181837 | PMC:PMC13195390 | DOI:10.14218/JCTH.2025.00589

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

arXiv:2603.06652v1 Announce Type: cross Abstract: Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.
❌