❌

Normal view

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

arXiv:2510.18814v2 Announce Type: replace-cross Abstract: Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-training method that alternates between self-generation and training on self-generated responses. It repeatedly samples questions, uses the model itself to generate low-temperature responses, and then finetunes the model on the self-generated data. In this self-training loop, we use an online data refresh mechanism, where each new batch is generated by the most recently updated model. Across six math reasoning benchmarks, SePT improves a strong no-training baseline, defined as the untuned base model evaluated at its best swept decoding temperature, on several tested models. In some settings, SePT can even approach the performance of Reinforcement Learning with Verifiable Rewards (RLVR). Additional ablations demonstrate the importance of online data refresh and temperature decoupling. Overall, our results identify a practical regime in which reasoning can be improved using self-generated supervision alone. Our code is available at https://github.com/ElementQi/SePT.

NSUN2/ALYREF-mediated RNA m5c modification promotes anoikis resistance of prostate cancer through activating autophagy

Oncogene, Published online: 07 April 2026; doi:10.1038/s41388-026-03762-4

NSUN2/ALYREF-mediated RNA m5c modification promotes anoikis resistance of prostate cancer through activating autophagy

Unified modeling of 3D molecular generation via atomic interactions with PocketXMol

A versatile, atom-level generative AI model enables unified pocket-interacting tasks, from docking to de novo design, and demonstrates robust experimental validation for both small-molecule and peptide therapeutics.

Microbiome and metabolite signatures for cirrhosis to HCC risk stratification: progress, controversies, and gaps

Front Cell Infect Microbiol. 2026 Mar 16;16:1793213. doi: 10.3389/fcimb.2026.1793213. eCollection 2026.

ABSTRACT

The progression from cirrhosis to hepatocellular carcinoma (HCC) is a key outcome in the management of chronic liver disease. This process has a long incubation period and significant individual differences, making early warning still difficult. Clinical follow-up mainly relies on imaging examinations and alpha fetoprotein, but the ability to identify high risk precancerous states is limited. The imbalance of gut microbiota and its metabolites may occur earlier than the visible stage of tumors. They can affect barrier integrity, chronic inflammation, immune surveillance, and metabolic homeostasis through the gut liver axis, and participate in the formation of a pro tumor microenvironment. Therefore, such changes may provide more upstream risk stratification clues for the population with cirrhosis. This article summarizes previous research evidence and summarizes the common microbiome and metabolite characteristics of cirrhosis and high-risk populations, including a decrease in short chain fatty acid (SCFA) related symbiotic bacteria, an increase in inflammation related bacteria, bile acid spectrum shift, and other intestinal derived metabolite abnormalities. This article also outlines the key mechanisms that these features may correspond to, such as barrier damage and microbial translocation, immune suppression, etc. There are still significant uncertainties at present. The effect of SCFA is context dependent. Different etiologies, diets, medications, and complications can lead to significant confounding and affect cross cohort consistency. Subsequent research requires longitudinal cohort validation and the promotion of multi omics integration and the construction of interpretable predictive models to support clinical translation.

PMID:41918873 | PMC:PMC13033666 | DOI:10.3389/fcimb.2026.1793213

Electric dipole moment drives the dynamics of the TNFR1 complex I signalosome

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10304-1

Long-range interactions mediated by protein electric dipole moments have a role in driving the assembly and disassembly of super-signalling complex I for promoting NF-κB signalling.

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

arXiv:2602.13954v1 Announce Type: cross Abstract: We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.
❌