❌

Normal view

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

arXiv:2508.10530v2 Announce Type: replace Abstract: The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize the expected reward that reflects human preferences. Recently, Direct Preference Optimization~(DPO) was proposed as a LM alignment method that directly optimize the policy from static preference data, and further improved by incorporating on-policy sampling~(i.e., preference candidates generated during the training loop) for better LM alignment. However, we show on-policy data is not always optimal, with systematic effectiveness difference emerging between static and on-policy preference candidates. For example, on-policy data can result in a $3\times$ effectiveness compared with static data for Llama-3, and a $0.4\times$ effectiveness for Zephyr. To explain the phenomenon, we propose the alignment stage assumption, which divides the alignment process into two distinct stages: the preference injection stage, which benefits from diverse data, and the preference fine-tuning stage, which favors high-quality data. Through theoretical and empirical analysis, we characterize these stages and propose an effective algorithm to identify the boundaries between them. We perform experiments on $5$ models~(Llama, Zephyr, Phi-2, Qwen, Pythia) and $2$ alignment methods~(DPO, SLiC-HF) to show the generalizability of alignment stage assumption and the effectiveness of the boundary measurement algorithm.

Context matching is not reasoning when performing generalized clinical evaluation of generative language models

npj Digital Medicine, Published online: 27 December 2025; doi:10.1038/s41746-025-02253-2

Context matching is not reasoning when performing generalized clinical evaluation of generative language models

MedDCR: Learning to Design Agentic Workflows for Medical Coding

arXiv:2511.13361v1 Announce Type: new Abstract: Medical coding converts free-text clinical notes into standardized diagnostic and procedural codes, which are essential for billing, hospital operations, and medical research. Unlike ordinary text classification, it requires multi-step reasoning: extracting diagnostic concepts, applying guideline constraints, mapping to hierarchical codebooks, and ensuring cross-document consistency. Recent advances leverage agentic LLMs, but most rely on rigid, manually crafted workflows that fail to capture the nuance and variability of real-world documentation, leaving open the question of how to systematically learn effective workflows. We present MedDCR, a closed-loop framework that treats workflow design as a learning problem. A Designer proposes workflows, a Coder executes them, and a Reflector evaluates predictions and provides constructive feedback, while a memory archive preserves prior designs for reuse and iterative refinement. On benchmark datasets, MedDCR outperforms state-of-the-art baselines and produces interpretable, adaptable workflows that better reflect real coding practice, improving both the reliability and trustworthiness of automated systems.

Prompt Engineering in Clinical Practice: Tutorial for Clinicians

Large language models (LLMs), such as OpenAI’s GPT series and Google’s PaLM, are transforming healthcare by improving clinical decision-making, enhancing patient communication, and simplifying administrative tasks. However, their performance relies heavily on prompt design, where small changes in wording or structure can greatly impact output quality. This poses a challenge for clinicians who are not experts in natural language processing (NLP). This tutorial combines prompt engineering techniques tailored for clinical use, covering methods like zero-shot, few-shot, chain-of-thought, and meta-prompting. We examine four critical dimensions (accuracy, bias mitigation, privacy protection, and workflow integration) through clinical case studies grounded in real-world practice. We provide actionable guidance on defining objectives, applying core principles, iteratively refining prompts, and integrating them into interoperable electronic health record (EHR) systems. This framework helps clinicians leverage LLMs to improve decision-making, streamline documentation, and enhance patient communication while maintaining ethical standards and ensuring patient safety.

Strategies for discovering novel hepatocellular carcinoma biomarkers

World J Hepatol. 2025 Feb 27;17(2):101201. doi: 10.4254/wjh.v17.i2.101201.

ABSTRACT

Liver cancer, particularly hepatocellular carcinoma (HCC), remains a significant global health challenge due to its high mortality rate and late-stage diagnosis. The discovery of reliable biomarkers is crucial for improving early detection and patient outcomes. This review provides a comprehensive overview of current and emerging biomarkers for HCC, including alpha-fetoprotein, des-gamma-carboxy prothrombin, glypican-3, Golgi protein 73, osteopontin, and microRNAs. Despite advancements, the diagnostic limitations of existing biomarkers underscore the urgent need for novel markers that can detect HCC in its early stages. The review emphasizes the importance of integrating multi-omics approaches, combining genomics, proteomics, and metabolomics, to develop more robust biomarker panels. Such integrative methods have the potential to capture the complex molecular landscape of HCC, offering insights into disease mechanisms and identifying targets for personalized therapies. The significance of large-scale validation studies, collaboration between research institutions and clinical settings, and consideration of regulatory pathways for clinical implementation is also discussed. In conclusion, while substantial progress has been made in biomarker discovery, continued research and innovation are essential to address the remaining challenges. The successful translation of these discoveries into clinical practice will require rigorous validation, standardization of protocols, and cross-disciplinary collaboration. By advancing the development and application of novel biomarkers, we can improve the early detection and management of HCC, ultimately enhancing patient survival and quality of life.

PMID:40027561 | PMC:PMC11866143 | DOI:10.4254/wjh.v17.i2.101201

Improving molecular subtypes and prognosis of pancreatic cancer through multi group analysis and machine learning

Discov Oncol. 2025 Jan 28;16(1):96. doi: 10.1007/s12672-025-01841-8.

ABSTRACT

BACKGROUND: Pancreatic cancer (PAC) has a complex tumor immune microenvironment, and currently, there is a lack of accurate personalized treatment. Establishing a novel consensus machine learning driven signature (CMLS) that offers a unique predictive model and possible treatment targets for this condition was the goal of this study.

METHODS: This study integrated multiple omics data of PAC patients, applied ten clustering techniques and ten machine learning approaches to construct molecular subtypes for PAC, and created a new CMLS.

RESULTS: Using multi-omics clustering, we discovered two cancer subtypes (CSs) associated with prognosis, among which CS1 exhibited poor prognostic outcomes. Subsequently, 13 central genes were identified through screening, constituting CMLS with a significant prognostic ability. The low CMLS group had a better prognosis and was more likely to possess a "hot" tumor phenotype. The prognosis for the high CMLS group was dismal. Still, the tumor mutation burden (TMB) and tumor neoantigen burden (TNB) levels in this group of patients were higher than in the low CMLS group, which were more favorable for immune therapy response.

CONCLUSION: This study emphasizes that CMLS provides a beneficial instrument for early prediction of patient prognosis and screening of probable patients appropriate for immunotherapy and has broad implications for clinical practice.

PMID:39873820 | PMC:PMC11775367 | DOI:10.1007/s12672-025-01841-8

Global trends and risk factors in gastric cancer: a comprehensive analysis of the Global Burden of Disease Study 2021 and multi-omics data

Int J Med Sci. 2025 Jan 1;22(2):341-356. doi: 10.7150/ijms.104437. eCollection 2025.

ABSTRACT

Background: Gastric cancer (GC) remains a significant global health challenge. This study aimed to comprehensively analyze GC epidemiology and risk factors to inform prevention and intervention strategies. Methods: We analyzed the Global Burden of Disease Study 2021 data, conducted 16 different machine learning (ML) models of NHANES data, performed Mendelian randomization (MR) studies on disease phenotypes, dietary preferences, microbiome, blood-based markers, and integrated differential gene expression and expression quantitative trait loci (eQTL) data from multiple cohorts to identify factors associated with GC risk. Results: Global age-standardized disability-adjusted life year rates (ASDR) for GC declined from 886.24 to 358.42 per 100,000 population between 1990 and 2030, with significant regional disparities. Despite this decline, total disability-adjusted life years show a concerning upward trend from 2015, rising from approximately 22.9 million to a projected 24.3 million by 2030. The slope index of inequality shifted from 87 in 1990 to -184 in 2021, indicating a reversal in GC burden distribution, with higher ASDR now associated with lower socio-demographic index countries. The ML models analysis identified higher levels of clinical characteristics such as phosphorus, calcium, eosinophils percent, and triglycerides, as well as lower levels of iron and monocyte percent, may be associated with an increased risk of GC. MR analyses revealed causal associations between GC risk and disease phenotypes such as Helicobacter pylori infection, chronic gastritis, obesity, depression, and dietary preferences such as dairy and processed meats. Gut microbiome analysis showed associations with microbiome such as Phascolarctobacterium and Ruminococcaceae species. Blood-based markers analysis identified protective and risk effects for cortisol, glutamate, nicotinamide, Natural Killer %lymphocyte, CD4-CD8- T cell Absolute Count, Phosphatidylcholine (16:0_18:1), and Interleukin-1-alpha. Integrated genomic analysis identified 10 genes significantly associated with GC risk, with strong evidence for colocalization in genes such as CCR6 and PILRB. Conclusions: This systematic analysis reveals complex global trends in GC burden and identifies novel clinical, disease phenotypes, dietary preferences, microbial, blood-based, and genetic risk factors. These findings provide potential targets for improved risk stratification, prevention, and intervention strategies to reduce the global burden of GC.

PMID:39781526 | PMC:PMC11704698 | DOI:10.7150/ijms.104437

Genome-wide characterization of circulating metabolic biomarkers

Nature, Published online: 06 March 2024; doi:10.1038/s41586-024-07148-y

A meta-analysis of genome-wide association studies for 233 circulating metabolites from 33 cohorts reveals more than 400 loci and suggests probable causal genes, providing insights into metabolic pathways and disease aetiology.

A pan-cancer single-cell panorama of human natural killer cells

Integrative single-cell RNA sequencing analyses on natural killer (NK) cells from over 700 patients across 24 tumor types depict shared and tumor-type-specific NK cell features and highlight the potential of specific myeloid cell subpopulations in regulating NK cell anti-tumor function.
❌