❌

Normal view

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

arXiv:2508.10530v2 Announce Type: replace Abstract: The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize the expected reward that reflects human preferences. Recently, Direct Preference Optimization~(DPO) was proposed as a LM alignment method that directly optimize the policy from static preference data, and further improved by incorporating on-policy sampling~(i.e., preference candidates generated during the training loop) for better LM alignment. However, we show on-policy data is not always optimal, with systematic effectiveness difference emerging between static and on-policy preference candidates. For example, on-policy data can result in a $3\times$ effectiveness compared with static data for Llama-3, and a $0.4\times$ effectiveness for Zephyr. To explain the phenomenon, we propose the alignment stage assumption, which divides the alignment process into two distinct stages: the preference injection stage, which benefits from diverse data, and the preference fine-tuning stage, which favors high-quality data. Through theoretical and empirical analysis, we characterize these stages and propose an effective algorithm to identify the boundaries between them. We perform experiments on $5$ models~(Llama, Zephyr, Phi-2, Qwen, Pythia) and $2$ alignment methods~(DPO, SLiC-HF) to show the generalizability of alignment stage assumption and the effectiveness of the boundary measurement algorithm.

Companion Agents: A Table-Information Mining Paradigm for Text-to-SQL

arXiv:2601.08838v1 Announce Type: cross Abstract: Large-scale Text-to-SQL benchmarks such as BIRD typically assume complete and accurate database annotations as well as readily available external knowledge, which fails to reflect common industrial settings where annotations are missing, incomplete, or erroneous. This mismatch substantially limits the real-world applicability of state-of-the-art (SOTA) Text-to-SQL systems. To bridge this gap, we explore a database-centric approach that leverages intrinsic, fine-grained information residing in relational databases to construct missing evidence and improve Text-to-SQL accuracy under annotation-scarce conditions. Our key hypothesis is that when a query requires multi-step reasoning over extensive table information, existing methods often struggle to reliably identify and utilize the truly relevant knowledge. We therefore propose to "cache" query-relevant knowledge on the database side in advance, so that it can be selectively activated at inference time. Based on this idea, we introduce Companion Agents (CA), a new Text-to-SQL paradigm that incorporates a group of agents accompanying database schemas to proactively mine and consolidate hidden inter-table relations, value-domain distributions, statistical regularities, and latent semantic cues before query generation. Experiments on BIRD under the fully missing evidence setting show that CA recovers +4.49 / +4.37 / +14.13 execution accuracy points on RSL-SQL / CHESS / DAIL-SQL, respectively, with larger gains on the Challenging subset +9.65 / +7.58 / +16.71. These improvements stem from CA's automatic database-side mining and evidence construction, suggesting a practical path toward industrial-grade Text-to-SQL deployment without reliance on human-curated evidence.

LOSTdb: a manually curated multi-omics database for lung cancer research

BMC Bioinformatics. 2025 Dec 3;26(1):290. doi: 10.1186/s12859-025-06319-6.

ABSTRACT

Lung cancer is one of the most prevalent malignant tumors with high morbidity and mortality rates worldwide. Extensive multi-omics analyses have revealed significant intratumoral heterogeneity even within the same histopathological subtype. However, a database that systematically integrates multi-omics data for lung cancer research has long been lacking. Here, we developed LOSTdb, a molecular subtype annotation system for lung cancer that integrates multi-omics data and metadata. LOSTdb comprises 295 multi-omics datasets, including bulk RNA-seq, genomic, proteomic, methylation, and scRNA-seq data, with over 10,000 manually curated metadata entries. This resource encompasses high-quality clinical specimens, mouse models, and cell lines, totaling 34,393 samples and more than 1.2 million single cells. Each omics sample was annotated with both literature-based classical subtypes and NMF-derived meta-program (MP) subtypes. The platform supports cross-searching of omics and metadata at the gene and dataset levels, offers multiple visualization and analysis methods, and includes five tool modules, enabling functions such as integrated analysis, significance analysis between metadata as well as between genes and metadata, and target prediction for lung cancer molecular subtypes, serving as an essential tool for lung cancer precision medicine. LOSTdb is a user-friendly interactive database freely accessible at http://lostdbcancer.com:8080 .

PMID:41339793 | DOI:10.1186/s12859-025-06319-6

LOSTdb: a manually curated multi-omics database for lung cancer research

3 December 2025 at 19:00

BMC Bioinformatics. 2025 Dec 3;26(1):290. doi: 10.1186/s12859-025-06319-6.

ABSTRACT

Lung cancer is one of the most prevalent malignant tumors with high morbidity and mortality rates worldwide. Extensive multi-omics analyses have revealed significant intratumoral heterogeneity even within the same histopathological subtype. However, a database that systematically integrates multi-omics data for lung cancer research has long been lacking. Here, we developed LOSTdb, a molecular subtype annotation system for lung cancer that integrates multi-omics data and metadata. LOSTdb comprises 295 multi-omics datasets, including bulk RNA-seq, genomic, proteomic, methylation, and scRNA-seq data, with over 10,000 manually curated metadata entries. This resource encompasses high-quality clinical specimens, mouse models, and cell lines, totaling 34,393 samples and more than 1.2 million single cells. Each omics sample was annotated with both literature-based classical subtypes and NMF-derived meta-program (MP) subtypes. The platform supports cross-searching of omics and metadata at the gene and dataset levels, offers multiple visualization and analysis methods, and includes five tool modules, enabling functions such as integrated analysis, significance analysis between metadata as well as between genes and metadata, and target prediction for lung cancer molecular subtypes, serving as an essential tool for lung cancer precision medicine. LOSTdb is a user-friendly interactive database freely accessible at http://lostdbcancer.com:8080 .

PMID:41339793 | PMC:PMC12676782 | DOI:10.1186/s12859-025-06319-6

Identifying specific functional roles for senescence across cell types

A dual recombinase-mediated genetic system for cell-type-specific lineage tracing, ablation, and gene manipulation of senescent cells reveals distinct roles of senescence across cell types.
❌