❌

Reading view

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.
  •  

Integration of multi-omics profiling reveals an epigenetic-based molecular classification of lung adenocarcinoma: implications for drug sensitivity and immunotherapy response prediction

Front Pharmacol. 2025 Feb 19;16:1540477. doi: 10.3389/fphar.2025.1540477. eCollection 2025.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) remains a major cause of cancer-related mortality worldwide, with high heterogeneity and poor prognosis. Epigenetic dysregulation plays a crucial role in LUAD progression, yet its potential in molecular classification and therapeutic prediction remains largely unexplored.

METHODS: We performed an integrated multi-omics analysis of 432 LUAD patients from TCGA and 398 patients from GEO datasets. Using consensus clustering and random survival forest (RSF) algorithms, we established an epigenetic-based molecular classification system and constructed a prognostic model. The model's performance was validated in multiple independent cohorts, and its biological implications were investigated through comprehensive functional analyses.

RESULTS: We identified two distinct molecular subtypes (CS1 and CS2) with significant differences in epigenetic modification patterns, immune microenvironment, and clinical outcomes (P = 0.005). The RSF-based prognostic model demonstrated robust performance in both training (TCGA-LUAD) and validation (GSE72094) cohorts, with time-dependent AUC values ranging from 0.625 to 0.694. Low-risk patients exhibited enhanced immune cell infiltration, particularly CD8+ T cells and M1 macrophages, and showed better responses to immune checkpoint inhibitors. Drug sensitivity analysis revealed subtype-specific therapeutic vulnerabilities, with low-risk patients showing higher sensitivity to conventional chemotherapy and targeted therapy.

CONCLUSION: Our study establishes a novel epigenetic-based classification system and predictive model for LUAD, providing valuable insights into patient stratification and personalized treatment selection. The model's ability to predict immunotherapy response and drug sensitivity offers practical guidance for clinical decision-making, potentially improving patient outcomes through precision medicine approaches.

PMID:40046740 | PMC:PMC11879945 | DOI:10.3389/fphar.2025.1540477

  •  
❌