❌

Normal view

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.

Integrating deep learning and multi-omics features in radiation pneumonitis prediction for lung cancer patients using PET/CT

BMC Med Imaging. 2025 Oct 27;25(1):426. doi: 10.1186/s12880-025-01971-z.

ABSTRACT

BACKGROUND: To investigate the feasibility and accuracy of PET radiomics features, along with their combination with CT radiomics, dosiomics, and deep learning (DL) features, in predicting radiation pneumonitis (RP) in lung cancer patients treated with volumetric modulated arc therapy (VMAT).

METHODS: A total of 206 and 27 lung cancer patients who underwent VMAT with pre-treatment PET/CT imaging were enrolled from Hospital One and Hospital Two for model training and external validation, respectively. Four machine learning (ML) methods were applied to build radiomics models with features extracted from CT (R_CT), PET (R_PET), radiomics features fused PET/CT (R_fFU) and fused PET/CT images (R_ iFU), as well dosiomics features (D). Three DL models were built to extract features from PET (DL_PET), CT (DL_CT), and fused PET/CT images (DL_FU). The best-performing radiomics and DL models were combined with dosiomics to create the final joint model. ROC curves with AUC, accuracy, sensitivity, and specificity evaluated the performance. A nomogram was constructed using top-performing model features, parameters, and relevant clinical factors.

RESULTS: The extreme gradient boosting (XGBoost) and 18-layer residual neural network (Resnet-18) achieved the best performance. The R+D+DL model combined radiomics, dosiomics, and DL features achieved AUCs of 0.93, 0.92 and 0.89 in the training, internal validaiton and external validation cohorts, respectively. A nomogram constructed with gender, Adaptive RT, SUVp90, and XGBoost-score achieved an AUC of 0.94 for RP prediction in VMAT-treated lung cancer patients using PET/CT.

CONCLUSION: Integrating radiomics, DL, dosiomics features and SUVp90 is promising in the RP prediction for lung cancer patients underwent VMAT using PET/CT images.

PMID:41146084 | DOI:10.1186/s12880-025-01971-z

❌