❌

Normal view

aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists

arXiv:2508.15126v2 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have enabled AI agents to autonomously generate scientific proposals, conduct experiments, author papers, and perform peer reviews. Yet this flood of AI-generated research content collides with a fragmented and largely closed publication ecosystem. Traditional journals and conferences rely on human peer review, making them difficult to scale and often reluctant to accept AI-generated research content; existing preprint servers (e.g. arXiv) lack rigorous quality-control mechanisms. Consequently, a significant amount of high-quality AI-generated research lacks appropriate venues for dissemination, hindering its potential to advance scientific progress. To address these challenges, we introduce aiXiv, a next-generation open-access platform for human and AI scientists. Its multi-agent architecture allows research proposals and papers to be submitted, reviewed, and iteratively refined by both human and AI scientists. It also provides API and MCP interfaces that enable seamless integration of heterogeneous human and AI scientists, creating a scalable and extensible ecosystem for autonomous scientific discovery. Through extensive experiments, we demonstrate that aiXiv is a reliable and robust platform that significantly enhances the quality of AI-generated research proposals and papers after iterative revising and reviewing on aiXiv. Our work lays the groundwork for a next-generation open-access ecosystem for AI scientists, accelerating the publication and dissemination of high-quality AI-generated research content. Code: https://github.com/aixiv-org aiXiv: https://aixiv.science

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.

Developing an Evaluation System for Quality of Health Educational Short Videos on Social Media (LassVQ) Using Nominal Group Technique and Analytic Hierarchy Process: Qualitative Study

Background: With the increasing use of social media platforms for health communication, the quality of health educational short videos (HESVs) has become a key concern. However, no standardized framework exists to evaluate the quality of health videos on social media, highlighting the need for a comprehensive evaluation system. Objective: The aim of this study is to develop a valid and structured evaluation tool for assessing the quality of HESVs on social media. Methods: The initial evaluation indicators obtained from the literature review and brainstorming undertaken in the study group were provided to the nominal group reference Lasswell’s 5W communication model, and two rounds of nominal group technique (NGT) were carried out to screen, add, revise, and adjust indicators, and reach a consensus of evaluation system. The indicators were then ranked based on their significance, as scored by the experts using the analytic hierarchy process. The content validity was assessed by experts who rated the relevance of each indicator on a 4-point Likert scale. Results: The primary indicators include communicator, communication content, communication channel, and communication effect, along with 13 secondary indicators and 34 tertiary indicators. 11 experts were enrolled in the NGT, 45% of experts had a doctoral degree, 80% of them were ranked associate professor or professor. The average familiarity coefficient of each key indicator of the NGT was 0.85. The average values of the expert judgment coefficient and authority coefficient were 0.93 and 0.85, respectively. In Round 1 of NGT, the “Communication target” of 5 primary indicators, 7 of 20 secondary indicators, and 66 of 94 tertiary indicators did not reach a consensus, and therefore, they were not deleted and will proceed to the next round of NGT. In Round 2 NGT, 1 primary indicator, 7 secondary indicators, and 59 tertiary indicators were deleted based on the consensus criteria. After the two rounds of NGT, 4 primary indicators, 13 secondary indicators, and 34 tertiary indicators finally reached a consensus. Among primary indicators, communication content was found to be the most influential, accounting for 45.68%. Among secondary indicators, credibility, scientificity, availability, and social attention were the most influential indicators, with priorities of 56.67%, 24.26%, 74.62%, and 39.89% in their respective categories. Among tertiary indicators, ‘Become a hot search recommended by the platform’ was the most influential indicator with a weight of 0.07. The content validity of all the evaluation indicators were 0.73 – 1.0, and the scale-level content validity index (average) was 0.87, which was indicated as acceptable. Conclusions: The evaluation system for the quality of HESVs on social media (LassVQ) was developed, and its validity was acceptable. The proposed evaluation system can be used in conjunction with qualitative methods to gain a holistic perspective on the multidimensional quality of HESVs on social media.
❌