❌

Normal view

Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering

arXiv:2511.15061v1 Announce Type: new Abstract: Genomic question answering often requires complex reasoning and integration across diverse biomedical sources. GeneGPT addressed this challenge by combining domain-specific APIs with OpenAI's code-davinci-002 large language model to enable natural language interaction with genomic databases. However, its reliance on a proprietary model limits scalability, increases operational costs, and raises concerns about data privacy and generalization. In this work, we revisit and reproduce GeneGPT in a pilot study using open source models, including Llama 3.1, Qwen2.5, and Qwen2.5 Coder, within a monolithic architecture; this allows us to identify the limitations of this approach. Building on this foundation, we then develop OpenBioLLM, a modular multi-agent framework that extends GeneGPT by introducing agent specialization for tool routing, query generation, and response validation. This enables coordinated reasoning and role-based task execution. OpenBioLLM matches or outperforms GeneGPT on over 90% of the benchmark tasks, achieving average scores of 0.849 on Gene-Turing and 0.830 on GeneHop, while using smaller open-source models without additional fine-tuning or tool-specific pretraining. OpenBioLLM's modular multi-agent design reduces latency by 40-50% across benchmark tasks, significantly improving efficiency without compromising model capability. The results of our comprehensive evaluation highlight the potential of open-source multi-agent systems for genomic question answering. Code and resources are available at https://github.com/ielab/OpenBioLLM.

Diagnosing Hallucination Risk in AI Surgical Decision-Support: A Sequential Framework for Sequential Validation

arXiv:2511.00588v1 Announce Type: cross Abstract: Large language models (LLMs) offer transformative potential for clinical decision support in spine surgery but pose significant risks through hallucinations, which are factually inconsistent or contextually misaligned outputs that may compromise patient safety. This study introduces a clinician-centered framework to quantify hallucination risks by evaluating diagnostic precision, recommendation quality, reasoning robustness, output coherence, and knowledge alignment. We assessed six leading LLMs across 30 expert-validated spinal cases. DeepSeek-R1 demonstrated superior overall performance (total score: 86.03 $\pm$ 2.08), particularly in high-stakes domains such as trauma and infection. A critical finding reveals that reasoning-enhanced model variants did not uniformly outperform standard counterparts: Claude-3.7-Sonnet's extended thinking mode underperformed relative to its standard version (80.79 $\pm$ 1.83 vs. 81.56 $\pm$ 1.92), indicating extended chain-of-thought reasoning alone is insufficient for clinical reliability. Multidimensional stress-testing exposed model-specific vulnerabilities, with recommendation quality degrading by 7.4% under amplified complexity. This decline contrasted with marginal improvements in rationality (+2.0%), readability (+1.7%) and diagnosis (+4.7%), highlighting a concerning divergence between perceived coherence and actionable guidance. Our findings advocate integrating interpretability mechanisms (e.g., reasoning chain visualization) into clinical workflows and establish a safety-aware validation framework for surgical LLM deployment.

A full life cycle biological clock based on routine clinical data and its impact in health and diseases

Nature Medicine, Published online: 27 October 2025; doi:10.1038/s41591-025-04006-w

The biological clock model LifeClock predicts biological age across all life stages from routine clinical data, revealing distinct pediatric and adult disease risk patterns.

Development and validation of an integrative 54 biomarker-based risk identification model for multi-cancer in 42,666 individuals: a population-based prospective study to guide advanced screening strategies

Biomark Res. 2025 Aug 11;13(1):101. doi: 10.1186/s40364-025-00812-z.

ABSTRACT

BACKGROUND: Early identification of high-risk individuals is crucial for optimizing cancer screening, particularly when considering expensive and invasive methods such as multi-omics technologies and endoscopic procedures. However, developing a robust, practical multi-cancer risk prediction model that integrates diverse, multi-scale data and with proper validation remains a significant challenge.

METHODS: We initialized the FuSion study by recruiting 42,666 participants from Taizhou, China, with a discovery cohort (n = 16,340) and an independent validation cohort (n = 26,308) after exclusion criteria. We integrated multi-scale data from 54 blood-derived biomarkers and 26 epidemiological exposures to develop a risk prediction model for five common cancers, including lung, esophageal, liver, gastric, and colorectal cancer. Employing five supervised machine learning approaches, we used a LASSO-based feature selection strategy to identify the most informative predictors. The model was trained and internally validated in the discovery cohort, externally applied in the validation cohort, and further evaluated through a prospective clinical follow-up to assess cancer events via clinical examinations.

RESULTS: The final model comprising four key biomarkers along with age, sex, and smoking intensity, achieving an AUROC of 0.767 (95% CI: 0.723-0.814) for five-year risk prediction. High-risk individuals (17.19% of the cohort) accounted for 50.42% of incident cancer cases, with a 15.19-fold increased risk compared to the low-risk group. During follow-up of 2,863 high-risk subjects, 9.64% were newly diagnosed with cancer or precancerous lesions. Notably, cancer detection in the high-risk group was 5.02 times higher than in the low-risk group and 1.74 times higher than in the intermediate-risk group. In particular, the incidence of esophageal cancers in the high-risk group was 16.84 times that of the low-risk group.

CONCLUSIONS: This is the first population-based prospective study in a large Chinese cohort that leverage multi-scale data including biomarkers for multi-cancer risk prediction. Our effective risk stratification model not only enhances early cancer detection but also lays the foundation for the targeted application of advanced screening methods, including but not limited to multi-omics technologies and endoscopy. These findings support precision prevention strategies and the optimal allocation of healthcare resources.

PMID:40790537 | PMC:PMC12341305 | DOI:10.1186/s40364-025-00812-z

Synthetic lethality of mRNA quality control complexes in cancer

Nature, Published online: 05 February 2025; doi:10.1038/s41586-024-08398-6

PELO–HBS1L and SKI complexes in the human mRNA quality control pathway exhibit a synthetic lethal interaction and may represent novel targets for the development of cancer therapies.
❌