❌

Reading view

An integrated bioinformatics and multi-omics investigation of the sirtuin family to identify their prognostic importance in human cancers

Tumour Biol. 2025 Jan-Dec;47:14230380251410470. doi: 10.1177/14230380251410470. Epub 2025 Dec 24.

ABSTRACT

BackgroundIn recent years, the significance of sirtuins in cancer biology has become increasingly evident, but their molecular mechanisms and prognostic impacts remain elusive.ObjectiveThe present study aimed to investigate the differential expression of the sirtuin gene family across cancers and to evaluate their prognostic value.MethodsWe used various bioinformatics databases and methodologies, including Oncomine, GEPIA, OncoDB, cBioPortal, R2 Kaplan-Meier Scanner, STRING, etc., to determine the expression pattern of the sirtuin family genes, along with their mutations and prognostic values in human cancers.ResultsIn the current study, SIRT1, SIRT2, SIRT4, and SIRT5 were downregulated in lymphoma, whereas SIRT6 and SIRT7 were overexpressed. In breast cancer, SIRT3, SIRT5, and SIRT7 were overexpressed, and in terms of kidney cancer, higher expression of SIRT2, SIRT3, and SIRT5 was observed. In contrast, for leukemia, bladder, and brain cancers, most sirtuin family members showed reduced expression. We found that most mutations occurred in uterine cancer, chRCC (chromophobe renal cell carcinoma), DLBCL (diffuse large B-cell lymphoma), melanoma, pRCC (papillary renal cell carcinoma), and esophageal cancer. Moreover, we identified the relevant functional proteins through protein-protein interaction analysis to evaluate copy number alterations (CNAs) in sirtuins. The most frequent alterations were amplifications and deep deletions. Survival analysis demonstrated that SIRT1 and SIRT2 overexpression correlated with improved overall survival in low-grade glioma but predicted poorer outcomes in ovarian cancer. Downregulation of SIRT1, SIRT3, and SIRT5 was associated with better prognosis in DLBCL, while SIRT3 and SIRT4 upregulation predicted favorable survival in testicular germ cell tumors. SIRT6 overexpression was linked to favorable prognosis in esophageal carcinoma and sarcoma, while unfavorable outcomes were observed in hepatocellular carcinoma and cholangiocarcinoma. SIRT7 upregulation was significantly associated with reduced survival in esophageal, liver, and uterine cancers, but surprisingly correlated with improved outcomes in urothelial carcinoma and cervical squamous cell carcinoma.ConclusionsTogether, this multi-omics analysis reveals the correlation and prognostic values of sirtuins across multiple types of human cancers and suggests that sirtuins may serve as promising biomarkers for different cancers.

PMID:41439701 | DOI:10.1177/14230380251410470

  •  

Evaluating Peer Online Forums to Support Health: Ethical and Practical Challenges

Many people use peer online forums to seek support for health-related problems. More research is needed to understand the impacts of forum use, and how these are generated. However, there are significant ethical and practical challenges with the methods available to do the required research. We examine the key challenges associated with conducting each of the most commonly used online data collection methods: surveys, interviews, forum post analysis; and triangulation of these methods. Based on our learning from the Improving Peer Online Forums (iPOF) study, an inter-disciplinary realist informed mixed methods evaluation of peer online forums, we outline strategies that can be used to address key issues pertaining to assessing important outcomes, facilitating participation, validating participants (users who consent to take part in one or more parts of the study), protecting anonymity, gaining consent, managing risk, multi-stakeholder engagement, and triangulation. We share this learning to support researchers, reviewers, and ethics committees faced with deciding how best to address these challenges. We highlight the need for ongoing open, transparent discussion to ensure the research field keeps pace with evolving technology design and societal attitudes to online data use.
  •  
  •  

Comparison of liquid biopsy-based technologies for cancer screening

Crit Rev Clin Lab Sci. 2025 Dec 27:1-12. doi: 10.1080/10408363.2025.2606357. Online ahead of print.

ABSTRACT

Circulating plasma DNA has found important applications in diverse medical fields, including prenatal testing, transplantation, and especially cancer. Many companies have developed products for detecting minimal residual disease, selecting or monitoring therapy, assessing prognosis, and confirming diagnosis. One major application is in screening asymptomatic individuals for the presence of cancer. Screening may facilitate better clinical outcomes through earlier interventions. Collectively, these technologies are widely known as "liquid biopsies". After the extraction of free DNA from the circulation, it is analyzed by various molecular techniques to explore differences between DNA originating from normal cells and cancer cells. Circulating plasma DNA originating from tumors (ctDNA) is expected to harbor the same molecular changes as tumor tissue itself. Thus, ctDNA is considered a surrogate of cancer tissue, but without the need to perform invasive biopsies to obtain it. Many new diagnostic companies have taken advantage of this new biomarker and developed technologies for screening for one or multiple cancers. We previously estimated the amount of ctDNA in circulation, which is admixed with DNA originating from normal cells. We concluded that since only a small fraction of the whole plasma (3 liters) is used for testing (3 to 4 mL), it is possible that the retrieved ctDNA may not be enough for cancer diagnosis in all patients. This problem is more acute with small tumors. Here, we mention some companies in the "liquid biopsy" arena and analyze their clinical data to establish if their tests are close to entering the clinic. We conclude from this analysis that current data do not support the use of these technologies for population screening due to many false negative and false positive results.

PMID:41454842 | DOI:10.1080/10408363.2025.2606357

  •  

Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning

arXiv:2512.15943v1 Announce Type: new Abstract: As organizations scale adoption of generative AI, model cost optimization and operational efficiency have emerged as critical factors determining sustainability and accessibility. While Large Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, their extensive computational requirements make them cost-prohibitive for routine enterprise use. This limitation motivates the exploration of Small Language Models (SLMs), which can deliver comparable performance in targeted applications while drastically reducing infrastructure overhead (Irugalbandara et al., 2023). In this work, we investigate the feasibility of replacing LLM-driven workflows with optimized SLMs. We trained a domain-adapted SLM to execute representative tasks traditionally handled by LLMs, such as document summarization, query answering, and structured data interpretation. As part of the experiment, we investigated the fine-tuning of facebook/opt-350m model (single epoch only) using the Hugging Face TRL (Transformer Reinforcement Learning), specifically the Supervised Fine-Tuning (SFT) trainer. The OPT-350M model was released by Meta AI in 2022 as part of the OPT (Open Pretrained Transformer) family of models. Similar studies demonstrate that even models at the 350M parameter scale can meaningfully contribute to instruction-tuning pipelines (Mekala et al., 2024). Experimental results demonstrated that our fine-tuned SLM achieves exceptional performance with a 77.55\% pass rate on ToolBench evaluation, significantly outperforming all baseline models including ChatGPT-CoT (26.00\%), ToolLLaMA-DFS (30.18\%), and ToolLLaMA-CoT (16.27\%). These findings emphasize that thoughtful design and targeted training of SLMs can significantly lower barriers to adoption, enabling cost-effective, large-scale integration of generative AI into production systems.
  •  

Distributional AGI Safety

arXiv:2512.16856v1 Announce Type: new Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic Artificial General Intelligence (AGI). The alternative AGI emergence hypothesis, where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills and affordances, has received far less attention. Here we argue that this patchwork AGI hypothesis needs to be given serious consideration, and should inform the development of corresponding safeguards and mitigations. The rapid deployment of advanced AI agents with tool-use capabilities and the ability to communicate and coordinate makes this an urgent safety consideration. We therefore propose a framework for distributional AGI safety that moves beyond evaluating and aligning individual agents. This framework centers on the design and implementation of virtual agentic sandbox economies (impermeable or semi-permeable), where agent-to-agent transactions are governed by robust market mechanisms, coupled with appropriate auditability, reputation management, and oversight to mitigate collective risks.
  •  

Emergent Bias and Fairness in Multi-Agent Decision Systems

arXiv:2512.16433v1 Announce Type: cross Abstract: Multi-agent systems have demonstrated the ability to improve performance on a variety of predictive tasks by leveraging collaborative decision making. However, the lack of effective evaluation methodologies has made it difficult to estimate the risk of bias, making deployment of such systems unsafe in high stakes domains such as consumer finance, where biased decisions can translate directly into regulatory breaches and financial loss. To address this challenge, we need to develop fairness evaluation methodologies for multi-agent predictive systems and measure the fairness characteristics of these systems in the financial tabular domain. Examining fairness metrics using large-scale simulations across diverse multi-agent configurations, with varying communication and collaboration mechanisms, we reveal patterns of emergent bias in financial decision-making that cannot be traced to individual agent components, indicating that multi-agent systems may exhibit genuinely collective behaviors. Our findings highlight that fairness risks in financial multi-agent systems represent a significant component of model risk, with tangible impacts on tasks such as credit scoring and income estimation. We advocate that multi-agent decision systems must be evaluated as holistic entities rather than through reductionist analyses of their constituent components.
  •  

AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research

arXiv:2512.16455v1 Announce Type: cross Abstract: In this paper, we describe a federated compute platform dedicated to support Artificial Intelligence in scientific workloads. Putting the effort into reproducible deployments, it delivers consistent, transparent access to a federation of physically distributed e-Infrastructures. Through a comprehensive service catalogue, the platform is able to offer an integrated user experience covering the full Machine Learning lifecycle, including model development (with dedicated interactive development environments), training (with GPU resources, annotation tools, experiment tracking, and federated learning support) and deployment (covering a wide range of deployment options all along the Cloud Continuum). The platform also provides tools for traceability and reproducibility of AI models, integrates with different Artificial Intelligence model providers, datasets and storage resources, allowing users to interact with the broader Machine Learning ecosystem. Finally, it is easily customizable to lower the adoption barrier by external communities.
  •  

From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs

arXiv:2512.16795v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds large language models (LLMs) in external evidence, but fails when retrieved sources conflict or contain outdated or subjective information. Prior work address these issues independently but lack unified reasoning supervision. We propose a reasoning-trace-augmented RAG framework that adds structured, interpretable reasoning across three stages : (1) document-level adjudication, (2) conflict analysis, and (3) grounded synthesis, producing citation-linked answers or justified refusals. A Conflict-Aware Trust-Score (CATS) pipeline is introduced which evaluates groundedness, factual correctness, refusal accuracy, and conflict-behavior alignment using an LLM-as-a-Judge. Our 539-query reasoning dataset and evaluation pipeline establish a foundation for conflict-aware, interpretable RAG systems. Experimental results demonstrate substantial gains over baselines, most notably with Qwen, where Supervised Fine-Tuning improved End-to-End answer correctness from 0.069 to 0.883 and behavioral adherence from 0.074 to 0.722.
  •  

Constitutional Law and AI Governance: Constraints on Model Licensing and Research Classification

arXiv:2509.05361v2 Announce Type: replace-cross Abstract: Transformative AI systems may pose unprecedented catastrophic risks, but the U.S. Constitution places significant constraints on the government's ability to govern this technology. This paper examines how the First Amendment, administrative law, and the Fourteenth Amendment shape the legal vulnerability of two regulatory proposals: model licensing and AI research classification. While the First Amendment may provide some degree of protection for model algorithms or outputs, this protection does not foreclose regulation. Policymakers must also consider administrative legal requirements, due to both agency review and authority. Finally, while substantive due process and equal protection pose minimal obstacles, procedural due process requires the government to clearly define when developers vest a legal interest in their models. Given this analysis, effective AI governance requires careful implementation to avoid these legal challenges.
  •  

First, do NOHARM: towards clinically safe large language models

arXiv:2512.01241v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a benchmark using 100 real primary care-to-specialist consultation cases to measure frequency and severity of harm from LLM-generated medical recommendations. NOHARM covers 10 specialties, with 12,747 expert annotations for 4,249 clinical management options. Across 31 LLMs, potential for severe harm from LLM recommendations occurs in up to 22.2% (95% CI 21.6-22.8%) of cases, with harm of omission accounting for 76.6% (95% CI 76.4-76.8%) of errors. Safety performance is only moderately correlated (r = 0.61-0.64) with existing AI and medical knowledge benchmarks. The best models outperform generalist physicians on safety (mean difference 9.7%, 95% CI 7.0-12.5%), and a diverse multi-agent approach improves safety compared to solo models (mean difference 8.0%, 95% CI 4.0-12.1%). Therefore, despite strong performance on existing evaluations, widely used AI models can produce severely harmful medical advice at nontrivial rates, underscoring clinical safety as a distinct performance dimension necessitating explicit measurement.
  •  

A Decision-Theoretic Approach for Managing Misalignment

arXiv:2512.15584v1 Announce Type: new Abstract: When should we delegate decisions to AI systems? While the value alignment literature has developed techniques for shaping AI values, less attention has been paid to how to determine, under uncertainty, when imperfect alignment is good enough to justify delegation. We argue that rational delegation requires balancing an agent's value (mis)alignment with its epistemic accuracy and its reach (the acts it has available). This paper introduces a formal, decision-theoretic framework to analyze this tradeoff precisely accounting for a principal's uncertainty about these factors. Our analysis reveals a sharp distinction between two delegation scenarios. First, universal delegation (trusting an agent with any problem) demands near-perfect value alignment and total epistemic trust, conditions rarely met in practice. Second, we show that context-specific delegation can be optimal even with significant misalignment. An agent's superior accuracy or expanded reach may grant access to better overall decision problems, making delegation rational in expectation. We develop a novel scoring framework to quantify this ex ante decision. Ultimately, our work provides a principled method for determining when an AI is aligned enough for a given context, shifting the focus from achieving perfect alignment to managing the risks and rewards of delegation under uncertainty.
  •  

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

arXiv:2512.14896v1 Announce Type: cross Abstract: Objectives: To evaluate large language model (LLM) performance on pharmacy licensure-style question-answering (QA) tasks and develop an external knowledge integration method to improve their accuracy. Methods: We benchmarked eleven existing LLMs with varying parameter sizes (8 billion to 70+ billion) using a 141-question pharmacy dataset. We measured baseline accuracy for each model without modification. We then developed a three-step retrieval-augmented generation (RAG) pipeline, DrugRAG, that retrieves structured drug knowledge from validated sources and augments model prompts with evidence-based context. This pipeline operates externally to the models, requiring no changes to model architecture or parameters. Results: Baseline accuracy ranged from 46% to 92%, with GPT-5 (92%) and o3 (89%) achieving the highest scores. Models with fewer than 8 billion parameters scored below 50%. DrugRAG improved accuracy across all tested models, with gains ranging from 7 to 21 percentage points (e.g., Gemma 3 27B: 61% to 71%, Llama 3.1 8B: 46% to 67%) on the 141-item benchmark. Conclusion: We demonstrate that external structured drug knowledge integration through DrugRAG measurably improves LLM accuracy on pharmacy tasks without modifying the underlying models. This approach provides a practical pipeline for enhancing pharmacy-focused AI applications with evidence-based information.
  •  

MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models

arXiv:2506.07400v3 Announce Type: replace-cross Abstract: The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to medical imaging remains challenging due to hallucinations, limited interpretability, and insufficient domain-specific medical knowledge, which can potentially reduce clinical accuracy. Although recent approaches combining imaging models with LLM reasoning have improved reporting, they typically rely on a single generalist agent, restricting their capacity to emulate the diverse and complex reasoning found in multidisciplinary medical teams. To address these limitations, we propose MedChat, a multi-agent diagnostic framework and platform that combines specialized vision models with multiple role-specific LLM agents, all coordinated by a director agent. This design enhances reliability, reduces hallucination risk, and enables interactive diagnostic reporting through an interface tailored for clinical review and educational use. Code available at https://github.com/Purdue-M2/MedChat.
  •  

Multimodal Foundation Models for Early Disease Detection

arXiv:2510.01899v2 Announce Type: replace-cross Abstract: Healthcare data now span EHRs, medical imaging, genomics, and wearable sensors, but most diagnostic models still process these modalities in isolation. This limits their ability to capture early, cross-modal disease signatures. This paper introduces a multimodal foundation model built on a transformer architecture that integrates heterogeneous clinical data through modality-specific encoders and cross-modal attention. Each modality is mapped into a shared latent space and fused using multi-head attention with residual normalization. We implement the framework using a multimodal dataset that simulates early-stage disease patterns across EHR sequences, imaging patches, genomic profiles, and wearable signals, including missing-modality scenarios and label noise. The model is trained using supervised classification together with self-supervised reconstruction and contrastive alignment to improve robustness. Experimental evaluation demonstrates strong performance in early-detection settings, with stable classification metrics, reliable uncertainty estimates, and interpretable attention patterns. The approach moves toward a flexible, pretrain-and-fine-tune foundation model that supports precision diagnostics, handles incomplete inputs, and improves early disease detection across oncology, cardiology, and neurology applications.
  •  

A novel statistical feature selection framework for biomarker discovery and cancer classification via multiomics integration

BMC Med Res Methodol. 2025 Dec 17. doi: 10.1186/s12874-025-02713-z. Online ahead of print.

ABSTRACT

BACKGROUND: Early cancer diagnosis is essential for improving prognosis and guiding treatment. However, the high dimensionality and complexity of omics data present major challenges. Computational approaches that extract stable biomarkers and enable reliable classification across cancer types and stages are needed.

METHODS: A novel feature selection method, sDCFE (synergistic Discriminative Cluster-based Feature Extraction), was developed by extending Fisher-like variance analysis with a median absolute deviation (MAD) regularization term and a cluster separation component to enhance robustness and interpretability. Features selected by sDCFE were compared with those obtained from XGBoost, and the intersected set of 82 genes was evaluated through functional enrichment (KEGG, Reactome, GO BP), survival analysis (Kaplan-Meier, Cox regression), and biomarker novelty assessment against six external resources. Hybrid classification models integrating XGBoost, sDCFE, and deep learning were applied to pancancer classification, and the framework was further extended to lung squamous cell carcinoma (LUSC) staging using RNA-seq and methylation data.

RESULTS: The overlap between sDCFE and XGBoost yielded 82 candidate biomarkers enriched in cancer-related pathways, including cell cycle regulation, immune signalling, and DNA repair. Novelty assessment stratified these genes into established, emerging, and novel categories. Six genes-HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP-emerged as the most promising candidates, supported by enrichment and survival associations across multiple cancers. The hybrid model achieved near-perfect pancancer classification on TCGA (accuracy = 99.3%, MCC = 0.992, AUC = 1.0) and demonstrated strong generalizability on PCAWG (accuracy = 94%, MCC = 0.929, AUC = 0.997). In the LUSC staging task, multiomics integration improved classification performance: the CNN-based model reached 84% accuracy, while logistic regression applied to sDCFE-ranked features achieved 88.5% accuracy with superior calibration, highlighting the robustness of the selected features.

CONCLUSION: sDCFE provides a principled extension of Fisher-like methods, enabling stable and interpretable biomarker selection. When combined with XGBoost and deep learning, the framework achieves highly accurate and biologically grounded cancer classification across both cancer types and stages. The identification of novel and prognostic biomarkers, including HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP, underscores its translational potential. These results position the framework as a promising precision oncology tool to support early diagnosis, risk stratification, and treatment decision-making.

PMID:41408184 | DOI:10.1186/s12874-025-02713-z

  •  

Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport

arXiv:2512.13702v1 Announce Type: cross Abstract: Objective: To develop the AI Product Passport, a standards-based framework improving transparency, traceability, and compliance in healthcare AI via lifecycle-based documentation. Materials and Methods: The AI Product Passport was developed within the AI4HF project, focusing on heart failure AI tools. We analyzed regulatory frameworks (EU AI Act, FDA guidelines) and existing standards to design a relational data model capturing metadata across AI lifecycle phases: study definition, dataset preparation, model generation/evaluation, deployment/monitoring, and passport generation. MLOps/ModelOps concepts were integrated for operational relevance. Co-creation involved feedback from AI4HF consortium and a Lisbon workshop with 21 diverse stakeholders, evaluated via Mentimeter polls. The open-source platform was implemented with Python libraries for automated provenance tracking. Results: The AI Product Passport was designed based on existing standards and methods with well-defined lifecycle management and role-based access. Its implementation is a web-based platform with a relational data model supporting auditable documentation. It generates machine- and human-readable reports, customizable for stakeholders. It aligns with FUTURE-AI principles (Fairness, Universality, Traceability, Usability, Robustness, Explainability), ensuring fairness, traceability, and usability. Exported passports detail model purpose, data provenance, performance, and deployment context. GitHub-hosted backend/frontend codebases enhance accessibility. Discussion and Conclusion: The AI Product Passport addresses transparency gaps in healthcare AI, meeting regulatory and ethical demands. Its open-source nature and alignment with standards foster trust and adaptability. Future enhancements include FAIR data principles and FHIR integration for improved interoperability, promoting responsible AI deployment.
  •  

Graph AI generates neurological hypotheses validated in molecular, organoid, and clinical systems

arXiv:2512.13724v1 Announce Type: cross Abstract: Neurological diseases are the leading global cause of disability, yet most lack disease-modifying treatments. We present PROTON, a heterogeneous graph transformer that generates testable hypotheses across molecular, organoid, and clinical systems. To evaluate PROTON, we apply it to Parkinson's disease (PD), bipolar disorder (BD), and Alzheimer's disease (AD). In PD, PROTON linked genetic risk loci to genes essential for dopaminergic neuron survival and predicted pesticides toxic to patient-derived neurons, including the insecticide endosulfan, which ranked within the top 1.29% of predictions. In silico screens performed by PROTON reproduced six genome-wide $\alpha$-synuclein experiments, including a split-ubiquitin yeast two-hybrid system (normalized enrichment score [NES] = 2.30, FDR-adjusted $p
  •  
❌