❌

Normal view

AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads

arXiv:2506.20535v2 Announce Type: replace-cross Abstract: The rapid advancement of AI, particularly large language models (LLMs), has raised significant concerns about the energy use and carbon emissions associated with model training and inference. However, existing tools for measuring and reporting such impacts are often fragmented, lacking systematic metric integration and offering limited support for correlation analysis among them. This paper presents AIMeter, a comprehensive software toolkit for the measurement, analysis, and visualization of energy use, power draw, hardware performance, and carbon emissions across AI workloads. By seamlessly integrating with existing AI frameworks, AIMeter offers standardized reports and exports fine-grained time-series data to support benchmarking and reproducibility in a lightweight manner. It further enables in-depth correlation analysis between hardware metrics and model performance and thus facilitates bottleneck identification and performance enhancement. By addressing critical limitations in existing tools, AIMeter encourages the research community to weigh environmental impact alongside raw performance of AI workloads and advances the shift toward more sustainable "Green AI" practices. The code is available at https://github.com/SusCom-Lab/AIMeter.

Nanomaterial-assisted immunodiagnostic profiling and therapeutic targeting of hepatocellular carcinoma: from molecular biomarkers to clinical applications

Front Immunol. 2025 Oct 14;16:1668630. doi: 10.3389/fimmu.2025.1668630. eCollection 2025.

ABSTRACT

AIMS AND OBJECTIVES: This study aimed to identify immunologically relevant transcriptomic and proteomic biomarkers in hepatocellular carcinoma (HCC) and to characterize their B-cell epitopes for potential integration into nanomaterial-based biosensors and immunomodulatory platforms for early diagnosis and targeted therapy.

METHODS: We conducted a comprehensive multi-omics analysis by integrating transcriptomic (TCGA-LIHC) and proteomic data to identify differentially expressed genes (DEGs) in HCC. Protein-protein interaction networks and pathway enrichment were used to prioritize hub genes. Five candidate biomarkers, RFC2, HSP90AB1, YWHAZ, CYP2E1, and ADH4, were selected for qRT-PCR and serum ELISA validation in clinical cohorts comprising 85 HCC patients and 50 healthy controls. B-cell epitope prediction was performed using BepiPred 2.0 and validated through synthetic peptide-based ELISA in the same cohort to assess immunoreactivity. Diagnostic performance was evaluated using ROC curve analysis.

RESULTS: RFC2, HSP90AB1, and YWHAZ were significantly upregulated (|log2FC|>0.2) and showed high serological expression, whereas CYP2E1 and ADH4 were consistently downregulated. Predicted B-cell epitopes from RFC2, HSP90AB1, and YWHAZ exhibited strong immunoreactivity (AUC>0.84), indicating their diagnostic potential. Enrichment analysis revealed that upregulated DEGs were involved in cell cycle and mitotic progression, while downregulated genes were linked to immune suppression and metabolic dysfunction. These validated immunogenic epitopes offer promising anchors for nanomaterial-functionalized biosensors, such as gold nanoparticle-conjugated ELISA, graphene-based electrochemical platforms, and peptide-coated quantum dots, for ultrasensitive and multiplexed HCC detection.

CONCLUSION: By integrating transcriptomic and proteomic screening with epitope-level validation, we identified a novel panel of immunogenic biomarkers suitable for nanomaterial-enabled diagnostics in HCC. These findings support the translational potential of peptide-nano scaffold conjugates in developing minimally invasive, immune-responsive biosensing and therapeutic tools tailored for early-stage liver cancer management.

PMID:41164201 | PMC:PMC12558944 | DOI:10.3389/fimmu.2025.1668630

Prospective proteomics for discovering biomarkers in lung adenocarcinoma: a literature review

Transl Cancer Res. 2025 Sep 30;14(9):6102-6117. doi: 10.21037/tcr-2025-1092. Epub 2025 Sep 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung adenocarcinoma (LUAD), as the main subtype of non-small cell lung cancer (NSCLC), faces clinical challenges including molecular heterogeneity, late diagnosis, and aggressive growth, leading to a low 5-year survival rate. Biomarkers are critical for early detection, accurate differentiation of benign/malignant lesions, and guiding personalized treatment strategies. Proteomic technologies using liquid biopsy show potential by analyzing protein changes and post-translational modifications (PTMs) to identify novel biomarkers and unravel cancer mechanisms. This review examines proteomic advances in LUAD, compares platform strengths, lists validated protein markers, and discusses challenges like specificity and regulations. It aims to develop a precision medicine framework by integrating multi-omics data for improved diagnosis and treatment.

METHODS: This study conducted a literature review by searching the PubMed and Web of Science databases for original articles written in English from 2002 to 2025, using the keywords "lung adenocarcinoma" OR "LUAD" AND "biomarkers" AND "proteomics" OR "SomaScan" OR "spatial proteomics" to identify the latest research findings in the field of proteomics technology and LUAD biomarkers. The included studies mainly focused on the current landscape of biomarkers in the diagnosis, treatment, and prognosis of LUAD.

KEY CONTENT AND FINDINGS: This review discusses high-throughput methods for comprehensive protein profiling in accessible biospecimens (tissues, blood, urine) to identify biomarkers for LUAD. We systematically evaluate emerging proteomic strategies, including mass spectrometry (MS), proximity extension assays (PEAs), spatial proteomics techniques, and SomaScan platforms-coupled with innovative computational frameworks have revolutionized biomarkers discovery and their translational potential in developing precision diagnostics and targeted therapies. Additionally, the review addresses challenges in integrating proteomics with genomics, transcriptomics, and metabolomics, offering new methodologies and expanding research in life sciences. As technological advancements continue, it is anticipated that more potential biomarkers will be conducted to validate the broader application in LUAD treatment, addressing early-stage disease complexities and aiding in selecting more effective treatment strategies.

CONCLUSIONS: By synthesizing cutting-edge evidence on proteome-driven LUAD biomarkers, this review elucidates actionable strategies to refine early detection protocols and mechanism-informed personalized treatment frameworks, directly advancing precision oncology initiatives for this prevalent malignancy through biomarker-guided clinical decision-making and multi-omics integration.

PMID:41158224 | PMC:PMC12554480 | DOI:10.21037/tcr-2025-1092

From Detection to Discovery: A Closed-Loop Approach for Simultaneous and Continuous Medical Knowledge Expansion and Depression Detection on Social Media

arXiv:2510.23626v1 Announce Type: cross Abstract: Social media user-generated content (UGC) provides real-time, self-reported indicators of mental health conditions such as depression, offering a valuable source for predictive analytics. While prior studies integrate medical knowledge to improve prediction accuracy, they overlook the opportunity to simultaneously expand such knowledge through predictive processes. We develop a Closed-Loop Large Language Model (LLM)-Knowledge Graph framework that integrates prediction and knowledge expansion in an iterative learning cycle. In the knowledge-aware depression detection phase, the LLM jointly performs depression detection and entity extraction, while the knowledge graph represents and weights these entities to refine prediction performance. In the knowledge refinement and expansion phase, new entities, relationships, and entity types extracted by the LLM are incorporated into the knowledge graph under expert supervision, enabling continual knowledge evolution. Using large-scale UGC, the framework enhances both predictive accuracy and medical understanding. Expert evaluations confirmed the discovery of clinically meaningful symptoms, comorbidities, and social triggers complementary to existing literature. We conceptualize and operationalize prediction-through-learning and learning-through-prediction as mutually reinforcing processes, advancing both methodological and theoretical understanding in predictive analytics. The framework demonstrates the co-evolution of computational models and domain knowledge, offering a foundation for adaptive, data-driven knowledge systems applicable to other dynamic risk monitoring contexts.

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

arXiv:2510.21293v2 Announce Type: replace Abstract: Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. However, current research often adopts techno-centric approaches, focusing primarily on technical attributes such as reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.

Integrating deep learning and multi-omics features in radiation pneumonitis prediction for lung cancer patients using PET/CT

BMC Med Imaging. 2025 Oct 27;25(1):426. doi: 10.1186/s12880-025-01971-z.

ABSTRACT

BACKGROUND: To investigate the feasibility and accuracy of PET radiomics features, along with their combination with CT radiomics, dosiomics, and deep learning (DL) features, in predicting radiation pneumonitis (RP) in lung cancer patients treated with volumetric modulated arc therapy (VMAT).

METHODS: A total of 206 and 27 lung cancer patients who underwent VMAT with pre-treatment PET/CT imaging were enrolled from Hospital One and Hospital Two for model training and external validation, respectively. Four machine learning (ML) methods were applied to build radiomics models with features extracted from CT (R_CT), PET (R_PET), radiomics features fused PET/CT (R_fFU) and fused PET/CT images (R_ iFU), as well dosiomics features (D). Three DL models were built to extract features from PET (DL_PET), CT (DL_CT), and fused PET/CT images (DL_FU). The best-performing radiomics and DL models were combined with dosiomics to create the final joint model. ROC curves with AUC, accuracy, sensitivity, and specificity evaluated the performance. A nomogram was constructed using top-performing model features, parameters, and relevant clinical factors.

RESULTS: The extreme gradient boosting (XGBoost) and 18-layer residual neural network (Resnet-18) achieved the best performance. The R+D+DL model combined radiomics, dosiomics, and DL features achieved AUCs of 0.93, 0.92 and 0.89 in the training, internal validaiton and external validation cohorts, respectively. A nomogram constructed with gender, Adaptive RT, SUVp90, and XGBoost-score achieved an AUC of 0.94 for RP prediction in VMAT-treated lung cancer patients using PET/CT.

CONCLUSION: Integrating radiomics, DL, dosiomics features and SUVp90 is promising in the RP prediction for lung cancer patients underwent VMAT using PET/CT images.

PMID:41146084 | DOI:10.1186/s12880-025-01971-z

Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions

arXiv:2411.14571v2 Announce Type: replace-cross Abstract: Answering end user security questions is challenging. While large language models (LLMs) like GPT, LLAMA, and Gemini are far from error-free, they have shown promise in answering a variety of questions outside of security. We studied LLM performance in the area of end user security by qualitatively evaluating 3 popular LLMs on 900 systematically collected end user security questions. While LLMs demonstrate broad generalist ``knowledge'' of end user security information, there are patterns of errors and limitations across LLMs consisting of stale and inaccurate answers, and indirect or unresponsive communication styles, all of which impacts the quality of information received. Based on these patterns, we suggest directions for model improvement and recommend user strategies for interacting with LLMs when seeking assistance with security.

Multimodal 3D Genome Pre-training

arXiv:2504.09060v2 Announce Type: replace-cross Abstract: Deep learning techniques have driven significant progress in various analytical tasks within 3D genomics in computational biology. However, a holistic understanding of 3D genomics knowledge remains underexplored. Here, we propose MIX-HIC, the first multimodal foundation model of 3D genome that integrates both 3D genome structure and epigenomic tracks, which obtains unified and comprehensive semantics. For accurate heterogeneous semantic fusion, we design the cross-modal interaction and mapping blocks for robust unified representation, yielding the accurate aggregation of 3D genome knowledge. Besides, we introduce the first large-scale dataset comprising over 1 million pairwise samples of Hi-C contact maps and epigenomic tracks for high-quality pre-training, enabling the exploration of functional implications in 3D genomics. Extensive experiments show that MIX-HIC can significantly surpass existing state-of-the-art methods in diverse downstream tasks. This work provides a valuable resource for advancing 3D genomics research.

A full life cycle biological clock based on routine clinical data and its impact in health and diseases

Nature Medicine, Published online: 27 October 2025; doi:10.1038/s41591-025-04006-w

The biological clock model LifeClock predicts biological age across all life stages from routine clinical data, revealing distinct pediatric and adult disease risk patterns.

MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning

arXiv:2510.21093v1 Announce Type: new Abstract: Recently, large models have shown significant potential for smart healthcare. However, the deployment of Large Vision-Language Models (LVLMs) for clinical services is currently hindered by three critical challenges: a tendency to hallucinate answers not grounded in visual evidence, the inefficiency of fixed-depth reasoning, and the difficulty of multi-institutional collaboration. To address these challenges, in this paper, we develop MedAlign, a novel framework to ensure visually accurate LVLM responses for Medical Visual Question Answering (Med-VQA). Specifically, we first propose a multimodal Direct Preference Optimization (mDPO) objective to explicitly align preference learning with visual context. We then design a Retrieval-Aware Mixture-of-Experts (RA-MoE) architecture that utilizes image and text similarity to route queries to a specialized and context-augmented LVLM (i.e., an expert), thereby mitigating hallucinations in LVLMs. To achieve adaptive reasoning and facilitate multi-institutional collaboration, we propose a federated governance mechanism, where the selected expert, fine-tuned on clinical datasets based on mDPO, locally performs iterative Chain-of-Thought (CoT) reasoning via the local meta-cognitive uncertainty estimator. Extensive experiments on three representative Med-VQA datasets demonstrate that MedAlign achieves state-of-the-art performance, outperforming strong retrieval-augmented baselines by up to $11.85\%$ in F1-score, and simultaneously reducing the average reasoning length by $51.60\%$ compared with fixed-depth CoT approaches.

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

arXiv:2510.21293v1 Announce Type: new Abstract: Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. However, current research often adopts techno-centric approaches, focusing primarily on technical attributes such as reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.

Benchmarking GPT-5 for biomedical natural language processing

arXiv:2509.04462v2 Announce Type: replace-cross Abstract: Biomedical literature and clinical narratives pose multifaceted challenges for natural language understanding, from precise entity extraction and document synthesis to multi-step diagnostic reasoning. This study extends a unified benchmark to evaluate GPT-5 and GPT-4o under zero-, one-, and five-shot prompting across five core biomedical NLP tasks: named entity recognition, relation extraction, multi-label document classification, summarization, and simplification, and nine expanded biomedical QA datasets covering factual knowledge, clinical reasoning, and multimodal visual understanding. Using standardized prompts, fixed decoding parameters, and consistent inference pipelines, we assessed model performance, latency, and token-normalized cost under official pricing. GPT-5 consistently outperformed GPT-4o, with the largest gains on reasoning-intensive datasets such as MedXpertQA and DiagnosisArena and stable improvements in multimodal QA. In core tasks, GPT-5 achieved better chemical NER and ChemProt scores but remained below domain-tuned baselines for disease NER and summarization. Despite producing longer outputs, GPT-5 showed comparable latency and 30 to 50 percent lower effective cost per correct prediction. Fine-grained analyses revealed improvements in diagnosis, treatment, and reasoning subtypes, whereas boundary-sensitive extraction and evidence-dense summarization remain challenging. Overall, GPT-5 approaches deployment-ready performance for biomedical QA while offering a favorable balance of accuracy, interpretability, and economic efficiency. The results support a tiered prompting strategy: direct prompting for large-scale or cost-sensitive applications, and chain-of-thought scaffolds for analytically complex or high-stakes scenarios, highlighting the continued need for hybrid solutions where precision and factual fidelity are critical.

VaultGemma: A Differentially Private Gemma Model

arXiv:2510.15001v2 Announce Type: replace-cross Abstract: We introduce VaultGemma 1B, a 1 billion parameter model within the Gemma family, fully trained with differential privacy. Pretrained on the identical data mixture used for the Gemma 2 series, VaultGemma 1B represents a significant step forward in privacy-preserving large language models. We openly release this model to the community

MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration

arXiv:2510.19423v1 Announce Type: new Abstract: We introduce MSC-Bench, a large-scale benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents in a hierarchical Model-Context Protocol (MCP) ecosystem. Existing benchmarks often evaluate tools in isolation, ignoring challenges such as functional overlap and cross-server orchestration, leading to overly optimistic assessments. MSC-Bench addresses these gaps by constructing ground truth through 'equal function sets', allowing objective metrics such as F1 score and reducing the dependency on LLM-as-a-judge evaluation. Organized as a five-level curriculum, it systematically tests agent capabilities from single-tool orchestration to complex cross-server planning, and robustness to out-of-scope requests. Experiments reveal that rigid hierarchies can hinder performance without co-designed strategies, and even state-of-the-art agents exhibit systemic weaknesses in robustness. MSC-Bench provides a diagnostic framework to expose these limitations and guide the development of more capable and efficient tool-using agents. The benchmark and resources are publicly available at https://github.com/snooow1029/MSC_Bench.

R-loops in hepatocellular carcinoma: Bridging genomic instability and therapeutic opportunity (Review)

Mol Med Rep. 2026 Jan;33(1):6. doi: 10.3892/mmr.2025.13716. Epub 2025 Oct 17.

ABSTRACT

R‑loops, three‑stranded nucleic acid structures composed of an RNA:DNA hybrid and displaced single‑stranded DNA, have emerged as important regulators of gene expression and genome maintenance. Although physiological R‑loops participate in normal cellular processes, their dysregulation can threaten genomic integrity by inducing DNA damage and replication stress. The present review explores the role of R‑loops in hepatocellular carcinoma (HCC), a malignancy characterized by marked genomic instability. In the present review, the formation mechanisms of R‑loops, their dual functions in transcriptional regulation and DNA damage, and their specific implications for HCC pathophysiology were discussed. HCC cells exhibit altered R‑loop homeostasis with aberrant accumulation linked to hepatitis B virus infection, inflammatory signaling and oncogene activation. The present review highlighted how HCC cells exploit or manage R‑loops to promote tumor progression, particularly through the epigenetic silencing of differentiation genes and modulation of replication stress responses. Furthermore, emerging therapeutic strategies targeting R‑loop biology were examined, including small molecules that induce synthetic lethality, gene‑based interventions and combination approaches that exploit R‑loop vulnerabilities. Challenges in targeting R‑loops and future directions, including multi‑omics profiling and biomarker development, were also addressed. Understanding the complex interplay between R‑loops and HCC offers promising avenues for novel diagnostic and therapeutic approaches for this malignancy.

PMID:41104860 | DOI:10.3892/mmr.2025.13716

Scalable generation and functional classification of genetic variants in inborn errors of immunity to accelerate clinical diagnosis and treatment

In lieu of traditional genetic variant testing approaches, an approach using scalable variant classification in primary human T cells with a clinically relevant readout can inform rapid diagnosis and treatment of inborn errors of immunity.

Protein lipoylation in cancer: metabolic reprogramming and therapeutic potential

Cell Death Discovery, Published online: 02 September 2025; doi:10.1038/s41420-025-02718-z

Protein lipoylation in cancer: metabolic reprogramming and therapeutic potential

An eyecare foundation model for clinical assistance: a randomized controlled trial

Nature Medicine, Published online: 28 August 2025; doi:10.1038/s41591-025-03900-7

Trained and validated on multimodal data from 14.5 million images from multicountry datasets, a foundation model is shown to increase diagnostic and referral accuracy of clinicians when used as an assistant in a trial involving 16 ophthalmologists and 668 patients.
❌