❌

Reading view

CLINB: A Climate Intelligence Benchmark for Foundational Models

arXiv:2511.11597v1 Announce Type: new Abstract: Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.
  •  

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.
  •  

Wearable Artificial Intelligence for Epilepsy: Scoping Review

Background: Epilepsy affects approximately 50 million people globally and imposes a substantial clinical and societal burden, requiring continuous and personalized monitoring for effective management. Wearable artificial intelligence (AI) technologies offer a promising solution by leveraging physiological signals and machine learning for seizure detection and prediction. While various approaches have been proposed, a comprehensive overview summarizing these advances and challenges is still needed. Objective: This review aims to comprehensively explore and map the existing literature on AI-driven wearable technologies for epilepsy, identifying device characteristics, AI methodologies, biosignal measurements, validation approaches, and research gaps. Methods: A scoping review was conducted following the PRISMA-ScR guidelines. A systematic search was performed across six electronic databases (Scopus, MEDLINE, EMBASE, ACM Digital Library, IEEE Xplore, and Google Scholar) to identify relevant studies published up to December 2023. We included studies that developed AI algorithms for epilepsy using non-invasive wearable devices (e.g., smartwatches, smart clothing) and excluded those using non-wearables or in-body devices. Eligible publication types included journal articles, conference papers, and dissertations. Study selection and data extraction were performed independently by six reviewers. The extracted data was synthesized narratively. Results: A total of 68 studies met the inclusion criteria. Research in this domain has increased significantly since 2021, with India, the United States, and China leading contributions. The studies examined both commercial (45.6%) and non-commercial (47.1%) wearable devices, with Empatica smart bands being the most frequently used. The primary biosignals monitored included activity measures (54.4%), cardiovascular metrics (45.6%), brain activity (35.3%), and electrodermal activity (33.8%). The most common AI models were support vector machines (42.6%), random forests (22.1%), and convolutional neural networks (16.2%). Most models focused on seizure detection (77.5%) compared to seizure prediction (22.5%), reflecting a research imbalance that suggests the need for further development in predictive analytics. Sensitivity (80.9%) was the most frequently reported performance metric, indicating a focus on identifying seizures; however, comprehensive clinical validation remains limited. Closed-source data predominated (64.7%), limiting the generalizability of findings. The most used validation methods were leave-one-out cross-validation (30.9%) and k-fold cross-validation (29.4%), while video-EEG served as the primary reference standard (42.6%). Conclusions: Wearable AI technologies show significant promise in epilepsy management, offering real-time, continuous monitoring and early seizure detection. To realize clinical impact, future research should prioritize the standardization of validation methods, promote open data exchange for reproducibility, and develop energy-efficient algorithms that support real-world deployment in wearable devices.
  •  

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.
  •  

Epistemic Diversity and Knowledge Collapse in Large Language Models

arXiv:2510.04226v4 Announce Type: replace-cross Abstract: Large language models (LLMs) tend to generate lexically, semantically, and stylistically homogenous texts. This poses a risk of knowledge collapse, where homogenous LLMs mediate a shrinking in the range of accessible information over time. Existing works on homogenization are limited by a focus on closed-ended multiple-choice setups or fuzzy semantic features, and do not look at trends across time and cultural contexts. To overcome this, we present a new methodology to measure epistemic diversity, i.e., variation in real-world claims in LLM outputs, which we use to perform a broad empirical study of LLM knowledge collapse. We test 27 LLMs, 155 topics covering 12 countries, and 200 prompt variations sourced from real user chats. For the topics in our study, we show that while newer models tend to generate more diverse claims, nearly all models are less epistemically diverse than a basic web search. We find that model size has a negative impact on epistemic diversity, while retrieval-augmented generation (RAG) has a positive impact, though the improvement from RAG varies by the cultural context. Finally, compared to a traditional knowledge source (Wikipedia), we find that country-specific claims reflect the English language more than the local one, highlighting a gap in epistemic representation
  •  

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.
  •  
  •  

Effectiveness of a Digital Therapy on 6-Month Weight Loss in People With Obesity: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) Randomized Clinical Trial

Background: Obesity is a chronic, relapsing disease influenced by environmental, lifestyle, biological, and genetic factors, affecting over 1 billion people globally. Treatment for adults typically involves multicomponent lifestyle interventions—diet, physical activity, and behavior change—for at least 6-12 months. However, adherence is often low, and in-person sessions can be time-consuming and costly. Digital therapeutics (DTx), which enhance patient engagement and support long-term outcomes, have proven effective in managing chronic and mental health conditions. DTx offer scalable, evidence-based solutions with the potential to improve obesity management. Objective: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) study is a prospective, multicenter, pragmatic, randomized, double-arm, single-blind, placebo-controlled trial evaluating the 6-month efficacy of an innovative, multicomponent digital intervention for obesity, which combines dietary, physical activity, and behavioral strategies in people with obesity (primary objective). Secondary objectives were assessing changes in BMI, waist circumference, blood pressure, glucose metabolism, lipid profile, adherence, and factors associated with absolute 6-month weight loss. Methods: The trial was conducted at 2 obesity centers in Italy with 246 participants aged 18-65 years (BMI 30-45 kg/m2), randomly assigned to either the Digital Therapeutics for Obesity (DTxO) app or a placebo app. DTxO offered personalized diet plans, exercise routines, and psycho-behavioral support, while the placebo app only allowed users to log data without feedback. Both groups followed a Mediterranean-style low-calorie diet with an 800 kcal/day deficit. On average, participants used the DTxO app for 42 minutes/day and the placebo app for 35 minutes, primarily for physical activity tracking. Univariable and multivariable generalized linear models were used to assess associations with 6-month absolute weight change (primary end point) and percent weight change (secondary end point). Results: Overall, 207 participants (84.1%) completed the 6-month visit. Both arms achieved a statistically significant absolute (DtxO: –3.2 kg, IQR –6.0 kg to –0.9 kg; placebo: –4.0 kg, IQR –6.9 kg to –0.5 kg; P<.001) and percent loss in body weight (DtxO: –3.0%, IQR –5.7% to –0.8%; placebo: –4.0%, IQR –8.5% to –0.5%; P<.001) after 6 months, without significant between-group differences (univariable generalized linear models: P=.34 and P=.17, respectively). Univariable regression analyses showed a significant association between adherence to app use and 6-month absolute weight loss (β=–.06, SE 0.02, P=.01) as well as percent weight loss (β=–.05, SE 0.01, P=.01). Adherent participants, defined as those with overall adherence at or above the 75th percentile of daily usage, included 35 individuals in the intervention group and 10 in the placebo group. In this subgroup, the estimated 6-month mean absolute weight change was –7.02 kg (95% CI –9.45 to –4.59) in the DTxO-adherent group and –3.50 kg (95% CI –7.01 to 0.01) in the placebo-adherent group (P=.02). The estimated 6-month mean percent change in weight was –6.31% (95% CI –8.86 to –3.76) in the DTxO-adherent group and –2.78% (95% CI –6.48 to 0.92) in the placebo-adherent group (P=.03). A significantly greater weight loss (P=.01 for study arm, either on absolute or percent change in weight from baseline) among adherent participants randomized to the DTxO app was also confirmed by analyses using mixed linear models for repeated measures. Conclusions: Although overall weight loss did not differ significantly between the DTxO and placebo groups, participants who used the DTxO app for at least 40% of the expected time achieved significantly greater weight loss. These results suggest that higher engagement with DTx can improve obesity outcomes. Further research should explore combining DTxO with pharmacological treatments or bariatric surgery. Trial Registration: ClinicalTrials.gov NCT05394779; https://clinicaltrials.gov/ct2/show/NCT05394779
  •  

Genomically matched therapy in advanced solid tumors: the randomized phase 2 ROME trial

Nature Medicine, Published online: 29 September 2025; doi:10.1038/s41591-025-03918-x

In the proof-of-concept phase 2 ROME trial, comprehensive genomic profiling followed by molecular tumor board evaluation and randomization of patients with metastatic solid cancer to receive personalized therapy or standard of care led to a significantly higher objective response rate and longer progression-free survival in patients who received personalized therapy.
  •  

Multi-Omics Feature Selection to Identify Biomarkers for Hepatocellular Carcinoma

Metabolites. 2025 Aug 28;15(9):575. doi: 10.3390/metabo15090575.

ABSTRACT

INTRODUCTION: Hepatocellular carcinoma (HCC), the most prevalent form of liver cancer, ranks as the third leading cause of mortality globally. Patients diagnosed with HCC exhibit a dismal prognosis mostly due to the emergence of symptoms in the advanced stages of the disease. Moreover, conventional biomarkers demonstrate insufficient efficacy in the early detection of HCC, hence highlighting the need for the identification of novel and more effective biomarkers.

METHODS: In this paper, we investigate methods for integration of multi-omics data we generated by both untargeted and targeted mass spectrometric analysis of serum samples from HCC cases and patients with liver cirrhosis. Specifically, the performances of several feature selection methods are evaluated on their abilities to identify a panel of multi-omics features that distinguish HCC cases from cirrhotic controls.

RESULTS: The integrative analysis identified key molecules associated with liver including such as leucine and isoleucine as well as SERPINA1, which is involved in LXR/RXR Activation and Acute Response signaling. A new method that uses recursive feature selection in conjunction with a transformer-based deep learning model as an estimator led to more promising results compared to other deep learning methods that perform disease classification and feature selection sequentially.

CONCLUSIONS: The findings in this study reinforce the importance of adapting or extending deep learning models to support robust feature selection, especially for integration of multi-omics data with limited sample size to avoid the risk of overfitting and the need for evaluation of the multi-omics features discovered in this study via blood samples from a larger and independent cohort to identify robust biomarkers for HCC.

PMID:41002959 | PMC:PMC12471784 | DOI:10.3390/metabo15090575

  •  

Single-cell multiome and spatial profiling reveals pancreas cell type-specific gene regulatory programs of type 1 diabetes progression

Sci Adv. 2025 Sep 12;11(37):eady0080. doi: 10.1126/sciadv.ady0080. Epub 2025 Sep 10.

ABSTRACT

Cell type-specific regulatory programs that drive type 1 diabetes (T1D) in the pancreas are poorly understood. Here, we performed single-nucleus multiomics and spatial transcriptomics in up to 32 nondiabetic (ND), autoantibody-positive (AAB+), and T1D pancreas donors. Genomic profiles from 853,005 cells mapped to 12 pancreatic cell types, including multiple exocrine subtypes. β, Acinar, and other cell types, and related cellular niches, had altered abundance and gene activity in T1D progression, including distinct pathways altered in AAB+ compared to T1D. We identified epigenomic drivers of gene activity in T1D and AAB+ which, combined with genetic association, revealed causal pathways of T1D risk including antigen presentation in β cells. Last, single-cell and spatial profiles together revealed widespread changes in cell-cell signaling in T1D including signals affecting β cell regulation. Overall, these results revealed drivers of T1D in the pancreas, which form the basis for therapeutic targets for disease prevention.

PMID:40929272 | PMC:PMC12422192 | DOI:10.1126/sciadv.ady0080

  •  

The WHO global landscape of cancer clinical trials

Nature Medicine, Published online: 09 September 2025; doi:10.1038/s41591-025-03926-x

This Review of the WHO’s International Clinical Trials Registry Platform presents a snapshot of the global cancer trial landscape and provides critical empirical evidence to inform policy, practice and investment.
  •  

Scalable generation and functional classification of genetic variants in inborn errors of immunity to accelerate clinical diagnosis and treatment

In lieu of traditional genetic variant testing approaches, an approach using scalable variant classification in primary human T cells with a clinically relevant readout can inform rapid diagnosis and treatment of inborn errors of immunity.
  •  

Systema: a framework for evaluating genetic perturbation response prediction beyond systematic variation

Nature Biotechnology, Published online: 25 August 2025; doi:10.1038/s41587-025-02777-8

An evaluation framework isolates perturbation-specific effects in perturbation datasets.
  •  

NAVIGATOR: A regional multimodal imaging biobank initiative powered by AI tools for precision medicine in oncology

Eur J Radiol. 2025 Jul 22;191:112327. doi: 10.1016/j.ejrad.2025.112327. Online ahead of print.

ABSTRACT

The NAVIGATOR project established an Italian regional imaging biobank and interactive research platform designed to support precision oncology through the integration of multimodal imaging, clinical, and omics data. The platform goes beyond a static repository, offering a secure Virtual Research Environment (VRE) where users can upload data, test AI algorithms, and execute complete analytical pipelines. The platform incorporates artificial intelligence (AI)-driven radiomics and deep learning methodologies to enable biomarker extraction, disease stratification, and predictive modeling. This manuscript presents the development and implementation of the NAVIGATOR infrastructure, including its data governance framework, ethical and legal considerations, and application to three oncological use cases: prostate, rectal, and gastric cancers. To date, the biobank has collected imaging and clinical data from over 700 patients across these cohorts. AI models were deployed within a dedicated VRE to facilitate image analysis, feature extraction, and classification tasks. The project addresses critical challenges related to data harmonization, regulatory compliance, privacy safeguards and fairness in AI systems. NAVIGATOR demonstrates the feasibility of integrating AI methodologies within imaging biobanks and provides a scalable framework to advance oncological research and support clinical decision-making.

PMID:40743874 | DOI:10.1016/j.ejrad.2025.112327

  •  

Liquid biopsy in breast cancer: Redefining precision medicine

J Liq Biopsy. 2025 Jul 16;9:100312. doi: 10.1016/j.jlb.2025.100312. eCollection 2025 Sep.

ABSTRACT

Breast cancer (BC) is the most frequent cancer and the leading cause of cancer-related death among women worldwide. It represents a heterogeneous group of diseases with distinct morphological, immunophenotypic, and molecular profiles, which significantly impact clinical behavior and therapeutic response. Moreover, under treatment pressure, tumor cells may undergo molecular changes and phenotypic plasticity, leading to resistance and therapeutic failure. Although tissue biopsy remains the gold standard for diagnosis and molecular characterization, it has several limitations, including invasiveness, sampling bias, and the inability to dynamically capture tumor evolution over time. Hence, a non-invasive and repeatable approach capable of real-time monitoring is increasingly needed. Liquid biopsy (LB), through the analysis of circulating tumor cells (CTCs) and circulating tumor DNA (ctDNA), has emerged as a powerful tool to complement tissue biopsy. It allows for longitudinal assessment of tumor burden, detection of minimal residual disease, and identification of molecular alterations relevant to targeted therapies. Despite promising results, the integration of LB into clinical practice is still limited by methodological heterogeneity, standardization gaps, and regulatory issues. Nonetheless, LB represents a key advancement toward precision oncology and may become essential in the personalized management of BC patients. In this review, we explore the current applications, benefits, and technical limitations of LB in different BC settings. We provide a comprehensive overview of the biological and clinical significance of CTCs and ctDNA, emphasizing their diagnostic, prognostic, and predictive roles. Finally, we present an updated summary of ongoing clinical trials that incorporate LB for clinical decision-making.

PMID:40740670 | PMC:PMC12308030 | DOI:10.1016/j.jlb.2025.100312

  •  
❌