❌

Reading view

MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers

arXiv:2504.08999v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly augmented with external tools through standardized interfaces like the Model Context Protocol (MCP). However, current MCP implementations face critical limitations: they typically require local process execution through STDIO transports, making them impractical for resource-constrained environments like mobile devices, web browsers, and edge computing. We present MCP Bridge, a lightweight RESTful proxy that connects to multiple MCP servers and exposes their capabilities through a unified API. Unlike existing solutions, MCP Bridge is fully LLM-agnostic, supporting any backend regardless of vendor. The system implements a risk-based execution model with three security levels-standard execution, confirmation workflow, and Docker isolation - while maintaining backward compatibility with standard MCP clients. However, reliable execution within this framework requires models that can strictly adhere to protocol schemas. To this end, we also fine-tuned the Qwen3 4B and 8B model family on the Agent-Ark/Toucan-1.5M dataset using four Reinforcement Learning techniques: Group Relative Policy Optimization (GRPO), Dr. GRPO, Beta Normalization Policy Optimization (BNPO), and Decoupled Alignment Policy Optimization (DAPO). Evaluated on the MCPToolBench++ benchmark, our optimized model achieves an F1 score of 73.0% that outperforms GPT-OSS-120B (62.17%) and remains competitive with the 70B+ parameter baselines. Evaluation demonstrates that MCP Bridge successfully addresses the constraints of direct MCP connections while providing enhanced security controls and cross-platform compatibility, enabling sophisticated LLM-powered applications in previously inaccessible environments.
  •  

Key Information Influencing Patient Decision-Making About AI in Health Care: Survey Experiment Study

Background: Artificial Intelligence (AI)-enabled devices are increasingly used in healthcare. However, there has been limited research on patients’ informational preferences, including which elements of AI device labeling enhance patient understanding, trust, and acceptance. Clear and effective patient-facing communication is essential to address patient concerns and support informed decision-making regarding AI-enabled care. Objective: Using simulated AI device labels in a cardiovascular context, we evaluated three aims. First, we identified key information elements that influence patient trust and acceptance of an AI device. Second, we examined how these effects varied based on patient characteristics. Third, we explored how patients evaluated informational content of AI labels and their perceived effectiveness of the AI labels in informing decision-making about the use of AI device, building trust in the device, and shaping their intention to use it in their healthcare. Methods: We recruited 340 US patients from ResearchMatch.org to participate in a web-based survey that contained two experiments. In the discrete choice experiment (DCE), participants indicated preferences in terms of trust and acceptance regarding 16 pairs of simulated AI device labels that varied across eight types of information needs identified in our previous qualitative work. In the single profile factorial experiment (SPFE), participants evaluated four randomly assigned label prototypes regarding the label’s legibility, comprehensibility, information overload, credibility, and perceived effectiveness in informing about the AI device, as well as participants’ trust in the AI device and intention to use the device in their healthcare. Data was analyzed using mixed effects binary or ordinal logistic regression. Results: The DCE showed that information about regulatory approval, high device performance, provider oversight, and AI’s value added to usual care significantly increased the likelihood of patient trust by 14.1-19.3% and acceptance by 13.3-17.9%. Subgroup analyses revealed variations based on patient characteristics such as familiarity with AI, health literacy, and recency of last medical checkup. The SPFE showed that patients reported good label comprehension, and that information about provider oversight, regulatory approval, device performance, and AI’s added value improved perceived credibility and effectiveness of the AI label (odds ratios [ORs] range 1.35-2.05), reduced doubts in the AI device (ORs range 0.61- 0.77), and increased trust and intention to use the AI device (ORs range 1.47-1.73). However, information about data privacy and safety management protocols are less influential. Conclusions: Patients value information about an AI device’s performance, provider oversight, regulatory status, and added value during decision-making. Providing transparent, easily understandable information about these aspects is critical to support patient determinations of trust and acceptance of AI-enabled healthcare. Information elements’ impact on patient trust and acceptance varies by patient characteristics, highlighting the need for a tailored approach to address the concerns of diverse patient groups about AI in healthcare.
  •  

AI-exposed jobs deteriorated before ChatGPT

arXiv:2601.02554v1 Announce Type: cross Abstract: Public debate links worsening job prospects for AI-exposed occupations to the release of ChatGPT in late 2022. Using monthly U.S. unemployment insurance records, we measure occupation- and location-specific unemployment risk and find that risk rose in AI-exposed occupations beginning in early 2022, months before ChatGPT. Analyzing millions of LinkedIn profiles, we show that graduate cohorts from 2021 onward entered AI-exposed jobs at lower rates than earlier cohorts, with gaps opening before late 2022. Finally, from millions of university syllabi, we find that graduates taking more AI-exposed curricula had higher first-job pay and shorter job searches after ChatGPT. Together, these results point to forces pre-dating generative AI and to the ongoing value of LLM-relevant education.
  •  

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images

arXiv:2512.14640v1 Announce Type: cross Abstract: Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, skilled personnel, and causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information from routinely available HE-stained slides, yet comprehensive benchmarks for lymphoma subtyping on multicenter data are lacking. In this work, we present the first multicenter lymphoma benchmarking dataset covering four common lymphoma subtypes and healthy control tissue. We systematically evaluate five publicly available pathology foundation models (H-optimus-1, H0-mini, Virchow2, UNI2, Titan) combined with attention-based (AB-MIL) and transformer-based (TransMIL) multiple instance learning aggregators across three magnifications (10x, 20x, 40x). On in-distribution test sets, models achieve multiclass balanced accuracies exceeding 80% across all magnifications, with all foundation models performing similarly and both aggregation methods showing comparable results. The magnification study reveals that 40x resolution is sufficient, with no performance gains from higher resolutions or cross-magnification aggregation. However, on out-of-distribution test sets, performance drops substantially to around 60%, highlighting significant generalization challenges. To advance the field, larger multicenter studies covering additional rare lymphoma subtypes are needed. We provide an automated benchmarking pipeline to facilitate such future research.
  •  

Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems

arXiv:2512.12791v2 Announce Type: replace-cross Abstract: Recent advances in agentic AI have shifted the focus from standalone Large Language Models (LLMs) to integrated systems that combine LLMs with tools, memory, and other agents to perform complex tasks. These multi-agent architectures enable coordinated reasoning, planning, and execution across diverse domains, allowing agents to collaboratively automate complex workflows. Despite these advances, evaluation and assessment of LLM agents and the multi-agent systems they constitute remain a fundamental challenge. Although various approaches have been proposed in the software engineering literature for evaluating conventional software components, existing methods for AI-based systems often overlook the non-deterministic nature of models. This non-determinism introduces behavioral uncertainty during execution, yet existing evaluations rely on binary task completion metrics that fail to capture it. Evaluating agentic systems therefore requires examining additional dimensions, including the agent ability to invoke tools, ingest and retrieve memory, collaborate with other agents, and interact effectively with its environment. These challenges emerged during our ongoing industry collaboration with MontyCloud Inc., when we deployed an agentic system in production. These limitations surfaced during deployment, highlighting practical gaps in the current evaluation methods and the need for a systematic assessment of agent behavior beyond task outcomes. Informed by these observations and established definitions of agentic systems, we propose an end-to-end Agent Assessment Framework with four evaluation pillars encompassing LLMs, Memory, Tools, and Environment. We validate the framework on a representative Autonomous CloudOps use case, where experiments reveal behavioral deviations overlooked by conventional metrics, demonstrating its effectiveness in capturing runtime uncertainties.
  •  

Somatic evolution following cancer treatment in normal tissue

Nature, Published online: 10 December 2025; doi:10.1038/s41586-025-09792-4

High-depth sequencing of non-cancerous tissue from patients with metastatic cancer reveals single-base mutational signatures of alcohol, smoking and cancer treatments, and reveals how exogenous factors, including cancer therapies, affect somatic cell evolution.
  •  

Multi-omic profiling provides insights into the heterogeneity, microenvironmental features, and biomarker landscape of small-cell lung cancer

Mol Cancer. 2025 Dec 2. doi: 10.1186/s12943-025-02514-4. Online ahead of print.

ABSTRACT

BACKGROUND: Greater understanding of differential therapeutic sensitivity, specifically to immunotherapy, in small-cell lung cancer (SCLC) is required.

METHODS: We explored SCLC heterogeneity through integrated molecular characterization of tumor tissue samples from 159 treatment-naive patients, utilizing genetic, epigenetic, transcriptional, and proteomic profiling, immunohistochemistry staining for multiple biologically relevant markers including transcriptional subtype-defining proteins, and spatial immune profiling using multiplex immunofluorescence.

RESULTS: Multi-omics analysis confirmed high heterogeneity across/within neuroendocrine and non-neuroendocrine subtypes. Methylomics analysis identified four methylome clusters that may enhance subtype prediction, prognosis, and longitudinal monitoring of subtype evolution. Immunohistochemistry analysis showed high MHC-I expression in non-neuroendocrine subtypes, which have greatest potential benefit from adding immunotherapy to chemotherapy; high DLL3 expression associated with neuroendocrine subtypes and an immune-cold tumor microenvironment. Multiplex immunofluorescence demonstrated associations of MHC-I with spatial arrangement and phenotypic features of immune cells in the tumor microenvironment of high-MHC-I-expressing SCLC, providing mechanistic rationale for MHC-I as a potential biomarker of immunotherapy response.

CONCLUSIONS: This multimodal profiling analysis provides further insights into the biologic complexity of SCLC and highlights potential therapeutic vulnerabilities of distinct disease subtypes.

PMID:41331472 | DOI:10.1186/s12943-025-02514-4

  •  

Integrative Analysis of Multi-Omics Data for Biomarker Discovery

Annu Int Conf IEEE Eng Med Biol Soc. 2025 Jul;2025:1-7. doi: 10.1109/EMBC58623.2025.11254134.

ABSTRACT

The complexity of biological systems and the limitations of analyzing individual omics studies for biomarker discovery have raised the need for a holistic approach by multi-omics integration. By integrating data from multiple layers, researchers can gain insights into the entire system rather than just individual components. Also, integrative analysis can help identify molecular signatures that are more accurate in predicting disease onset, progression, and response to treatment, leading to better-targeted therapies and personalized medicine. In this paper, we explored statistical and deep learning methods for integrative analysis of metabolomics, lipidomics, peptidomics, proteomics, and glycoproteomics data acquired by LC-MS/MS analysis of serum samples from 20 hepatocellular carcinoma (HCC) cases and 20 patients with liver cirrhosis (CIRR). The goal is to identify a panel of multi-omics features that distinguish HCC cases from cirrhotic controls. A pathway analysis using these features identified biological pathways such as LXR/RXR Activation and Acute Response signaling as significantly enriched in our multi-omics datasets.

PMID:41336317 | PMC:PMC12694951 | DOI:10.1109/EMBC58623.2025.11254134

  •  

DNA-Based Liquid Biopsy for Evaluating Surgical and Postsurgical Outcomes in Gynecologic Malignancies: A Systematic Review

J Clin Lab Anal. 2025 Dec 1:e70139. doi: 10.1002/jcla.70139. Online ahead of print.

ABSTRACT

INTRODUCTION: DNA-based liquid biopsies, including circulating tumor DNA (ctDNA) and cell-free DNA (cfDNA), are emerging as minimally invasive biomarkers for monitoring surgical and postsurgical outcomes in gynecologic malignancies. These tools offer the potential to guide early intervention, refine risk stratification, and improve prognostic accuracy. This systematic review aimed to assess the clinical utility of DNA-based liquid biopsies in evaluating recurrence, surgical success, and preoperative diagnosis in gynecologic cancers.

METHODS: A systematic review was conducted in accordance with PRISMA guidelines, covering studies published from 2017 to 2025. Literature searches were performed in PubMed, Scopus, and Web of Science. A total of 32 eligible observational studies involving 3210 patients with ovarian, endometrial, uterine, and other gynecologic malignancies were included. Study quality was assessed using the Newcastle-Ottawa Scale (NOS).

RESULTS: The studies showed a broad geographic and methodological diversity, with a median NOS score of 7. CtDNA and cfDNA demonstrated promise in three key areas: (1) Recurrence prediction-postoperative ctDNA positivity was associated with higher relapse rates and reduced disease-free survival; (2) Monitoring surgical outcomes and treatment response-ctDNA dynamics more accurately reflected tumor burden than traditional markers like CA125; (3) Preoperative diagnostic support-cfDNA methylation profiling and cfDNA/CA125 models enhanced malignancy detection and risk stratification. Ovarian and endometrial cancers were most frequently studied.

CONCLUSIONS: DNA-based liquid biopsies show strong potential in perioperative care for gynecologic cancers. Their integration into clinical workflows could improve the detection of minimal residual disease and inform individualized surgical planning.

PMID:41327898 | DOI:10.1002/jcla.70139

  •  

CLINB: A Climate Intelligence Benchmark for Foundational Models

arXiv:2511.11597v1 Announce Type: new Abstract: Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.
  •  

REFA: Reference Free Alignment for multi-preference optimization

arXiv:2412.16378v4 Announce Type: replace-cross Abstract: To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effective against this bias, we demonstrate that length normalization itself introduces a failure mode: the URSLA shortcut. Here models learn to satisfy the alignment objective by prematurely truncating low-quality responses rather than learning from their semantic content. To address this, we introduce REFA, a new alignment framework that proposes probabilistic control on a structural token that controls termination. Our core innovation is a new class of regularizers that operate directly on the probability of the End-of-Sequence (EOS) token, a previously unexploited control lever. This token-level intervention provides a principled solution to the URSLA shortcut, ensuring genuine quality improvements. Furthermore, it unlocks a versatile mechanism for managing the alignment-efficiency tradeoff, enabling practitioners to fine-tune models that adhere to specific token budgets. Empirically, REFA achieves a 60.29% win rate and a 52.17% length-controlled win rate on AlpacaEval2 with Llama-3-8B-Instruct, demonstrating the power of our token-level control paradigm.
  •  

Wearable Artificial Intelligence for Epilepsy: Scoping Review

Background: Epilepsy affects approximately 50 million people globally and imposes a substantial clinical and societal burden, requiring continuous and personalized monitoring for effective management. Wearable artificial intelligence (AI) technologies offer a promising solution by leveraging physiological signals and machine learning for seizure detection and prediction. While various approaches have been proposed, a comprehensive overview summarizing these advances and challenges is still needed. Objective: This review aims to comprehensively explore and map the existing literature on AI-driven wearable technologies for epilepsy, identifying device characteristics, AI methodologies, biosignal measurements, validation approaches, and research gaps. Methods: A scoping review was conducted following the PRISMA-ScR guidelines. A systematic search was performed across six electronic databases (Scopus, MEDLINE, EMBASE, ACM Digital Library, IEEE Xplore, and Google Scholar) to identify relevant studies published up to December 2023. We included studies that developed AI algorithms for epilepsy using non-invasive wearable devices (e.g., smartwatches, smart clothing) and excluded those using non-wearables or in-body devices. Eligible publication types included journal articles, conference papers, and dissertations. Study selection and data extraction were performed independently by six reviewers. The extracted data was synthesized narratively. Results: A total of 68 studies met the inclusion criteria. Research in this domain has increased significantly since 2021, with India, the United States, and China leading contributions. The studies examined both commercial (45.6%) and non-commercial (47.1%) wearable devices, with Empatica smart bands being the most frequently used. The primary biosignals monitored included activity measures (54.4%), cardiovascular metrics (45.6%), brain activity (35.3%), and electrodermal activity (33.8%). The most common AI models were support vector machines (42.6%), random forests (22.1%), and convolutional neural networks (16.2%). Most models focused on seizure detection (77.5%) compared to seizure prediction (22.5%), reflecting a research imbalance that suggests the need for further development in predictive analytics. Sensitivity (80.9%) was the most frequently reported performance metric, indicating a focus on identifying seizures; however, comprehensive clinical validation remains limited. Closed-source data predominated (64.7%), limiting the generalizability of findings. The most used validation methods were leave-one-out cross-validation (30.9%) and k-fold cross-validation (29.4%), while video-EEG served as the primary reference standard (42.6%). Conclusions: Wearable AI technologies show significant promise in epilepsy management, offering real-time, continuous monitoring and early seizure detection. To realize clinical impact, future research should prioritize the standardization of validation methods, promote open data exchange for reproducibility, and develop energy-efficient algorithms that support real-world deployment in wearable devices.
  •  

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.
  •  

Epistemic Diversity and Knowledge Collapse in Large Language Models

arXiv:2510.04226v4 Announce Type: replace-cross Abstract: Large language models (LLMs) tend to generate lexically, semantically, and stylistically homogenous texts. This poses a risk of knowledge collapse, where homogenous LLMs mediate a shrinking in the range of accessible information over time. Existing works on homogenization are limited by a focus on closed-ended multiple-choice setups or fuzzy semantic features, and do not look at trends across time and cultural contexts. To overcome this, we present a new methodology to measure epistemic diversity, i.e., variation in real-world claims in LLM outputs, which we use to perform a broad empirical study of LLM knowledge collapse. We test 27 LLMs, 155 topics covering 12 countries, and 200 prompt variations sourced from real user chats. For the topics in our study, we show that while newer models tend to generate more diverse claims, nearly all models are less epistemically diverse than a basic web search. We find that model size has a negative impact on epistemic diversity, while retrieval-augmented generation (RAG) has a positive impact, though the improvement from RAG varies by the cultural context. Finally, compared to a traditional knowledge source (Wikipedia), we find that country-specific claims reflect the English language more than the local one, highlighting a gap in epistemic representation
  •  

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.
  •  
  •  

Effectiveness of a Digital Therapy on 6-Month Weight Loss in People With Obesity: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) Randomized Clinical Trial

Background: Obesity is a chronic, relapsing disease influenced by environmental, lifestyle, biological, and genetic factors, affecting over 1 billion people globally. Treatment for adults typically involves multicomponent lifestyle interventions—diet, physical activity, and behavior change—for at least 6-12 months. However, adherence is often low, and in-person sessions can be time-consuming and costly. Digital therapeutics (DTx), which enhance patient engagement and support long-term outcomes, have proven effective in managing chronic and mental health conditions. DTx offer scalable, evidence-based solutions with the potential to improve obesity management. Objective: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) study is a prospective, multicenter, pragmatic, randomized, double-arm, single-blind, placebo-controlled trial evaluating the 6-month efficacy of an innovative, multicomponent digital intervention for obesity, which combines dietary, physical activity, and behavioral strategies in people with obesity (primary objective). Secondary objectives were assessing changes in BMI, waist circumference, blood pressure, glucose metabolism, lipid profile, adherence, and factors associated with absolute 6-month weight loss. Methods: The trial was conducted at 2 obesity centers in Italy with 246 participants aged 18-65 years (BMI 30-45 kg/m2), randomly assigned to either the Digital Therapeutics for Obesity (DTxO) app or a placebo app. DTxO offered personalized diet plans, exercise routines, and psycho-behavioral support, while the placebo app only allowed users to log data without feedback. Both groups followed a Mediterranean-style low-calorie diet with an 800 kcal/day deficit. On average, participants used the DTxO app for 42 minutes/day and the placebo app for 35 minutes, primarily for physical activity tracking. Univariable and multivariable generalized linear models were used to assess associations with 6-month absolute weight change (primary end point) and percent weight change (secondary end point). Results: Overall, 207 participants (84.1%) completed the 6-month visit. Both arms achieved a statistically significant absolute (DtxO: –3.2 kg, IQR –6.0 kg to –0.9 kg; placebo: –4.0 kg, IQR –6.9 kg to –0.5 kg; P<.001) and percent loss in body weight (DtxO: –3.0%, IQR –5.7% to –0.8%; placebo: –4.0%, IQR –8.5% to –0.5%; P<.001) after 6 months, without significant between-group differences (univariable generalized linear models: P=.34 and P=.17, respectively). Univariable regression analyses showed a significant association between adherence to app use and 6-month absolute weight loss (β=–.06, SE 0.02, P=.01) as well as percent weight loss (β=–.05, SE 0.01, P=.01). Adherent participants, defined as those with overall adherence at or above the 75th percentile of daily usage, included 35 individuals in the intervention group and 10 in the placebo group. In this subgroup, the estimated 6-month mean absolute weight change was –7.02 kg (95% CI –9.45 to –4.59) in the DTxO-adherent group and –3.50 kg (95% CI –7.01 to 0.01) in the placebo-adherent group (P=.02). The estimated 6-month mean percent change in weight was –6.31% (95% CI –8.86 to –3.76) in the DTxO-adherent group and –2.78% (95% CI –6.48 to 0.92) in the placebo-adherent group (P=.03). A significantly greater weight loss (P=.01 for study arm, either on absolute or percent change in weight from baseline) among adherent participants randomized to the DTxO app was also confirmed by analyses using mixed linear models for repeated measures. Conclusions: Although overall weight loss did not differ significantly between the DTxO and placebo groups, participants who used the DTxO app for at least 40% of the expected time achieved significantly greater weight loss. These results suggest that higher engagement with DTx can improve obesity outcomes. Further research should explore combining DTxO with pharmacological treatments or bariatric surgery. Trial Registration: ClinicalTrials.gov NCT05394779; https://clinicaltrials.gov/ct2/show/NCT05394779
  •  
❌