❌

Normal view

Toward explainable AI approaches for breast imaging: adapting foundation models to diverse populations

arXiv:2511.17828v1 Announce Type: cross Abstract: Foundation models hold promise for specialized medical imaging tasks, though their effectiveness in breast imaging remains underexplored. This study leverages BiomedCLIP as a foundation model to address challenges in model generalization. BiomedCLIP was adapted for automated BI-RADS breast density classification using multi-modality mammographic data (synthesized 2D images, digital mammography, and digital breast tomosynthesis). Using 96,995 images, we compared single-modality (s2D only) and multi-modality training approaches, addressing class imbalance through weighted contrastive learning. Both approaches achieved similar accuracy (multi-modality: 0.74, single-modality: 0.73), with the multi-modality model offering broader applicability across different imaging modalities and higher AUC values consistently above 0.84 across BI-RADS categories. External validation on the RSNA and EMBED datasets showed strong generalization capabilities (AUC range: 0.80-0.93). GradCAM visualizations confirmed consistent and clinically relevant attention patterns, highlighting the models interpretability and robustness. This research underscores the potential of foundation models for breast imaging applications, paving the way for future extensions for diagnostic tasks.

Towards Automating Data Access Permissions in AI Agents

arXiv:2511.17959v1 Announce Type: cross Abstract: As AI agents attempt to autonomously act on users' behalf, they raise transparency and control issues. We argue that permission-based access control is indispensable in providing meaningful control to the users, but conventional permission models are inadequate for the automated agentic execution paradigm. We therefore propose automated permission management for AI agents. Our key idea is to conduct a user study to identify the factors influencing users' permission decisions and to encode these factors into an ML-based permission management assistant capable of predicting users' future decisions. We find that participants' permission decisions are influenced by communication context but importantly individual preferences tend to remain consistent within contexts, and align with those of other participants. Leveraging these insights, we develop a permission prediction model achieving 85.1% accuracy overall and 94.4% for high-confidence predictions. We find that even without using permission history, our model achieves an accuracy of 66.9%, and a slight increase of training samples (i.e., 1-4) can substantially increase the accuracy by 10.8%.

Reinforcement Learning for Portfolio Optimization with a Financial Goal and Defined Time Horizons

arXiv:2511.18076v1 Announce Type: cross Abstract: This research proposes an enhancement to the innovative portfolio optimization approach using the G-Learning algorithm, combined with parametric optimization via the GIRL algorithm (G-learning approach to the setting of Inverse Reinforcement Learning) as presented by. The goal is to maximize portfolio value by a target date while minimizing the investor's periodic contributions. Our model operates in a highly volatile market with a well-diversified portfolio, ensuring a low-risk level for the investor, and leverages reinforcement learning to dynamically adjust portfolio positions over time. Results show that we improved the Sharpe Ratio from 0.42, as suggested by recent studies using the same approach, to a value of 0.483 a notable achievement in highly volatile markets with diversified portfolios. The comparison between G-Learning and GIRL reveals that while GIRL optimizes the reward function parameters (e.g., lambda = 0.0012 compared to 0.002), its impact on portfolio performance remains marginal. This suggests that reinforcement learning methods, like G-Learning, already enable robust optimization. This research contributes to the growing development of reinforcement learning applications in financial decision-making, demonstrating that probabilistic learning algorithms can effectively align portfolio management strategies with investor needs.

The Alignment Paradox of Medical Large Language Models in Infertility Care: Decoupling Algorithmic Improvement from Clinical Decision-making Quality

arXiv:2511.18084v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly adopted in clinical decision support, yet aligning them with the multifaceted reasoning pathways of real-world medicine remains a major challenge. Using more than 8,000 infertility treatment records, we systematically evaluate four alignment strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and In-Context Learning (ICL) through a dual-layer framework combining automatic benchmarks with blinded doctor-in-the-loop assessments. GRPO achieves the highest algorithmic accuracy across multiple decision layers, confirming the value of reinforcement-based optimization for structured prediction tasks. However, clinicians consistently prefer the SFT model, citing clearer reasoning processes (p = 0.035) and higher therapeutic feasibility (p = 0.019). In blinded pairwise comparisons, SFT attains the highest winning rate (51.2%), outperforming both GRPO (26.2%) and even physicians' original decisions (22.7%). These results reveal an alignment paradox: algorithmic improvements do not necessarily translate into higher clinical trust, and may diverge from human-centered preferences. Our findings highlight the need for alignment strategies that prioritize clinically interpretable and practically feasible reasoning, rather than solely optimizing decision-level accuracy.

Clinician-in-the-Loop Smart Home System to Detect Urinary Tract Infection Flare-Ups via Uncertainty-Aware Decision Support

arXiv:2511.18334v1 Announce Type: cross Abstract: Urinary tract infection (UTI) flare-ups pose a significant health risk for older adults with chronic conditions. These infections often go unnoticed until they become severe, making early detection through innovative smart home technologies crucial. Traditional machine learning (ML) approaches relying on simple binary classification for UTI detection offer limited utility to nurses and practitioners as they lack insight into prediction uncertainty, hindering informed clinical decision-making. This paper presents a clinician-in-the-loop (CIL) smart home system that leverages ambient sensor data to extract meaningful behavioral markers, train robust predictive ML models, and calibrate them to enable uncertainty-aware decision support. The system incorporates a statistically valid uncertainty quantification method called Conformal-Calibrated Interval (CCI), which quantifies uncertainty and abstains from making predictions ("I don't know") when the ML model's confidence is low. Evaluated on real-world data from eight smart homes, our method outperforms baseline methods in recall and other classification metrics while maintaining the lowest abstention proportion and interval width. A survey of 42 nurses confirms that our system's outputs are valuable for guiding clinical decision-making, underscoring their practical utility in improving informed decisions and effectively managing UTIs and other condition flare-ups in older adults.

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

arXiv:2511.19220v1 Announce Type: cross Abstract: Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs demonstrate genuine visual grounding when answering Italian medical questions by testing four state-of-the-art models: Claude Sonnet 4.5, GPT-4o, GPT-5-mini, and Gemini 2.0 flash exp. Using 60 questions from the EuropeMedQA Italian dataset that explicitly require image interpretation, we substitute correct medical images with blank placeholders to test whether models truly integrate visual and textual information. Our results reveal striking variability in visual dependency: GPT-4o shows the strongest visual grounding with a 27.9pp accuracy drop (83.2% [74.6%, 91.7%] to 55.3% [44.1%, 66.6%]), while GPT-5-mini, Gemini, and Claude maintain high accuracy with modest drops of 8.5pp, 2.4pp, and 5.6pp respectively. Analysis of model-generated reasoning reveals confident explanations for fabricated visual interpretations across all models, suggesting varying degrees of reliance on textual shortcuts versus genuine visual analysis. These findings highlight critical differences in model robustness and the need for rigorous evaluation before clinical deployment.

AI-Assisted Cardiovascular Risk Assessment by General Practitioners in Resource-Constrained Indonesian Settings Using a Conceptual Prototype: Randomized Controlled Study

Background: Preventive strategies integrated with digital health and artificial intelligence (AI), have significant potential to mitigate the global burden of atherosclerotic cardiovascular disease (ASCVD). AI-enabled clinical decision support (CDS) systems increasingly provide patient-specific insights beyond traditional risk factors. Despite these advances, their capacity to enhance clinical decision-making in resource-constrained settings remains largely unexplored. Objective: We conducted a randomised controlled study to assess the effect of AI-based CDS on 10-year ASCVD risk assessment and management in primary prevention. Methods: In a three-way within-subject randomised design, doctors completed nine clinical vignettes representative of primary care presentations in a resource-constrained outpatient setting. For each vignette, participants assessed 10-year ASCVD risk and made management decisions using either a conceptual prototype of AI-based CDS, automated CDS, or no decision support. The conceptual prototype represented contemporary risk calculators based on traditional machine learning models (e.g., random forest, neural networks, logistic regression) that incorporate additional predictors alongside traditional risk factors. Primary outcomes were correct risk assessment and patient management (prescription of aspirin, statins, and anti-hypertensives; referral for advanced examinations). Decision-making time and perceptions about AI utility were also measured. Results: 102 doctors from all seven geographical regions of Indonesia participated. Most participants were 26–35 years (83%), 56% male, with a median of six years of clinical experience (IQR=4.75). AI-based CDS improved risk assessment by 27% (2(2, n=102) = 48.875, P<.001 compared to unassisted or one additional correct risk classification for every patients where doctors use ai needed treat nnt="3.7;" ci the prescription of statins also improved by n="102)=" p in pairwise comparisons assisted with ai-based cds correctly assessed significantly more cases adjusted and prescribed appropriate statin often medium effect size r=".35)" control. ai-assisted required less time marginal means sec vs f however improvements aspirin anti-hypertensives did not reach statistical significance. no improvement was observed referral decisions. participants generally viewed positively agreeing strongly that they would follow its recommendations indicating it if given access. believed could enhance efficiency assessment particularly high-volume primary care settings while noting need verify against clinical guidelines each patient. conclusions: coupled reduced decision-making highlight potential utility ascvd resource-constrained efficient healthcare resources is crucial. further research ascertain whether this online study translate real-world low-resource settings.>
❌