❌

Normal view

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

arXiv:2512.00496v1 Announce Type: cross Abstract: As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Slovak Conceptual Dictionary Miroslav Bl\v{s}t\'ak
    arXiv:2512.00579v1 Announce Type: cross Abstract: When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages, especially in the case of low resourced languages. In this article, we introduce a new conceptual dictionary for the Slovak language as the first linguistic tool of this kind. Since Slovak language is a langu
     

Slovak Conceptual Dictionary

arXiv:2512.00579v1 Announce Type: cross Abstract: When solving tasks in the field of natural language processing, we sometimes need dictionary tools, such as lexicons, word form dictionaries or knowledge bases. However, the availability of dictionary data is insufficient in many languages, especially in the case of low resourced languages. In this article, we introduce a new conceptual dictionary for the Slovak language as the first linguistic tool of this kind. Since Slovak language is a language with limited linguistic resources and there are currently not available any machine-readable linguistic data sources with a sufficiently large volume of data, many tasks which require automated processing of Slovak text achieve weaker results compared to other languages and are almost impossible to solve.

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

arXiv:2512.00590v1 Announce Type: cross Abstract: Knowledge graphs (KGs) provide structured, verifiable grounding for large language models (LLMs), but current LLM-based systems commonly use KGs as auxiliary structures for text retrieval, leaving their intrinsic quality underexplored. In this work, we propose Wikontic, a multi-stage pipeline that constructs KGs from open-domain text by extracting candidate triplets with qualifiers, enforcing Wikidata-based type and relation constraints, and normalizing entities to reduce duplication. The resulting KGs are compact, ontology-consistent, and well-connected; on MuSiQue, the correct answer entity appears in 96% of generated triplets. On HotpotQA, our triplets-only setup achieves 76.0 F1, and on MuSiQue 59.8 F1, matching or surpassing several retrieval-augmented generation baselines that still require textual context. In addition, Wikontic attains state-of-the-art information-retention performance on the MINE-1 benchmark (86%), outperforming prior KG construction methods. Wikontic is also efficient at build time: KG construction uses less than 1,000 output tokens, about 3$\times$ fewer than AriGraph and $

Deep Learning-Based Computer Vision Models for Early Cancer Detection Using Multimodal Medical Imaging and Radiogenomic Integration Frameworks

arXiv:2512.00714v1 Announce Type: cross Abstract: Early cancer detection remains one of the most critical challenges in modern healthcare, where delayed diagnosis significantly reduces survival outcomes. Recent advancements in artificial intelligence, particularly deep learning, have enabled transformative progress in medical imaging analysis. Deep learning-based computer vision models, such as convolutional neural networks (CNNs), transformers, and hybrid attention architectures, can automatically extract complex spatial, morphological, and temporal patterns from multimodal imaging data including MRI, CT, PET, mammography, histopathology, and ultrasound. These models surpass traditional radiological assessment by identifying subtle tissue abnormalities and tumor microenvironment variations invisible to the human eye. At a broader scale, the integration of multimodal imaging with radiogenomics linking quantitative imaging features with genomics, transcriptomics, and epigenetic biomarkers has introduced a new paradigm for personalized oncology. This radiogenomic fusion allows the prediction of tumor genotype, immune response, molecular subtypes, and treatment resistance without invasive biopsies.

Multi-Modal AI for Remote Patient Monitoring in Cancer Care

arXiv:2512.00949v1 Announce Type: cross Abstract: For patients undergoing systemic cancer therapy, the time between clinic visits is full of uncertainties and risks of unmonitored side effects. To bridge this gap in care, we developed and prospectively trialed a multi-modal AI framework for remote patient monitoring (RPM). This system integrates multi-modal data from the HALO-X platform, such as demographics, wearable sensors, daily surveys, and clinical events. Our observational trial is one of the largest of its kind and has collected over 2.1 million data points (6,080 patient-days) of monitoring from 84 patients. We developed and adapted a multi-modal AI model to handle the asynchronous and incomplete nature of real-world RPM data, forecasting a continuous risk of future adverse events. The model achieved an accuracy of 83.9% (AUROC=0.70). Notably, the model identified previous treatments, wellness check-ins, and daily maximum heart rate as key predictive features. A case study demonstrated the model's ability to provide early warnings by outputting escalating risk profiles prior to the event. This work establishes the feasibility of multi-modal AI RPM for cancer care and offers a path toward more proactive patient support.(Accepted at Europe NeurIPS 2025 Multimodal Representation Learning for Healthcare Workshop)
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Will Humanity Be Rendered Obsolete by AI? Mohamed El Louadi · Emna Ben Romdhane
    arXiv:2510.22814v3 Announce Type: replace Abstract: This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existenti
     

Will Humanity Be Rendered Obsolete by AI?

arXiv:2510.22814v3 Announce Type: replace Abstract: This article analyzes the existential risks artificial intelligence (AI) poses to humanity, tracing the trajectory from current AI to ultraintelligence. Drawing on Irving J. Good and Nick Bostrom's theoretical work, plus recent publications (AI 2027; If Anyone Builds It, Everyone Dies), it explores AGI and superintelligence. Considering machines' exponentially growing cognitive power and hypothetical IQs, it addresses the ethical and existential implications of an intelligence vastly exceeding humanity's, fundamentally alien. Human extinction may result not from malice, but from uncontrollable, indifferent cognitive superiority.

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

arXiv:2502.07299v3 Announce Type: replace-cross Abstract: The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-trained models have achieved great success in analyzing these macromolecules individually, their interconnected nature remains underexplored. This paper follows the guidance of the central dogma to redesign both the data and model pipeline and offers a comprehensive framework, Life-Code, that spans different biological functions. As for data flow, we propose a unified pipeline to integrate multi-omics data by reverse-transcribing RNA and reverse-translating amino acids into nucleotide-based sequences. As for the model, we design a codon tokenizer and a hybrid long-sequence architecture to encode the interactions between coding and non-coding regions through masked modeling pre-training. To model the translation and folding process with coding sequences, Life-Code learns protein structures of the corresponding amino acids by knowledge distillation from off-the-shelf protein language models. Such designs enable Life-Code to capture complex interactions within genetic sequences, providing a more comprehensive understanding of multi-omics with the central dogma. Extensive experiments show that Life-Code achieves state-of-the-art results on various tasks across three omics, highlighting its potential for advancing multi-omics analysis and interpretation.

The AI Productivity Index (APEX)

arXiv:2509.25721v3 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.
  • ✇STAT
  • Opinion: Racial bias in medicine can be as simple as dismissing Black patients as a ‘hard stick’ Jahidah La Roche
    I was moments away from a routine screening colonoscopy when it happened again. The warm and professional pre-procedure nurse began preparing for intravenous insertion. She tied the tourniquet loosely around my arm, took a quick glance, and untied it within seconds. “I can’t find a vein. You must be dehydrated,” she said, moving immediately to the back of my hand. I paused. I didn’t feel dehydrated. Yes, I had followed the bowel prep instructions, consuming only liquids the day before, but I
     

Opinion: Racial bias in medicine can be as simple as dismissing Black patients as a ‘hard stick’

2 December 2025 at 17:30

I was moments away from a routine screening colonoscopy when it happened again. The warm and professional pre-procedure nurse began preparing for intravenous insertion. She tied the tourniquet loosely around my arm, took a quick glance, and untied it within seconds. “I can’t find a vein. You must be dehydrated,” she said, moving immediately to the back of my hand.

I paused. I didn’t feel dehydrated. Yes, I had followed the bowel prep instructions, consuming only liquids the day before, but I had no signs of dehydration. I knew my body. I knew my veins.

Read the rest…

© AIZAR RALDES/AFP via Getty Images

DNA-Based Liquid Biopsy for Evaluating Surgical and Postsurgical Outcomes in Gynecologic Malignancies: A Systematic Review

J Clin Lab Anal. 2025 Dec 1:e70139. doi: 10.1002/jcla.70139. Online ahead of print.

ABSTRACT

INTRODUCTION: DNA-based liquid biopsies, including circulating tumor DNA (ctDNA) and cell-free DNA (cfDNA), are emerging as minimally invasive biomarkers for monitoring surgical and postsurgical outcomes in gynecologic malignancies. These tools offer the potential to guide early intervention, refine risk stratification, and improve prognostic accuracy. This systematic review aimed to assess the clinical utility of DNA-based liquid biopsies in evaluating recurrence, surgical success, and preoperative diagnosis in gynecologic cancers.

METHODS: A systematic review was conducted in accordance with PRISMA guidelines, covering studies published from 2017 to 2025. Literature searches were performed in PubMed, Scopus, and Web of Science. A total of 32 eligible observational studies involving 3210 patients with ovarian, endometrial, uterine, and other gynecologic malignancies were included. Study quality was assessed using the Newcastle-Ottawa Scale (NOS).

RESULTS: The studies showed a broad geographic and methodological diversity, with a median NOS score of 7. CtDNA and cfDNA demonstrated promise in three key areas: (1) Recurrence prediction-postoperative ctDNA positivity was associated with higher relapse rates and reduced disease-free survival; (2) Monitoring surgical outcomes and treatment response-ctDNA dynamics more accurately reflected tumor burden than traditional markers like CA125; (3) Preoperative diagnostic support-cfDNA methylation profiling and cfDNA/CA125 models enhanced malignancy detection and risk stratification. Ovarian and endometrial cancers were most frequently studied.

CONCLUSIONS: DNA-based liquid biopsies show strong potential in perioperative care for gynecologic cancers. Their integration into clinical workflows could improve the detection of minimal residual disease and inform individualized surgical planning.

PMID:41327898 | DOI:10.1002/jcla.70139

Acceptability of Health Information Technology by Health Care Professionals: Where We Are Now and How We Can Fill the Gap

Digital health is expected to improve efficiency and quality of health. Health information technologies (HIT) imply allocated time, appropriate training and new types of responsibility whose physical and mental impact on healthcare professionals (HCPs) has emerged as an important issue. The present review provides updated data and opinions about such potential impact and discusses the relevance of ongoing programs established to better characterize barriers and facilitators of HIT implementation. The extent of Internet-based healthcare information and digital apps imposes new responsibilities on HCPs in helping patients select reliable sources and incorporate them in the understanding and self-management of the disease. Several reviews also identified exhaustion, depersonalization, workload, over-alerting, poor work-life integration and job unsatisfaction as potential drivers of electronic health report (EHR)-associated clinician burnout and/or HIT unacceptability. Paradoxically, the increasing use of generative artificial intelligence in the decision-making process may in turn introduce an additional layer of complexity due to required specific skills and associated cognitive overload and stress. Regarding EHRs, various approaches like more proportionate use, better adequation of available commercial tools, or multidisciplinary workflows within the clinic and building of new specialty-specific tools are expected to reduce clinician burden. The ongoing E-health Efficiency Evaluation (E3) project has been developed to define the factors and dimensions impacting overall digital environment and to identify relevant ways of optimizing its acceptability by HCPs. The way of preventing and alleviating the adverse effects of digital health is a major challenge that all HIT stakeholders should be aware of.

AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education: Observational Crossover Cohort Study

Background: Virtual patient (VP) simulations can be used to practice clinical reasoning (CR) in controlled learning environments. Traditional computer-based VP platforms often lack the authenticity and interactivity required for effective CR training. Artificial intelligence (AI)–enhanced social robotic VPs can enhance realism and engagement; however, quantitative evidence comparing them with conventional VP platforms remains limited. Objective: We compared medical students’ experience of an AI-enhanced social robotic versus a conventional computer-based VP platform regarding the extent to which the design characteristics of the respective platform facilitate CR skill training. Methods: This observational crossover cohort study involved 178 sixth-semester medical students at Karolinska Institutet, Stockholm, Sweden (response rate: 42.3%; 178 of 421 invited students; Spring 2024-Spring 2025), who experienced both a large language model–enhanced social robotic VP platform supporting dialogue (social artificial intelligence–enhanced robotic interface [SARI]) and a conventional computer-based VP platform (virtual interactive case [VIC]) during their clinical rotation within rheumatology. Platform order was determined by clinical rotation scheduling. VP design was evaluated using a validated questionnaire across 5 domains: authenticity, professional approach, coaching quality, learning effects, and overall judgment. Students’ CR training preferences were assessed using categorical responses and a Visual Analogue Scale, where a lower score favored SARI and a score of 5 indicated equal preference between platforms. Results: SARI outperformed VIC across all 5 VP design domains. Students rated SARI higher for authenticity (median 4.0, IQR 3.5-4.5 vs 3.0, IQR 2.5-3.5; P<.001 professional approach iqr vs>P<.001 coaching quality iqr vs>P<.001 learning effect iqr vs>P<.001 and overall judgment vs iqr>P<.001 students strongly preferred sari for cr training vs odds ratio ci>P<.001 with visual analogue scale scores confirming this preference iqr>P<.001 preferences were consistent across most subgroups prior vp experience and platform order in the difference was not significant that is students with vs or ci>P=.11) and students first introduced to VIC (55% vs 45%; OR 1.5; 95% CI 0.7-2.9; P=.33). Conclusions: Our findings provide the first quantitative evidence that AI-enhanced social robotic VPs offer superior design characteristics than conventional computer-based platforms for CR training in medical education. These results support the use of AI-driven social robots for VP simulations to better prepare medical students for real clinical encounters, and warrant future research on objective CR skill outcomes and long-term transfer to clinical practice. Unlike previous qualitative studies examining each platform separately, this study provides the first quantitative comparison of design characteristics between AI-enhanced social robotic and conventional computer-based VPs.

Monitoring of circulating tumor DNA allows early detection of disease relapse in patients with operable breast cancer

Mol Oncol. 2025 Nov 27. doi: 10.1002/1878-0261.70170. Online ahead of print.

ABSTRACT

Breast cancer is known for late recurrences, yet current follow-up lacks radiological or blood-based monitoring for systemic relapse. This study evaluated circulating tumor DNA (ctDNA) monitoring for early detection of systemic relapse after curative treatment. In this case-control study of 70 patients with operable breast cancer (35 with relapse and 35 without relapse), blood samples were collected every 6-12 months during a median 8.3-year follow-up. ctDNA was analyzed by targeted DNA sequencing using Oncomine™ Breast cfDNA Research Assay v2, and results were compared to genetic analysis of tumor and metastasis biopsies. ctDNA was detected at relapse in 19 of 35 (54%) patients with disease relapse and preceded clinical or radiological relapse detection in 17, with a median lead time of 10.3 months. In 13 (68%) patients, there was concordance with tumor mutations, and in seven patients, there was also concordance with metastasis. Among the relapse-free patients, seven were ctDNA-positive postsurgery, and only one of them had a match among the tumor variants. These findings suggest serial ctDNA analysis may enable earlier detection of systemic relapse in patients with operable breast cancer.

PMID:41307327 | DOI:10.1002/1878-0261.70170

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

arXiv:2506.14652v2 Announce Type: replace-cross Abstract: In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception -- in addition to a more expansive understanding of (1) methodological rigor -- should include aspects related to (2) what background knowledge informs what to work on (epistemic rigor); (3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); (4) how clearly articulated the theoretical constructs under use are (conceptual rigor); (5) what is reported and how (reporting rigor); and (6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much-needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders.

Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues

Nat Cell Biol. 2025 Nov 26. doi: 10.1038/s41556-025-01811-w. Online ahead of print.

ABSTRACT

Spatial omics technologies have transformed biomedical research by enabling high-resolution molecular profiling while preserving the native tissue architecture. These advances provide unprecedented insights into tissue structure and function. However, the high cost and time-intensive nature of spatial omics experiments necessitate careful experimental design, particularly in selecting regions of interest (ROIs) from large tissue sections. Currently, ROI selection is performed manually, which introduces subjectivity, inconsistency and a lack of reproducibility. Previous studies have shown strong correlations between spatial molecular patterns and histological features, suggesting that readily available and cost-effective histology images can be leveraged to guide spatial omics experiments. Here we present Smart Spatial omics (S2-omics), an end-to-end workflow that automatically selects ROIs from histology images with the goal of maximizing molecular information content in the ROIs. Through comprehensive evaluations across multiple spatial omics platforms and tissue types, we demonstrate that S2-omics enables systematic and reproducible ROI selection and enhances the robustness and impact of downstream biological discovery.

PMID:41298871 | DOI:10.1038/s41556-025-01811-w

Information content as a health system screening tool for rare diseases

npj Digital Medicine, Published online: 25 November 2025; doi:10.1038/s41746-025-02096-x

Information content as a health system screening tool for rare diseases

Toward explainable AI approaches for breast imaging: adapting foundation models to diverse populations

arXiv:2511.17828v1 Announce Type: cross Abstract: Foundation models hold promise for specialized medical imaging tasks, though their effectiveness in breast imaging remains underexplored. This study leverages BiomedCLIP as a foundation model to address challenges in model generalization. BiomedCLIP was adapted for automated BI-RADS breast density classification using multi-modality mammographic data (synthesized 2D images, digital mammography, and digital breast tomosynthesis). Using 96,995 images, we compared single-modality (s2D only) and multi-modality training approaches, addressing class imbalance through weighted contrastive learning. Both approaches achieved similar accuracy (multi-modality: 0.74, single-modality: 0.73), with the multi-modality model offering broader applicability across different imaging modalities and higher AUC values consistently above 0.84 across BI-RADS categories. External validation on the RSNA and EMBED datasets showed strong generalization capabilities (AUC range: 0.80-0.93). GradCAM visualizations confirmed consistent and clinically relevant attention patterns, highlighting the models interpretability and robustness. This research underscores the potential of foundation models for breast imaging applications, paving the way for future extensions for diagnostic tasks.
❌