❌

Normal view

CACARA: Cross-Modal Alignment Leveraging a Text-Centric Approach for Cost-Effective Multimodal and Multilingual Learning

arXiv:2512.00496v1 Announce Type: cross Abstract: As deep learning models evolve, new applications and challenges are rapidly emerging. Tasks that once relied on a single modality, such as text, images, or audio, are now enriched by seamless interactions between multimodal data. These connections bridge information gaps: an image can visually materialize a text, while audio can add context to an image. Researchers have developed numerous multimodal models, but most rely on resource-intensive training across multiple modalities. Similarly, extending these models to new languages often follows the same resource-heavy training strategy. In this work, we propose a multimodal and multilingual architecture, CACARA, trained through emergent alignment learning, enabling the seamless integration of new modalities into an existing bimodal/multimodal model without requiring full retraining. This work breaks new ground by demonstrating that this emergent alignment paradigm can unlock multilingual capabilities from monolingual training. By fine-tuning the newly incorporated modality only on data aligned with the English language, our model develops support for over 100 languages without explicit multilingual pretraining or tuning of the text encoder. Such emergent multimodal and multilingual properties are gained efficiently, preserving previously learned knowledge at a training cost comparable to that of a monolingual model. Our strategy achieves up to a 14.24 percentage points improvement in R@1 audio-to-text retrieval, outperforming state-of-the-art multimodal models -- all without the heavy computational cost of retraining across every modality and language.

Deep Learning-Based Computer Vision Models for Early Cancer Detection Using Multimodal Medical Imaging and Radiogenomic Integration Frameworks

arXiv:2512.00714v1 Announce Type: cross Abstract: Early cancer detection remains one of the most critical challenges in modern healthcare, where delayed diagnosis significantly reduces survival outcomes. Recent advancements in artificial intelligence, particularly deep learning, have enabled transformative progress in medical imaging analysis. Deep learning-based computer vision models, such as convolutional neural networks (CNNs), transformers, and hybrid attention architectures, can automatically extract complex spatial, morphological, and temporal patterns from multimodal imaging data including MRI, CT, PET, mammography, histopathology, and ultrasound. These models surpass traditional radiological assessment by identifying subtle tissue abnormalities and tumor microenvironment variations invisible to the human eye. At a broader scale, the integration of multimodal imaging with radiogenomics linking quantitative imaging features with genomics, transcriptomics, and epigenetic biomarkers has introduced a new paradigm for personalized oncology. This radiogenomic fusion allows the prediction of tumor genotype, immune response, molecular subtypes, and treatment resistance without invasive biopsies.

Multi-Modal AI for Remote Patient Monitoring in Cancer Care

arXiv:2512.00949v1 Announce Type: cross Abstract: For patients undergoing systemic cancer therapy, the time between clinic visits is full of uncertainties and risks of unmonitored side effects. To bridge this gap in care, we developed and prospectively trialed a multi-modal AI framework for remote patient monitoring (RPM). This system integrates multi-modal data from the HALO-X platform, such as demographics, wearable sensors, daily surveys, and clinical events. Our observational trial is one of the largest of its kind and has collected over 2.1 million data points (6,080 patient-days) of monitoring from 84 patients. We developed and adapted a multi-modal AI model to handle the asynchronous and incomplete nature of real-world RPM data, forecasting a continuous risk of future adverse events. The model achieved an accuracy of 83.9% (AUROC=0.70). Notably, the model identified previous treatments, wellness check-ins, and daily maximum heart rate as key predictive features. A case study demonstrated the model's ability to provide early warnings by outputting escalating risk profiles prior to the event. This work establishes the feasibility of multi-modal AI RPM for cancer care and offers a path toward more proactive patient support.(Accepted at Europe NeurIPS 2025 Multimodal Representation Learning for Healthcare Workshop)

A TinyML Reinforcement Learning Approach for Energy-Efficient Light Control in Low-Cost Greenhouse Systems

arXiv:2512.01167v1 Announce Type: cross Abstract: This study presents a reinforcement learning (RL)-based control strategy for adaptive lighting regulation in controlled environments using a low-power microcontroller. A model-free Q-learning algorithm was implemented to dynamically adjust the brightness of a Light-Emitting Diode (LED) based on real-time feedback from a light-dependent resistor (LDR) sensor. The system was trained to stabilize at 13 distinct light intensity levels (L1 to L13), with each target corresponding to a specific range within the 64-state space derived from LDR readings. A total of 130 trials were conducted, covering all target levels with 10 episodes each. Performance was evaluated in terms of convergence speed, steps taken, and time required to reach target states. Box plots and histograms were generated to analyze the distribution of training time and learning efficiency across targets. Experimental validation demonstrated that the agent could effectively learn to stabilize at varying light levels with minimal overshooting and smooth convergence, even in the presence of environmental perturbations. This work highlights the feasibility of lightweight, on-device RL for energy-efficient lighting control and sets the groundwork for multi-modal environmental control applications in resource-constrained agricultural systems.

PhySense: Sensor Placement Optimization for Accurate Physics Sensing

arXiv:2505.18190v4 Announce Type: replace-cross Abstract: Physics sensing plays a central role in many scientific and engineering domains, which inherently involves two coupled tasks: reconstructing dense physical fields from sparse observations and optimizing scattered sensor placements to observe maximum information. While deep learning has made rapid advances in sparse-data reconstruction, existing methods generally omit optimization of sensor placements, leaving the mutual enhancement between reconstruction and placement on the shelf. To change this suboptimal practice, we propose PhySense, a synergistic two-stage framework that learns to jointly reconstruct physical fields and to optimize sensor placements, both aiming for accurate physics sensing. The first stage involves a flow-based generative model enhanced by cross-attention to adaptively fuse sparse observations. Leveraging the reconstruction feedback, the second stage performs sensor placement via projected gradient descent to satisfy spatial constraints. We further prove that the learning objectives of the two stages are consistent with classical variance-minimization principles, providing theoretical guarantees. Extensive experiments across three challenging benchmarks, especially a 3D geometry dataset, indicate PhySense achieves state-of-the-art physics sensing accuracy and discovers informative sensor placements previously unconsidered. Code is available at this repository: https://github.com/thuml/PhySense.

The AI Productivity Index (APEX)

arXiv:2509.25721v3 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Exosome-Mediated RUNX3 DNA Delivery for Lung Cancer Therapy

ACS Appl Mater Interfaces. 2025 Dec 1. doi: 10.1021/acsami.5c15987. Online ahead of print.

ABSTRACT

Gene therapy represents a promising strategy for treating lung cancer, with the potential to inhibit the proliferation of cancerous cells and induce apoptosis. However, current gene therapy for lung cancer encounters challenges with delivery, targeting, and safety, such as off-target effects, immune responses, and the necessity for better delivery methods. Here, we introduce gene therapy using the key regulator in lung adenocarcinoma, runt-related transcription factor 3 (RUNX3), within exosomes (Exos), which are known for their biocompatibility and ability to selectively target cancer cells. We packaged the RUNX3 plasmid DNA into human exosomes (hExo-Rs), designed to target and induce apoptosis in cancer cells, resulting in a viability decrease to 43.3%. Normal fibroblasts remained viable at 96.0%, confirming the safety of hExo-Rs for future therapies. We delivered hExo-Rs to cancer spheroids, examined their effects, and found that cytokines from treated cells promote M1 macrophage polarization, emphasizing their potential for immunotherapy. We developed a hydrogel platform for the targeted 14-day release of RUNX3 pDNA by attaching hExo-Rs to gelatin using microbial transglutaminase, which enables the selective decrease in cancer cell viability and confirms apoptosis. Our demonstration of RUNX3 gene therapy with Exos presents selective anticancer effectiveness and the promise of clinical use through localized, sustained release using the hydrogel.

PMID:41325015 | DOI:10.1021/acsami.5c15987

DNA-Based Liquid Biopsy for Evaluating Surgical and Postsurgical Outcomes in Gynecologic Malignancies: A Systematic Review

J Clin Lab Anal. 2025 Dec 1:e70139. doi: 10.1002/jcla.70139. Online ahead of print.

ABSTRACT

INTRODUCTION: DNA-based liquid biopsies, including circulating tumor DNA (ctDNA) and cell-free DNA (cfDNA), are emerging as minimally invasive biomarkers for monitoring surgical and postsurgical outcomes in gynecologic malignancies. These tools offer the potential to guide early intervention, refine risk stratification, and improve prognostic accuracy. This systematic review aimed to assess the clinical utility of DNA-based liquid biopsies in evaluating recurrence, surgical success, and preoperative diagnosis in gynecologic cancers.

METHODS: A systematic review was conducted in accordance with PRISMA guidelines, covering studies published from 2017 to 2025. Literature searches were performed in PubMed, Scopus, and Web of Science. A total of 32 eligible observational studies involving 3210 patients with ovarian, endometrial, uterine, and other gynecologic malignancies were included. Study quality was assessed using the Newcastle-Ottawa Scale (NOS).

RESULTS: The studies showed a broad geographic and methodological diversity, with a median NOS score of 7. CtDNA and cfDNA demonstrated promise in three key areas: (1) Recurrence prediction-postoperative ctDNA positivity was associated with higher relapse rates and reduced disease-free survival; (2) Monitoring surgical outcomes and treatment response-ctDNA dynamics more accurately reflected tumor burden than traditional markers like CA125; (3) Preoperative diagnostic support-cfDNA methylation profiling and cfDNA/CA125 models enhanced malignancy detection and risk stratification. Ovarian and endometrial cancers were most frequently studied.

CONCLUSIONS: DNA-based liquid biopsies show strong potential in perioperative care for gynecologic cancers. Their integration into clinical workflows could improve the detection of minimal residual disease and inform individualized surgical planning.

PMID:41327898 | DOI:10.1002/jcla.70139

Monitoring of circulating tumor DNA allows early detection of disease relapse in patients with operable breast cancer

Mol Oncol. 2025 Nov 27. doi: 10.1002/1878-0261.70170. Online ahead of print.

ABSTRACT

Breast cancer is known for late recurrences, yet current follow-up lacks radiological or blood-based monitoring for systemic relapse. This study evaluated circulating tumor DNA (ctDNA) monitoring for early detection of systemic relapse after curative treatment. In this case-control study of 70 patients with operable breast cancer (35 with relapse and 35 without relapse), blood samples were collected every 6-12 months during a median 8.3-year follow-up. ctDNA was analyzed by targeted DNA sequencing using Oncomine™ Breast cfDNA Research Assay v2, and results were compared to genetic analysis of tumor and metastasis biopsies. ctDNA was detected at relapse in 19 of 35 (54%) patients with disease relapse and preceded clinical or radiological relapse detection in 17, with a median lead time of 10.3 months. In 13 (68%) patients, there was concordance with tumor mutations, and in seven patients, there was also concordance with metastasis. Among the relapse-free patients, seven were ctDNA-positive postsurgery, and only one of them had a match among the tumor variants. These findings suggest serial ctDNA analysis may enable earlier detection of systemic relapse in patients with operable breast cancer.

PMID:41307327 | DOI:10.1002/1878-0261.70170

Circulating Tumor DNA (ctDNA) in Gastroesophageal Adenocarcinoma (GEA): Evidence and Emerging Applications

27 November 2025 at 19:00

Cancers (Basel). 2025 Nov 18;17(22):3692. doi: 10.3390/cancers17223692.

ABSTRACT

The role of circulating tumor DNA (ctDNA) in gastroesophageal adenocarcinoma (GEA) has expanded in recent years. In resectable disease, postoperative ctDNA is able to detect patients at highest risk of recurrence months before scans. Tumor-informed assays provide the best sensitivity and emerging methylation assays are useful when tissue is scarce. In metastatic GEA, baseline ctDNA burden correlates with prognosis, and a decrease in ctDNA level after treatment initiation reflects therapeutic response. It can also uncover actionable targets, including ERBB2, FGFR2, and MSI-H, and detect resistance that can arise after starting treatment. Limitations include variable assay performance, low shedding in some tumors, clonal hematopoiesis confounding, and a lack of randomized data showing that ctDNA-guided changes improve outcomes. Ongoing trials are testing MRD-guided escalation/de-escalation and ctDNA-directed biomarker therapy. In this review, we evaluate the role of ctDNA in GEA cancers over recent years.

PMID:41301057 | PMC:PMC12650754 | DOI:10.3390/cancers17223692

Morality in AI. A plea to embed morality in LLM architectures and frameworks

arXiv:2511.20689v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate human decision-making and behaviour. Ensuring LLM processing of moral meaning therefore has become a critical challenge. Current approaches rely predominantly on bottom-up methods such as fine-tuning and reinforcement learning from human feedback. We propose a fundamentally different approach: embedding moral meaning processing directly into the architectural mechanisms and frameworks of transformer-based models through top-down design principles. We first sketch a framework that conceptualizes attention as a dynamic interface mediating between structure and processing, contrasting with existing linear attention frameworks in psychology. We start from established biological-artificial attention analogies in neural architecture design to improve cognitive processing. We extend this analysis to moral processing, using Iris Murdoch's theory of loving attention (sustained, just observation that enables moral transformation by reseeing others with clarity and compassion) to philosophically discuss functional analogies between human and LLM moral processing. We formulate and evaluate potentially promising technical operationalizations to embed morality in LLM architectures and frameworks. We acknowledge the limitations of our exploration and give three key contributions. (1) We conceptualize attention as a dynamic system mechanism mediating between structure and processing. (2) Drawing on the Murdoch notion of loving attention, we outline technical pathways for embedding morality in LLMs, through modified training objectives, runtime weight adjustments, and architectural refinements to attention. (3) We argue that integrating morality into architectures and frameworks complements external, constraint-based methods. We conclude with a call for collaboration between transformer designers and philosophers engaged in AI ethics.

From Prediction to Foresight: The Role of AI in Designing Responsible Futures

arXiv:2511.21570v1 Announce Type: new Abstract: In an era marked by rapid technological advancements and complex global challenges, responsible foresight has emerged as an essential framework for policymakers aiming to navigate future uncertainties and shape the future. Responsible foresight entails the ethical anticipation of emerging opportunities and risks, with a focus on fostering proactive, sustainable, and accountable future design. This paper coins the term "responsible computational foresight", examining the role of human-centric artificial intelligence and computational modeling in advancing responsible foresight, establishing a set of foundational principles for this new field and presenting a suite of AI-driven foresight tools currently shaping it. AI, particularly in conjunction with simulations and scenario analysis, enhances policymakers' ability to address uncertainty, evaluate risks, and devise strategies geared toward sustainable, resilient futures. However, responsible foresight extends beyond mere technical forecasting; it demands a nuanced understanding of the interdependencies within social, environmental, economic and political systems, alongside a commitment to ethical, long-term decision-making that supports human intelligence. We argue that AI will play a role as a supportive tool in responsible, human-centered foresight, complementing rather than substituting policymaker judgment to enable the proactive shaping of resilient and ethically sound futures. This paper advocates for the thoughtful integration of AI into foresight practices to empower policymakers and communities as they confront the grand challenges of the 21st century.

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

How Do Companies Manage the Environmental Sustainability of AI? An Interview Study About Green AI Efforts and Regulations

arXiv:2505.07317v2 Announce Type: replace-cross Abstract: With the ever-growing adoption of artificial intelligence (AI), AI-based software and its negative impact on the environment are no longer negligible, and studying and mitigating this impact has become a critical area of research. However, it is currently unclear which role environmental sustainability plays during AI adoption in industry and how AI regulations influence Green AI practices and decision-making in industry. We therefore aim to investigate the Green AI perception and management of industry practitioners. To this end, we conducted a total of 11 interviews with participants from 10 different organizations that adopted AI-based software. The interviews explored three main themes: AI adoption, current efforts in mitigating the negative environmental impact of AI, and the influence of the EU AI Act and the Corporate Sustainability Reporting Directive (CSRD). Our findings indicate that 9 of 11 participants prioritized business efficiency during AI adoption, with minimal consideration of environmental sustainability. Monitoring and mitigation of AI's environmental impact were very limited. Only one participant monitored negative environmental effects. Regarding applied mitigation practices, six participants reported no actions, with the others sporadically mentioning techniques like prompt engineering, relying on smaller models, or not overusing AI. Awareness and compliance with the EU AI Act are low, with only one participant reporting on its influence, while the CSRD drove sustainability reporting efforts primarily in larger companies. All in all, our findings reflect a lack of urgency and priority for sustainable AI among these companies. We suggest that current regulations are not very effective, which has implications for policymakers. Additionally, there is a need to raise industry awareness, but also to provide user-friendly techniques and tools for Green AI practices.

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

arXiv:2506.14652v2 Announce Type: replace-cross Abstract: In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception -- in addition to a more expansive understanding of (1) methodological rigor -- should include aspects related to (2) what background knowledge informs what to work on (epistemic rigor); (3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); (4) how clearly articulated the theoretical constructs under use are (conceptual rigor); (5) what is reported and how (reporting rigor); and (6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much-needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders.

Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues

Nat Cell Biol. 2025 Nov 26. doi: 10.1038/s41556-025-01811-w. Online ahead of print.

ABSTRACT

Spatial omics technologies have transformed biomedical research by enabling high-resolution molecular profiling while preserving the native tissue architecture. These advances provide unprecedented insights into tissue structure and function. However, the high cost and time-intensive nature of spatial omics experiments necessitate careful experimental design, particularly in selecting regions of interest (ROIs) from large tissue sections. Currently, ROI selection is performed manually, which introduces subjectivity, inconsistency and a lack of reproducibility. Previous studies have shown strong correlations between spatial molecular patterns and histological features, suggesting that readily available and cost-effective histology images can be leveraged to guide spatial omics experiments. Here we present Smart Spatial omics (S2-omics), an end-to-end workflow that automatically selects ROIs from histology images with the goal of maximizing molecular information content in the ROIs. Through comprehensive evaluations across multiple spatial omics platforms and tissue types, we demonstrate that S2-omics enables systematic and reproducible ROI selection and enhances the robustness and impact of downstream biological discovery.

PMID:41298871 | DOI:10.1038/s41556-025-01811-w

Information content as a health system screening tool for rare diseases

npj Digital Medicine, Published online: 25 November 2025; doi:10.1038/s41746-025-02096-x

Information content as a health system screening tool for rare diseases

Human Experts' Evaluation of Generative AI for Contextualizing STEAM Education in the Global South

arXiv:2511.19482v2 Announce Type: replace-cross Abstract: This study investigates how human experts evaluate the capacity of Generative AI (GenAI) to contextualize STEAM education in the Global South, with a focus on Ghana. Using a convergent mixed-methods design, four STEAM specialists assessed GenAI-generated lesson plans created with a customized Culturally Responsive Lesson Planner (CRLP) and compared them to standardized lesson plans from the Ghana National Council for Curriculum and Assessment (NaCCA). Quantitative ratings were based on a validated 25-item Culturally Responsive Pedagogy Rubric measuring bias awareness, cultural representation, contextual relevance, linguistic responsiveness, and teacher agency. Qualitative reflections provided additional insight into how GenAI handles cultural and pedagogical appropriateness. Findings show that GenAI, when paired with the CRLP tool, can support contextualized STEAM instruction by linking abstract curriculum standards to learners' cultural knowledge, community practices, and everyday experiences. Experts rated GenAI-assisted lessons as more culturally grounded and pedagogically responsive than NaCCA plans, integrating Indigenous knowledge, bilingual elements, and locally relevant examples. However, GenAI struggled to represent Ghana's cultural pluralism, often offering surface-level references to language, history, and identity. These weaknesses were most evident in Mathematics and Computing, where cultural nuance was limited. The results highlight the need for continued teacher mediation, community involvement, and culturally attuned refinement of AI outputs. Future work should include classroom trials, expanded expert participation, and model fine-tuning using Indigenous language corpora to strengthen cultural fidelity in Global South contexts.
❌