❌

Reading view

Toward an AI Reasoning-Enabled System for Patient-Clinical Trial Matching

arXiv:2512.08026v1 Announce Type: new Abstract: Screening patients for clinical trial eligibility remains a manual, time-consuming, and resource-intensive process. We present a secure, scalable proof-of-concept system for Artificial Intelligence (AI)-augmented patient-trial matching that addresses key implementation challenges: integrating heterogeneous electronic health record (EHR) data, facilitating expert review, and maintaining rigorous security standards. Leveraging open-source, reasoning-enabled large language models (LLMs), the system moves beyond binary classification to generate structured eligibility assessments with interpretable reasoning chains that support human-in-the-loop review. This decision support tool represents eligibility as a dynamic state rather than a fixed determination, identifying matches when available and offering actionable recommendations that could render a patient eligible in the future. The system aims to reduce coordinator burden, intelligently broaden the set of trials considered for each patient and guarantee comprehensive auditability of all AI-generated outputs.
  •  

Beyond Traditional Diagnostics: Transforming Patient-Side Information into Predictive Insights with Knowledge Graphs and Prototypes

arXiv:2512.08261v1 Announce Type: new Abstract: Predicting diseases solely from patient-side information, such as demographics and self-reported symptoms, has attracted significant research attention due to its potential to enhance patient awareness, facilitate early healthcare engagement, and improve healthcare system efficiency. However, existing approaches encounter critical challenges, including imbalanced disease distributions and a lack of interpretability, resulting in biased or unreliable predictions. To address these issues, we propose the Knowledge graph-enhanced, Prototype-aware, and Interpretable (KPI) framework. KPI systematically integrates structured and trusted medical knowledge into a unified disease knowledge graph, constructs clinically meaningful disease prototypes, and employs contrastive learning to enhance predictive accuracy, which is particularly important for long-tailed diseases. Additionally, KPI utilizes large language models (LLMs) to generate patient-specific, medically relevant explanations, thereby improving interpretability and reliability. Extensive experiments on real-world datasets demonstrate that KPI outperforms state-of-the-art methods in predictive accuracy and provides clinically valid explanations that closely align with patient narratives, highlighting its practical value for patient-centered healthcare delivery.
  •  

Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology

arXiv:2512.08674v1 Announce Type: new Abstract: Multimodal clinical reasoning in the field of gastrointestinal (GI) oncology necessitates the integrated interpretation of endoscopic imagery, radiological data, and biochemical markers. Despite the evident potential exhibited by Multimodal Large Language Models (MLLMs), they frequently encounter challenges such as context dilution and hallucination when confronted with intricate, heterogeneous medical histories. In order to address these limitations, a hierarchical Multi-Agent Framework is proposed, which emulates the collaborative workflow of a human Multidisciplinary Team (MDT). The system attained a composite expert evaluation score of 4.60/5.00, thereby demonstrating a substantial improvement over the monolithic baseline. It is noteworthy that the agent-based architecture yielded the most substantial enhancements in reasoning logic and medical accuracy. The findings indicate that mimetic, agent-based collaboration provides a scalable, interpretable, and clinically robust paradigm for automated decision support in oncology.
  •  

Towards Foundation Models with Native Multi-Agent Intelligence

arXiv:2512.08743v1 Announce Type: new Abstract: Foundation models (FMs) are increasingly assuming the role of the "brain" of AI agents. While recent efforts have begun to equip FMs with native single-agent abilities -- such as GUI interaction or integrated tool use -- we argue that the next frontier is endowing FMs with native multi-agent intelligence. We identify four core capabilities of FMs in multi-agent contexts: understanding, planning, efficient communication, and adaptation. Contrary to assumptions about the spontaneous emergence of such abilities, we provide extensive empirical evidence across 41 large language models showing that strong single-agent performance alone does not automatically yield robust multi-agent intelligence. To address this gap, we outline key research directions -- spanning dataset construction, evaluation, training paradigms, and safety considerations -- for building FMs with native multi-agent intelligence.
  •  

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture

arXiv:2512.08130v1 Announce Type: cross Abstract: Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons. An important element of such efforts is the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper describes the first component of a novel Biothreat Benchmark Generation (BBG) Framework. The BBG approach is designed to help model developers and evaluators reliably measure and assess the biosecurity risk uplift and general harm potential of existing and future AI models, while accounting for key aspects of the threat itself that are often overlooked in other benchmarking efforts, including different actor capability levels, and operational (in addition to purely technical) risk factors. As a pilot, the BBG is first being developed to address bacterial biological threats only. The BBG is built upon a hierarchical structure of biothreat categories, elements and tasks, which then serves as the basis for the development of task-aligned queries. This paper outlines the development of this biothreat task-query architecture, which we have named the Bacterial Biothreat Schema, while future papers will describe follow-on efforts to turn queries into model prompts, as well as how the resulting benchmarks can be implemented for model evaluation. Overall, the BBG Framework, including the Bacterial Biothreat Schema, seeks to offer a robust, re-usable structure for evaluating bacterial biological risks arising from LLMs across multiple levels of aggregation, which captures the full scope of technical and operational requirements for biological adversaries, and which accounts for a wide spectrum of biological adversary capabilities.
  •  

ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Access

arXiv:2512.08193v1 Announce Type: cross Abstract: We present ClinicalTrialsHub, an interactive search-focused platform that consolidates all data from ClinicalTrials.gov and augments it by automatically extracting and structuring trial-relevant information from PubMed research articles. Our system effectively increases access to structured clinical trial data by 83.8% compared to relying on ClinicalTrials.gov alone, with potential to make access easier for patients, clinicians, researchers, and policymakers, advancing evidence-based medicine. ClinicalTrialsHub uses large language models such as GPT-5.1 and Gemini-3-Pro to enhance accessibility. The platform automatically parses full-text research articles to extract structured trial information, translates user queries into structured database searches, and provides an attributed question-answering system that generates evidence-grounded answers linked to specific source sentences. We demonstrate its utility through a user study involving clinicians, clinical researchers, and PhD students of pharmaceutical sciences and nursing, and a systematic automatic evaluation of its information extraction and question answering capabilities.
  •  

Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset

arXiv:2512.08459v1 Announce Type: cross Abstract: The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons has generated significant policy, academic, and public concern. Both model developers and policymakers seek to quantify and mitigate any risk, with an important element of such efforts being the development of model benchmarks that can assess the biosecurity risk posed by a particular model. This paper discusses the pilot implementation of the Bacterial Biothreat Benchmark (B3) dataset. It is the third in a series of three papers describing an overall Biothreat Benchmark Generation (BBG) framework, with previous papers detailing the development of the B3 dataset. The pilot involved running the benchmarks through a sample frontier AI model, followed by human evaluation of model responses, and an applied risk analysis of the results along several dimensions. Overall, the pilot demonstrated that the B3 dataset offers a viable, nuanced method for rapidly assessing the biosecurity risk posed by a LLM, identifying the key sources of that risk and providing guidance for priority areas of mitigation priority.
  •  

AI-powered virtual tissues from spatial proteomics for clinical diagnostics and biomedical discovery

arXiv:2501.06039v2 Announce Type: replace-cross Abstract: Spatial proteomics technologies have transformed our understanding of complex tissue architecture in cancer but present unique challenges for computational analysis. Each study uses a different marker panel and protocol, and most methods are tailored to single cohorts, which limits knowledge transfer and robust biomarker discovery. Here we present Virtual Tissues (VirTues), a general-purpose foundation model for spatial proteomics that learns marker-aware, multi-scale representations of proteins, cells, niches and tissues directly from multiplex imaging data. From a single pretrained backbone, VirTues supports marker reconstruction, cell typing and niche annotation, spatial biomarker discovery, and patient stratification, including zero-shot annotation across heterogeneous panels and datasets. In triple-negative breast cancer, VirTues-derived biomarkers predict anti-PD-L1 chemo-immunotherapy response and stratify disease-free survival in an independent cohort, outperforming state-of-the-art biomarkers derived from the same datasets and current clinical stratification schemes.
  •  

OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities

arXiv:2505.23856v2 Announce Type: replace-cross Abstract: The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse. The core approach to mitigate these concerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptible to attacks that exploit mismatched generalization of model capabilities (e.g., prompts in low-resource languages or prompts provided in non-text modalities such as image and audio). To tackle this challenge, we propose Omniguard, an approach for detecting harmful prompts across languages and modalities. Our approach (i) identifies internal representations of an LLM/MLLM that are aligned across languages or modalities and then (ii) uses them to build a language-agnostic or modality-agnostic classifier for detecting harmful prompts. Omniguard improves harmful prompt classification accuracy by 11.57\% over the strongest baseline in a multilingual setting, by 20.44\% for image-based prompts, and sets a new SOTA for audio-based prompts. By repurposing embeddings computed during generation, Omniguard is also very efficient ($\approx\!120 \times$ faster than the next fastest baseline). Code and data are available at: https://github.com/vsahil/OmniGuard.
  •  

Somatic evolution following cancer treatment in normal tissue

Nature, Published online: 10 December 2025; doi:10.1038/s41586-025-09792-4

High-depth sequencing of non-cancerous tissue from patients with metastatic cancer reveals single-base mutational signatures of alcohol, smoking and cancer treatments, and reveals how exogenous factors, including cancer therapies, affect somatic cell evolution.
  •  

Lung Cancer Diagnosis and Prognostic Monitoring Through Cell-Free RNA via Liquid Biopsy

Ther Clin Risk Manag. 2025 Dec 2;21:1615-1636. doi: 10.2147/TCRM.S542338. eCollection 2025.

ABSTRACT

Lung cancer remains a leading cause of cancer-related mortality worldwide, largely due to challenges in its early detection and effective management. Despite advances in treatment modalities, the complex nature of lung cancer, characterized by its molecular heterogeneity and resistance mechanisms, underscores the need for innovative approaches. Cell-free RNA (cfRNA) has emerged as a promising biomarker with significant clinical applications in lung cancer diagnosis, monitoring, and precision medicine. We explore key themes including the utility of cfRNA in early detection, differentiation between benign and malignant lung nodules, molecular subtyping, and real-time therapeutic monitoring. Advances in liquid biopsy technologies, particularly non-invasive cfRNA analysis, provide dynamic means of tracking tumor evolution. cfRNA biomarkers such as miRNA, long non-coding RNAs, and circular RNAs offer unique insights into tumor biology, paving the way for personalized treatment strategies. Further, we discuss the application of cutting-edge technologies such as AI-driven analytics, next-generation sequencing, and multi-omics integration, which are enhancing the clinical utility of cfRNA in identifying treatment resistance and improving outcomes in immunotherapy, targeted therapy, and chemotherapy. The review addresses significant challenges facing cfRNA applications, including pre-analytical variability, technical limitations in detection methods, economic constraints, and the lack of standardization in clinical protocols. Through multidisciplinary collaborations and standardized methodologies, significant progress can be made toward integrating cfRNA into routine clinical practice. Emphasis is placed on future research directions, which include validating cfRNA biomarkers across diverse populations, streamlining workflows, and addressing scalability issues for real-world applications. This comprehensive exploration positions cfRNA at the forefront of innovations in lung cancer management, offering a pathway for improved diagnostic accuracy and individualized care.

PMID:41367889 | PMC:PMC12682701 | DOI:10.2147/TCRM.S542338

  •  

Decoding the enigma of multiple primary lung cancers: from mechanism to bedside-a narrative review

Transl Lung Cancer Res. 2025 Nov 30;14(11):5181-5197. doi: 10.21037/tlcr-2025-957. Epub 2025 Nov 27.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer is the leading cause of global cancer mortality. Multiple primary lung cancer (MPLC) represents a clinically challenging subtype characterized by independent tumor foci. Distinguishing MPLC from intrapulmonary metastases is crucial for prognosis and treatment. This review integrates current evidence on MPLC's etiology, molecular mechanisms, diagnosis, and management, aiming to provide a clinical reference and highlight future precision medicine directions.

METHODS: We searched PubMed/MEDLINE, Web of Science, and Google Scholar for articles published between January 2000 and September 2024. Search terms included "multiple primary lung cancer", "diagnosis", "molecular characteristics", and "treatment". The selection focused on English-language research and reviews addressing MPLC pathogenesis, diagnosis, or management.

KEY CONTENT AND FINDINGS: The review delineates the multifactorial pathogenesis of MPLC, encompassing genetic susceptibility, somatic heterogeneity, clonal evolution, and epigenetic dysregulation. It frames these mechanisms against a backdrop of "field cancerization" and dynamic tumor microenvironment interactions. The evolution of diagnosis from histology to integrated molecular-artificial intelligence (AI) models is detailed, alongside treatment strategies that must overcome the challenge of inter-lesional heterogeneity.

CONCLUSIONS: MPLC is a distinct entity arising from genetic, epigenetic, and microenvironmental interplay. Advancing its management requires multi-omics integration to decipher pathology and identify biomarkers. Future work should develop AI-enhanced diagnostics and lesion-specific treatment strategies. This review synthesizes current evidence to inform and direct future research and clinical innovation in MPLC.

PMID:41367572 | PMC:PMC12683420 | DOI:10.21037/tlcr-2025-957

  •  

Causal relationship of immune cell characteristics in hepatocellular carcinoma: A multi-omics analysis based on Mendelian randomization

Medicine (Baltimore). 2025 Dec 5;104(49):e45942. doi: 10.1097/MD.0000000000045942.

ABSTRACT

The tumor immune microenvironment of hepatocellular carcinoma (HCC) is complex, yet the causal relationship between immune cell subpopulations and HCC risk remains incompletely elucidated. This study aims to systematically evaluate the causal association between immune cell subpopulations and HCC using Mendelian randomization (MR) analysis, and to validate the biological mechanisms underlying these associations through multi-omics data. Bidirectional two-sample MR analysis was performed to examine causal relationships between 731 immune cell subpopulations and HCC. Inverse-variance weighting (IVW) served as the primary analysis method, with robustness validation through Bayesian weighted MR (BWMR) and machine learning algorithms. Therefore, for significantly associated immune subpopulations, independent analyses of gene expression, prognosis, and tumor immune microenvironment were conducted using HCC data from the Cancer Genome Atlas (TCGA) LIHC cohort. MR analysis and validation identified 21 immune cell subpopulations with significant causal associations to HCC risk. Among these, 12 were identified as risk factors, and 9 as protective factors. Validation in the TCGA cohort revealed that risk-associated immune subpopulations were predominantly enriched for markers of T cell exhaustion and immunosuppressive microenvironments, whereas protective subpopulations likely represented a distinct regulatory B cell subset whose function was associated with the anti-inflammatory factor interleukin-10. This study genetically confirms that specific immune cell functional subpopulations constitute causal risk factors for HCC. These subpopulations exert their effects by shaping distinct tumor immune microenvironments. These findings provide novel mechanisms for understanding the immunopathogenesis of HCC and identify potential targets for developing novel immune intervention strategies.

PMID:41366997 | DOI:10.1097/MD.0000000000045942

  •  

Integrative modeling of longitudinal cell-free DNA and tumor volume dynamics: a multimodal quantitative prognostic framework

Transl Lung Cancer Res. 2025 Nov 30;14(11):4746-4755. doi: 10.21037/tlcr-2025-940. Epub 2025 Nov 27.

ABSTRACT

BACKGROUND: Liquid biopsy based on cell-free DNA (cfDNA) in oncology has emerged as a promising technique for tracking cancer dynamics, especially for detecting minimal residual disease. To date, most studies have used cfDNA for static evaluations of tumor burden. In this study, we propose a novel approach integrating serial cfDNA and computed tomography (CT) tumor volume to fully reflect the dynamic nature of tumor response after treatment.

METHODS: This prospective study involved 25 patients treated with curative-intent radiotherapy for localized non-small cell lung cancer (NSCLC) between June 2019 and November 2020, with 17 subsequently included in final analysis. Longitudinal blood samples were divided into two phases relative to day 3 after treatment initiation, and kinetic parameters, such as velocity and acceleration of cfDNA levels, were calculated. To complement sparse samplings in later days, volume data from routine CT scans were incorporated. K-means clustering using two different variable sets (cfDNA only and cfDNA with volume parameters) and conventional assessment using Response Evaluation Criteria in Solid Tumors (RECIST) v1.1 were applied to stratify patients, and their performance was compared.

RESULTS: The model incorporating both cfDNA and volume parameters effectively separated responders (mean progression-free survival, 44.2 months) from non-responders [16.6 months, P=0.02; area under the receiver operating characteristic curve (AUC) =0.955], outperforming cfDNA only model (36.0 vs. 14.5 months, P=0.04; AUC =0.848). In contrast, RECIST v1.1-based conventional assessment showed no significant difference (P=0.62, AUC =0.70).

CONCLUSIONS: Therefore, our study demonstrates that integration of longitudinal cfDNA and tumor volume dynamics yielded improved assessment of treatment response and prognosis in NSCLC.

PMID:41367558 | PMC:PMC12683446 | DOI:10.21037/tlcr-2025-940

  •  

The Download: a peek at AI’s future

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

The State of AI: A vision of the world in 2030  

There are huge gulfs of opinion when it comes to predicting the near-future impacts of generative AI. In one camp there are those who predict that over the next decade the impact of AI will exceed that of the Industrial Revolution—a 150-year period of economic and social upheaval so great that we still live in the world it wrought. 

At the other end of the scale we have team ‘Normal Technology’: experts who push back not only on these sorts of predictions but on their foundational worldview. That’s not how technology works, they argue.

Advances at the cutting edge may come thick and fast, but change across the wider economy, and society as a whole, moves at human speed. Widespread adoption of new technologies can be slow; acceptance slower. AI will be no different. What should we make of these extremes? 

Read the full conversation between MIT Technology Review’s senior AI editor Will Douglas Heaven and Tim Bradshaw, FT global tech correspondent, about where AI will go next, and what our world will look like in the next five years.

This is the final edition of The State of AI, a collaboration between the Financial Times and MIT Technology Review. Read the rest of the series, and if you want to keep up-to-date with what’s going on in the world of AI, sign up to receive our free Algorithm newsletter every Monday.

How AI is changing the economy

There’s a lot at stake when it comes to understanding how AI is changing the economy at large. What’s the right outlook to have? Join Mat Honan, editor in chief, David Rotman, editor at large, and Richard Waters, FT columnist, at 1pm ET today to hear them discuss what’s happening across industries and the market. Sign up now to be part of this exclusive subscriber-only event.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Trump says he’ll sign an order blocking states from regulating AI
But he’s facing a lot of pushback, including from members of his own party. (CNN)
+ The whole debacle can be traced back to congressional inaction. (Semafor)

2 Google’s new smart glasses are getting rave reviews 👓
You’ll be able to get your hands on a pair in 2026. Watch out, Apple and Meta. (Tech Radar)

3 Trump gave the go-ahead for Nvidia to sell powerful AI chips to China
The US gets a 25% cut of the sales—but what does it lose longer-term? (WP $)
+ And how much could China stand to gain? (NYT $)
+ How a top Chinese AI model overcame US sanctions. (MIT Technology Review)

4 America’s data center backlash is here
Republican and Democrat alike, local residents are sick of rapidly rising power bills. (Vox $)
+ More than 200 environmental groups are demanding a US-wide moratorium on new data centers. (The Guardian)
+ The data center boom in the desert. (MIT Technology Review)

5 A quarter of teens are turning to AI chatbots for mental health support
Given the lack of real-world help, can you really blame them? (The Guardian)
+ Therapists are secretly using ChatGPT. Clients are triggered. (MIT Technology Review)

6 ICEBlock is suing the US government over its App Store removal 
Its creator is arguing that the Department of Justice’s demands to Apple violated his First Amendment rights. (404 Media)
+ It’s one of a number of ICE-tracking initiatives to be pulled by tech platforms this year. (MIT Technology Review)

7 This band quit Spotify, but it’s been replaced by AI knockoffs
The platform seems to be struggling against the tide of slop. (Futurism) 
+ AI is coming for music, too. (MIT Technology Review)

8 Think you’re immune to online ads? Think again
If you’re scrolling on social media, you’re being sold to. Relentlessly. (The Verge $)

9 People really do not like Microsoft Copilot
It’s like Clippy all over again, except it’s even less avoidable. (Quartz $)

10 The longest solar eclipse for 100 years is coming
And we’ll only have to wait until 2027 to see it! (Wired $)

Quote of the day

“Governments and MPs are shooting themselves in the foot by pandering to tech giants, because that just tells young people that they don’t care about our future.”

—Adele Zeynep Walton, founding member of online safety campaign group Ctrl+Alt+Reclaim, tells The Guardian why young activists are taking matters into their own hands. 

One more thing

fleet of ships at sea
COURTESY OF OCEANBIRD

Inside the long quest to advance Chinese writing technology

Every second of every day, someone is typing in Chinese. Though the mechanics look a little different from typing in English—people usually type the pronunciation of a character and then pick it out of a selection that pops up, autocomplete-style—it’s hard to think of anything more quotidian. The software that allows this exists beneath the awareness of pretty much everyone who uses it. It’s just there.

What’s largely been forgotten is that a large cast of eccentrics and linguists, engineers and polymaths, spent much of the 20th century torturing themselves over how Chinese was ever going to move away from the ink brush to any other medium. Read the full story.

—Veronique Greenwood

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Pantone chose a ‘calming’ shade of white for its Color of 2026… and people are fuming. 
+ Ozempic needles on the Christmas tree, anyone? Here’s why we’re going crazy for weird baubles. 
+ Can relate to this baby seal for instinctively heading to the nearest pub.
+ Thrilled to see One Battle After Another get so many Golden Globes nominations.

  •  

AI Application in Anti-Money Laundering for Sustainable and Transparent Financial Systems

arXiv:2512.06240v1 Announce Type: new Abstract: Money laundering and financial fraud remain major threats to global financial stability, costing trillions annually and challenging regulatory oversight. This paper reviews how artificial intelligence (AI) applications can modernize Anti-Money Laundering (AML) workflows by improving detection accuracy, lowering false-positive rates, and reducing the operational burden of manual investigations, thereby supporting more sustainable development. It further highlights future research directions including federated learning for privacy-preserving collaboration, fairness-aware and interpretable AI, reinforcement learning for adaptive defenses, and human-in-the-loop visualization systems to ensure that next-generation AML architectures remain transparent, accountable, and robust. In the final part, the paper proposes an AI-driven KYC application that integrates graph-based retrieval-augmented generation (RAG Graph) with generative models to enhance efficiency, transparency, and decision support in KYC processes related to money-laundering detection. Experimental results show that the RAG-Graph architecture delivers high faithfulness and strong answer relevancy across diverse evaluation settings, thereby enhancing the efficiency and transparency of KYC CDD/EDD workflows and contributing to more sustainable, resource-optimized compliance practices.
  •  

KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening

arXiv:2512.05994v1 Announce Type: cross Abstract: With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.
  •  

Data Visualization Support for Interdisciplinary Team Treatment Planning in Clinical Oncology: Scoping Review

Background: Complex and expanding datasets in clinical oncology applications require flexible and interactive visualization of patient data to provide physicians and other medical professionals with maximum amount of information. In particular, interdisciplinary tumor conferences profit from customized tools to integrate, link, and visualize relevant data from all professions involved. Objective: Our objective was to identify and present currently available data visualization tools for tumor boards and related areas. We wanted to provide an overview of not only the digital tools currently used in tumor board settings but also of the data they include, their respective visualization solutions, and their integration into hospital processes. Methods: This scoping review was based on the scoping study framework by Arksey and O’Malley and attempted to answer the following research question: “What are the key features of data visualization solutions used in molecular and organ tumor boards, and how are these elements integrated and used within the clinical setting?” The following electronic databases were searched for articles: PubMed, Web of Science, and Scopus. Articles were deemed eligible if published in English in the last 10 years. Eligible articles were first deduplicated, followed by screening of titles and abstracts. Full-text screening was then conducted to decide on article selection. All included articles were analyzed using a data extraction template. The template included a variety of meta-information, as well as specific fields aiming to answer the research question. Results: The review process started with 2049 articles, of which 1014 (49.49%) were included in the title and abstract screening. A total of 5.47% (112/2049) of the publications were eligible for full-text screening, leading to 2.93% (60/2049) of the publications being eligible for final inclusion. They covered 49 distinct visualization tools and applications. We discovered a variety of innovative visualization solutions, most often driven by the complexity of omics data, represented in 96% (47/49) of the tools. Tables remained the most used tool for the visualization of data types described in the articles. Approximately one-third of the identified tools (16/49, 33%) were systematically evaluated in some form. For most discovered tools (37/49, 76%), there was no documentation of implementation into the clinical routine. A significant number of applications (21/49, 43%) were available through open-source access. Conclusions: There is a wide range of projects providing visualization solutions for tumor boards and clinical oncology applications. Among the few tools that have made their way into clinical routine settings, there are both commercial and academic solutions. While tables for a variety of data types remain the dominant visualization strategy, the complexity of omics data appears to be the driving force behind many visualization innovations in the domain of tumor boards. Trial Registration:
  •  

Digital Health Technologies: Learnings and Perspectives From a Patient Engagement Stakeholder Expectations Matrix Study

As digital health technologies become increasingly integrated into health care systems worldwide, there is growing recognition that their full potential can be realized only when development is rooted in patient engagement (PE). Despite its proven value in clinical research and health care delivery, PE remains insufficiently embedded in digital health design and implementation. This perspective paper explores the current state of PE in digital health through findings from the Stakeholder Expectations Matrix program developed by Patient Focused Medicines Development. Drawing from 37 in-depth interviews across 6 key stakeholder groups, complemented by insights gathered during a multisession cocreation track at the Patient Engagement Open Forum, this paper highlights differing perspectives on digital health, the barriers to meaningful engagement, and the fragmented nature of data governance and technology adoption. Findings point not only to significant gaps in shared understanding, infrastructure, and policy but also to clear opportunities for collaboration, including early recommendations for building a more inclusive and patient-centered digital health ecosystem, one that supports sustainable innovation, trust, and systemwide impact.
  •  
❌