❌

Normal view

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

arXiv:2602.06609v1 Announce Type: cross Abstract: Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.

Exploring AI-Augmented Sensemaking of Patient-Generated Health Data: A Mixed-Method Study with Healthcare Professionals in Cardiac Risk Reduction

arXiv:2602.05687v2 Announce Type: replace-cross Abstract: Individuals are increasingly generating substantial personal health and lifestyle data, e.g. through wearables and smartphones. While such data could transform preventative care, its integration into clinical practice is hindered by its scale, heterogeneity and the time pressure and data literacy of healthcare professionals (HCPs). We explore how large language models (LLMs) can support sensemaking of patient-generated health data (PGHD) with automated summaries and natural language data exploration. Using cardiovascular disease (CVD) risk reduction as a use case, 16 HCPs reviewed multimodal PGHD in a mixed-methods study with a prototype that integrated common charts, LLM-generated summaries, and a conversational interface. Findings show that AI summaries provided quick overviews that anchored exploration, while conversational interaction supported flexible analysis and bridged data-literacy gaps. However, HCPs raised concerns about transparency, privacy, and overreliance. We contribute empirical insights and sociotechnical design implications for integrating AI-driven summarization and conversation into clinical workflows to support PGHD sensemaking.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Nature Medicine, Published online: 09 February 2026; doi:10.1038/s41591-025-04074-y

In a randomized controlled study involving 1,298 participants from a general sample, performance of humans when assisted by a large language model (LLM) was sensibly inferior to that of the LLM alone when assessing ten medical scenarios leading to disease identification and recommendations for treatment.

Integrating liquid biopsies in non-small cell lung cancer diagnosis and management: opportunities and challenges

8 February 2026 at 19:00

Expert Rev Anticancer Ther. 2026 Feb 15:1-12. doi: 10.1080/14737140.2026.2630026. Online ahead of print.

ABSTRACT

INTRODUCTION: Liquid biopsy has emerged as an important approach to capture tumor-derived material from blood and other body fluids, offering a minimally invasive window into cancer biology. In non - small cell lung cancer (NSCLC), it enables comprehensive molecular profiling that informs patient management, from guiding therapy choices to monitoring disease status and assessing minimal residual disease (MRD).

AREAS COVERED: Its main advantages over tissue biopsy lie in being noninvasive, capable of reflecting tumor heterogeneity and real-time biological changes. These strengths allow liquid biopsy to be applied at different clinical timepoints, including diagnosis, treatment decision-making, evaluation during therapy, detection of resistance, and surveillance for recurrence. Although circulating tumor DNA (ctDNA) remains the most established analyte, the scope is broadening to include circulating RNAs, circulating tumor cells, exosomes, DNA methylation signatures, and tumor-educated platelets, each providing complementary insights. A literature search of PubMed, EMBASE, and Web of Science was conducted without restrictions, supplemented by screening reference lists and major oncology conference abstracts.

EXPERT OPINION: While significant progress has been made integrating liquid biopsies in NSCLC, challenges persist, encompassing issues of standardization, cost, and clinical integration.

PMID:41656166 | DOI:10.1080/14737140.2026.2630026

The Feasibility of Smartwatch Micro–Ecological Momentary Assessment for Tracking Eating Patterns of Malaysian Children and Adolescents in the South-East Asian Community Observatory Child Health Update 2020: Cross-Sectional Study

Background: Mobile phone ecological momentary assessment (EMA) methods are a well-established measure of eating and drinking behaviors, but compliance can be poor. Micro-EMA (μEMA), which collects information with a single tap response to brief questions on smartwatches, offers a novel application that may improve response rates. To our knowledge, there is no data evaluating μEMA to measure eating habits in children or in low-to-middle-income countries. Objective: In this study, we investigated the feasibility of micro-EMA to measure eating patterns in Malaysian children and adolescents. Methods: We invited 100 children and adolescents aged 7-18 years in Segamat, Malaysia, to participate in 2021-2022. Smartwatches were distributed to 83 children and adolescents who agreed to participate. Participants were asked to wear the smartwatch for 8 days and respond to 12 prompts per day, hourly, from 9AM to 8PM, asking for information on their meals, snacks, and drinks consumed. A questionnaire captured their experiences using the smartwatch and μEMA interface. Response rate (proportion of prompts responded to) assessed participants’ adherence. We explored associations between response rate with time of day, across days, age, and sex using multilevel binomial logistic regression modeling. Results: Eighty-two participants provided usable smartwatch data. The median number (IQR) of meals, drinks, and snacks per day was 2 (2-4), 3 (1-5), and 1 (0-2), respectively, on the first day of the study. The median response rate across the study was 68% (IQR 50-83). The response rate decreased across study days from 74% (68-78) on Day 1 to 40% (30-50) on Day 7 (odds ratio [OR] per study day 0.73, 95% CI 0.64-0.83). Response rate was lowest at the start of the day and highest between the hours of 12 PM and 2 PM. Female participants responded to more prompts than male participants (OR 1.72, 95% CI 1.03-2.86). There was no evidence of differential response by age (OR 0.73, 95% CI 0.41-1.28). Most participants (65%) rated their experience using the smartwatch positively, with 33% saying they were happy to participate in future studies using the smartwatch. For children that did not wear the smartwatch for the full study duration (n=22), discomfort was the most common complaint (41%). Conclusions: In this study of the feasibility of μEMA on smartwatches to measure eating in Malaysian children, we found the method was acceptable. However, response rates declined across study days, resulting in substantial missingness. Future studies (eg, through focus groups) should explore approaches to improving response to event prompts, trial alternative devices to increase children’s comfort, and evaluate revised protocols for reporting of intake events.

EcDNA-borne structural variants drive oncogenic fusion transcript amplification

Extrachromosomal DNA (ecDNA) is a major source of oncogenic fusions across cancer types, generating tissue-specific fusion landscapes with diagnostic potential. EcDNA-borne PVT1 5′-end fusions stabilize partner RNAs and boost oncogene output.
❌