❌

Normal view

The challenge of generating and evolving real-life like synthetic test data without accessing real-world raw data -- a Systematic Review

arXiv:2602.06609v1 Announce Type: cross Abstract: Background: High-level system testing of applications that use data from e-Government services as input requires test data that is real-life-like but where the privacy of personal information is guaranteed. Applications with such strong requirement include information exchange between countries, medicine, banking, etc. This review aims to synthesize the current state-of-the-practice in this domain. Objectives: The objective of this Systematic Review is to identify existing approaches for creating and evolving synthetic test data without using real-life raw data. Methods: We followed well-known methodologies for conducting systematic literature reviews, including the ones from Kitchenham as well as guidelines for analysing the limitations of our review and its threats to validity. Results: A variety of methods and tools exist for creating privacy-preserving test data. Our search found 1,013 publications in IEEE Xplore, ACM Digital Library, and SCOPUS. We extracted data from 75 of those publications and identified 37 approaches that answer our research question partly. A common prerequisite for using these methods and tools is direct access to real-life data for data anonymization or synthetic test data generation. Nine existing synthetic test data generation approaches were identified that were closest to answering our research question. Nevertheless, further work would be needed to add the ability to evolve synthetic test data to the existing approaches. Conclusions: None of the publications really covered our requirements completely, only partially. Synthetic test data evolution is a field that has not received much attention from researchers but needs to be explored in Digital Government Solutions, especially since new legal regulations are being placed in force in many countries.

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning

arXiv:2602.05183v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly trained in complex Reinforcement Learning, multi-agent environments, making it difficult to understand how behavior changes over training. Sparse Autoencoders (SAEs) have recently shown to be useful for data-centric interpretability. In this work, we analyze large-scale reinforcement learning training runs from the sophisticated environment of Full-Press Diplomacy by applying pretrained SAEs, alongside LLM-summarizer methods. We introduce Meta-Autointerp, a method for grouping SAE features into interpretable hypotheses about training dynamics. We discover fine-grained behaviors including role-playing patterns, degenerate outputs, language switching, alongside high-level strategic behaviors and environment-specific bugs. Through automated evaluation, we validate that 90% of discovered SAE Meta-Features are significant, and find a surprising reward hacking behavior. However, through two user studies, we find that even subjectively interesting and seemingly helpful SAE features may be worse than useless to humans, along with most LLM generated hypotheses. However, a subset of SAE-derived hypotheses are predictively useful for downstream tasks. We further provide validation by augmenting an untrained agent's system prompt, improving the score by +14.2%. Overall, we show that SAEs and LLM-summarizer provide complementary views into agent behavior, and together our framework forms a practical starting point for future data-centric interpretability work on ensuring trustworthy LLM behavior throughout training.

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

Nature Medicine, Published online: 09 February 2026; doi:10.1038/s41591-025-04074-y

In a randomized controlled study involving 1,298 participants from a general sample, performance of humans when assisted by a large language model (LLM) was sensibly inferior to that of the LLM alone when assessing ten medical scenarios leading to disease identification and recommendations for treatment.

The Feasibility of Smartwatch Micro–Ecological Momentary Assessment for Tracking Eating Patterns of Malaysian Children and Adolescents in the South-East Asian Community Observatory Child Health Update 2020: Cross-Sectional Study

Background: Mobile phone ecological momentary assessment (EMA) methods are a well-established measure of eating and drinking behaviors, but compliance can be poor. Micro-EMA (μEMA), which collects information with a single tap response to brief questions on smartwatches, offers a novel application that may improve response rates. To our knowledge, there is no data evaluating μEMA to measure eating habits in children or in low-to-middle-income countries. Objective: In this study, we investigated the feasibility of micro-EMA to measure eating patterns in Malaysian children and adolescents. Methods: We invited 100 children and adolescents aged 7-18 years in Segamat, Malaysia, to participate in 2021-2022. Smartwatches were distributed to 83 children and adolescents who agreed to participate. Participants were asked to wear the smartwatch for 8 days and respond to 12 prompts per day, hourly, from 9AM to 8PM, asking for information on their meals, snacks, and drinks consumed. A questionnaire captured their experiences using the smartwatch and μEMA interface. Response rate (proportion of prompts responded to) assessed participants’ adherence. We explored associations between response rate with time of day, across days, age, and sex using multilevel binomial logistic regression modeling. Results: Eighty-two participants provided usable smartwatch data. The median number (IQR) of meals, drinks, and snacks per day was 2 (2-4), 3 (1-5), and 1 (0-2), respectively, on the first day of the study. The median response rate across the study was 68% (IQR 50-83). The response rate decreased across study days from 74% (68-78) on Day 1 to 40% (30-50) on Day 7 (odds ratio [OR] per study day 0.73, 95% CI 0.64-0.83). Response rate was lowest at the start of the day and highest between the hours of 12 PM and 2 PM. Female participants responded to more prompts than male participants (OR 1.72, 95% CI 1.03-2.86). There was no evidence of differential response by age (OR 0.73, 95% CI 0.41-1.28). Most participants (65%) rated their experience using the smartwatch positively, with 33% saying they were happy to participate in future studies using the smartwatch. For children that did not wear the smartwatch for the full study duration (n=22), discomfort was the most common complaint (41%). Conclusions: In this study of the feasibility of μEMA on smartwatches to measure eating in Malaysian children, we found the method was acceptable. However, response rates declined across study days, resulting in substantial missingness. Future studies (eg, through focus groups) should explore approaches to improving response to event prompts, trial alternative devices to increase children’s comfort, and evaluate revised protocols for reporting of intake events.

Tumor microbiome differences in early-onset versus average-onset pancreatic adenocarcinoma

ESMO Gastrointest Oncol. 2025 Jul 7;9:100194. doi: 10.1016/j.esmogo.2025.100194. eCollection 2025 Sep.

ABSTRACT

BACKGROUND: Compelling evidence supports the biomarker potential of microbiome in pancreatic adenocarcinoma. Given the knowledge gap on the characteristics and significance of microbiome in early-onset pancreatic ductal adenocarcinoma (eoPDAC, age <50 years), we aimed to evaluate microbiome profiles in resected specimens from individuals with eoPDAC and average-onset PDAC (aoPDAC, age >50 years).

MATERIALS AND METHODS: We carried out shotgun metagenomic sequencing in resected specimens from individuals with eoPDAC (n = 24) and aoPDAC (n = 20). Statistical tests included Wilcoxon test, permutational analysis of variance, multiomic classifier modeling, differential abundance analysis, and linear regression. All P values were adjusted for multiple testing and P < 0.05 was considered statistically significant.

RESULTS: We successfully sequenced several bacteria and fungi in the tumor specimens from 44 individuals with resected PDAC (24 eoPDAC and 20 aoPDAC). The alpha diversity of the bacterial microbiome was higher in eoPDAC tumor tissue compared with aoPDAC (P = 0.04). In contrast, the fungal mycobiome's alpha diversity was higher for aoPDAC tumor tissue (P = 0.02). Key organisms with differential abundance between tumor tissue from individuals with eoPDAC and aoPDAC included Bacillus, Candida, Collimonas, Cupriavidus, Enterobacter, Escherichia, Klebsiella, Malasseiza, Mucilaginibacter, Neisseria, and Sphingomonas. Higher bacterial diversity in tumor tissue was associated with better overall survival for individuals with eoPDAC (R = 0.26, P = 0.02).

CONCLUSIONS: Shotgun metagenomic sequencing identified bacterial microbiome and fungal mycobiome in tumors from individuals with eoPDAC and aoPDAC. We observed significant differences in alpha and beta diversity and relative abundances of organisms suggesting distinct microbiome signatures. Microbiome associations with survival were observed in eoPDAC indicating unique potential as prognostic biomarker.

PMID:41647993 | PMC:PMC12836659 | DOI:10.1016/j.esmogo.2025.100194

EcDNA-borne structural variants drive oncogenic fusion transcript amplification

Extrachromosomal DNA (ecDNA) is a major source of oncogenic fusions across cancer types, generating tissue-specific fusion landscapes with diagnostic potential. EcDNA-borne PVT1 5′-end fusions stabilize partner RNAs and boost oncogene output.
❌