❌

Reading view

From Data to Behavior: Predicting Unintended Model Behaviors Before Training

arXiv:2602.04735v1 Announce Type: cross Abstract: Large Language Models (LLMs) can acquire unintended biases from seemingly benign training data even without explicit cues or malicious content. Existing methods struggle to detect such risks before fine-tuning, making post hoc evaluation costly and inefficient. To address this challenge, we introduce Data2Behavior, a new task for predicting unintended model behaviors prior to training. We also propose Manipulating Data Features (MDF), a lightweight approach that summarizes candidate data through their mean representations and injects them into the forward pass of a base model, allowing latent statistical signals in the data to shape model activations and reveal potential biases and safety risks without updating any parameters. MDF achieves reliable prediction while consuming only about 20% of the GPU resources required for fine-tuning. Experiments on Qwen3-14B, Qwen2.5-32B-Instruct, and Gemma-3-12b-it confirm that MDF can anticipate unintended behaviors and provide insight into pre-training vulnerabilities.
  •  

Contrastive Continual Learning for Model Adaptability in Internet of Things

arXiv:2602.04881v1 Announce Type: cross Abstract: Internet of Things (IoT) deployments operate in nonstationary, dynamic environments where factors such as sensor drift, evolving user behavior, and heterogeneous user privacy requirements can affect application utility. Continual learning (CL) addresses this by adapting models over time without catastrophic forgetting. Meanwhile, contrastive learning has emerged as a powerful representation-learning paradigm that improves robustness and sample efficiency in a self-supervised manner. This paper reviews the usage of \emph{contrastive continual learning} (CCL) for IoT, connecting algorithmic design (replay, regularization, distillation, prompts) with IoT system realities (TinyML constraints, intermittent connectivity, privacy). We present a unifying problem formulation, derive common objectives that blend contrastive and distillation losses, propose an IoT-oriented reference architecture for on-device, edge, and cloud-based CCL, and provide guidance on evaluation protocols and metrics. Finally, we highlight open unique challenges with respect to the IoT domain, such as spanning tabular and streaming IoT data, concept drift, federated settings, and energy-aware training.
  •  

DISCOVER: Identifying Patterns of Daily Living in Human Activities from Smart Home Data

arXiv:2503.01733v3 Announce Type: replace-cross Abstract: Smart homes equipped with ambient sensors offer a transformative approach to continuous health monitoring and assisted living. Traditional research in this domain primarily focuses on Human Activity Recognition (HAR), which relies on mapping sensor data to a closed set of predefined activity labels. However, the fixed granularity of these labels often constrains their practical utility, failing to capture the subtle, household-specific nuances essential, for example, for tracking individual health over time. To address this, we propose DISCOVER, a framework for discovering and annotating Patterns of Daily Living (PDL) - fine-grained, recurring sequences of sensor events that emerge directly from a resident's unique routines. DISCOVER utilizes a self-supervised feature extraction and representation-aware clustering pipeline, supported by a custom visualization interface that enables experts to interpret and label discovered patterns with minimal effort. Our evaluation across multiple smart-home environments demonstrates that DISCOVER identifies cohesive behavioral clusters with high inter-rater agreement while achieving classification performance comparable to fully-supervised baselines using only 0.01% of the labels. Beyond reducing annotation overhead, DISCOVER establishes a foundation for longitudinal analysis. By grounding behavior in a resident's specific environment rather than rigid semantic categories, our framework facilitates the observation of within-person habitual drift. This capability positions the system as a potential tool for identifying subtle behavioral indicators associated with early-stage cognitive decline in future longitudinal studies.
  •  

From guardrails to governance: A CEO’s guide for securing agentic systems

The previous article in this series, “Rules fail at the prompt, succeed at the boundary,” focused on the first AI-orchestrated espionage campaign and the failure of prompt-level control. This article is the prescription. The question every CEO is now getting from their board is some version of: What do we do about agent risk?

Across recent AI security guidance from standards bodies, regulators, and major providers, a simple idea keeps repeating: treat agents like powerful, semi-autonomous users, and enforce rules at the boundaries where they touch identity, tools, data, and outputs.

The following is an actionable eight-step plan one can ask teams to implement and report against:  

Eight controls, three pillars: govern agentic systems at the boundary. Source: Protegrity

Constrain capabilities

These steps help define identity and limit capabilities.

1. Identity and scope: Make agents real users with narrow jobs

Today, agents run under vague, over-privileged service identities. The fix is straightforward: treat each agent as a non-human principal with the same discipline applied to employees.

Every agent should run as the requesting user in the correct tenant, with permissions constrained to that user’s role and geography. Prohibit cross-tenant on-behalf-of shortcuts. Anything high-impact should require explicit human approval with a recorded rationale. That is how Google’s Secure AI Framework (SAIF) and NIST AI’s access-control guidance are meant to be applied in practice.

The CEO question: Can we show, today, a list of our agents and exactly what each is allowed to do?

2. Tooling control: Pin, approve, and bound what agents can use

The Anthropic espionage framework worked because the attackers could wire Claude into a flexible suite of tools (e.g., scanners, exploit frameworks, data parsers) through Model Context Protocol, and those tools weren’t pinned or policy-gated.

The defense is to treat toolchains like a supply chain:

  • Pin versions of remote tool servers.
  • Require approvals for adding new tools, scopes, or data sources.
  • Forbid automatic tool-chaining unless a policy explicitly allows it.

This is exactly what OWASP flags under excessive agency and what it recommends protecting against. Under the EU AI Act, designing for such cyber-resilience and misuse resistance is part of the Article 15 obligation to ensure robustness and cybersecurity.

The CEO question: Who signs off when an agent gains a new tool or a broader scope? How does one know?

3. Permissions by design: Bind tools to tasks, not to models

A common anti-pattern is to give the model a long-lived credential and hope prompts keep it polite. SAIF and NIST argue the opposite: credentials and scopes should be bound to tools and tasks, rotated regularly, and auditable. Agents then request narrowly scoped capabilities through those tools.

In practice, that looks like: “finance-ops-agent may read, but not write, certain ledgers without CFO approval.”

The CEO question: Can we revoke a specific capability from an agent without re-architecting the whole system?

Control data and behavior

These steps gate inputs, outputs, and constrain behavior.

4. Inputs, memory, and RAG: Treat external content as hostile until proven otherwise

Most agent incidents start with sneaky data: a poisoned web page, PDF, email, or repository that smuggles adversarial instructions into the system. OWASP’s prompt-injection cheat sheet and OpenAI’s own guidance both insist on strict separation of system instructions from user content and on treating unvetted retrieval sources as untrusted.

Operationally, gate before anything enters retrieval or long-term memory: new sources are reviewed, tagged, and onboarded; persistent memory is disabled when untrusted context is present; provenance is attached to each chunk.

The CEO question: Can we enumerate every external content source our agents learn from, and who approved them?

5. Output handling and rendering: Nothing executes “just because the model said so”

In the Anthropic case, AI-generated exploit code and credential dumps flowed straight into action. Any output that can cause a side effect needs a validator between the agent and the real world. OWASP’s insecure output handling category is explicit on this point, as are browser security best practices around origin boundaries.

The CEO question: Where, in our architecture, are agent outputs assessed before they run or ship to customers?

6. Data privacy at runtime: Protect the data first, then the model

Protect the data such that there is nothing dangerous to reveal by default. NIST and SAIF both lean toward “secure-by-default” designs where sensitive values are tokenized or masked and only re-hydrated for authorized users and use cases.

In agentic systems, that means policy-controlled detokenization at the output boundary and logging every reveal. If an agent is fully compromised, the blast radius is bounded by what the policy lets it see.

This is where the AI stack intersects not just with the EU AI Act but with GDPR and sector-specific regimes. The EU AI Act expects providers and deployers to manage AI-specific risk; runtime tokenization and policy-gated reveal are strong evidence that one is actively controlling those risks in production.

The CEO question: When our agents touch regulated data, is that protection enforced by architecture or by promises?

Prove governance and resilience

For the final steps, it’s important to show controls work and keep working.

7. Continuous evaluation: Don’t ship a one-time test, ship a test harness

Anthropic’s research about sleeper agents should eliminate all fantasies about single test dreams and show how critical continuous evaluation is. This means instrumenting agents with deep observability, regularly red teaming with adversarial test suites, and backing everything with robust logging and evidence, so failures become both regression tests and enforceable policy updates.

The CEO question: Who works to break our agents every week, and how do their findings change policy?

 8. Governance, inventory, and audit: Keep score in one place

AI security frameworks emphasize inventory and evidence: enterprises must know which models, prompts, tools, datasets, and vector stores they have, who owns them, and what decisions were taken about risk.

For agents, that means a living catalog and unified logs:

  • Which agents exist, on which platforms
  • What scopes, tools, and data each is allowed
  • Every approval, detokenization, and high-impact action, with who approved it and when

The CEO question: If asked how an agent made a specific decision, could we reconstruct the chain?

And don’t forget the system-level threat model: assume the threat actor GTG-1002 is already in your enterprise. To complete enterprise preparedness, zoom out and consider the MITRE ATLAS product, which exists precisely because adversaries attack systems, not models. Anthropic provides a case study of a state-based threat actor (GTG-1002) doing exactly that with an agentic framework.

Taken together, these controls do not make agents magically safe. They do something more familiar and more reliable: they put AI, its access, and actions back inside the same security frame used for any powerful user or system.

For boards and CEOs, the question is no longer “Do we have good AI guardrails?” It’s: Can we answer the CEO questions above with evidence, not assurances?

This content was produced by Protegrity. It was not written by MIT Technology Review’s editorial staff.

  •  

Phenome-wide analysis of copy number variants in 470,727 UK Biobank genomes

Nature, Published online: 04 February 2026; doi:10.1038/s41586-025-10087-x

A multiancestry phenome-wide analysis of copy number variants in the UK Biobank genomes increases power to detect genetic associations with complex traits across human populations.
  •  

Extrachromosomal DNA drives molecular and clinical heterogeneity in hepatocellular carcinoma: a multi-omics analysis and prognostic model development

Hum Genomics. 2026 Feb 3. doi: 10.1186/s40246-026-00927-w. Online ahead of print.

ABSTRACT

BACKGROUND: Extrachromosomal DNA (ecDNA) is an emerging hallmark of cancer that promotes tumor evolution and heterogeneity. However, the molecular characteristics and clinical significance of ecDNA in hepatocellular carcinoma (HCC) remain incompletely understood.

METHODS: The clinical outcomes, genomics, transcriptomics, proteomics, tumor microenvironment, and drug target landscapes of ecDNA-negative and ecDNA-positive HCC in the Cancer Genome Atlas (TCGA) were compared. Next, the least absolute shrinkage and selection operator (LASSO) and random survival forest (RSF) algorithms were used to screen the ecDNA gene signature. A nomogram was constructed and evaluated based on the risk score and clinicopathological features. Finally, the role of DNASE1L3 was validated through in vitro experiments.

RESULTS: EcDNA-positive tumors showed increased vascular invasion, higher AFP levels, and more TP53 mutations. These tumors displayed unique activation of proliferation pathways, decreased stromal infiltration, and heightened immune activation. Our validated six-gene signature (RNF186, BMP6, AOC1, FBLL1, MYBL2, and DNASE1L3) demonstrated strong prognostic value when combined with tumor stage in the nomogram. Notably, DNASE1L3 was downregulated in HCC, showed endothelial cell-specific expression, and suppressed the proliferation and migration of Hep3B2.1-7 cells.

CONCLUSION: Our study characterizes the molecular and clinical distinctions between ecDNA-negative and ecDNA-positive HCC and establishes a clinically applicable gene signature for patient prognosis. These findings advance our understanding of ecDNA-driven tumor heterogeneity and provide potential strategies for personalized HCC management.

PMID:41634868 | DOI:10.1186/s40246-026-00927-w

  •  

Trustworthy Blockchain-based Federated Learning for Electronic Health Records: Securing Participant Identity with Decentralized Identifiers and Verifiable Credentials

arXiv:2602.02629v1 Announce Type: cross Abstract: The digitization of healthcare has generated massive volumes of Electronic Health Records (EHRs), offering unprecedented opportunities for training Artificial Intelligence (AI) models. However, stringent privacy regulations such as GDPR and HIPAA have created data silos that prevent centralized training. Federated Learning (FL) has emerged as a promising solution that enables collaborative model training without sharing raw patient data. Despite its potential, FL remains vulnerable to poisoning and Sybil attacks, in which malicious participants corrupt the global model or infiltrate the network using fake identities. While recent approaches integrate Blockchain technology for auditability, they predominantly rely on probabilistic reputation systems rather than robust cryptographic identity verification. This paper proposes a Trustworthy Blockchain-based Federated Learning (TBFL) framework integrating Self-Sovereign Identity (SSI) standards. By leveraging Decentralized Identifiers (DIDs) and Verifiable Credentials (VCs), our architecture ensures only authenticated healthcare entities contribute to the global model. Through comprehensive evaluation using the MIMIC-IV dataset, we demonstrate that anchoring trust in cryptographic identity verification rather than behavioral patterns significantly mitigates security risks while maintaining clinical utility. Our results show the framework successfully neutralizes 100% of Sybil attacks, achieves robust predictive performance (AUC = 0.954, Recall = 0.890), and introduces negligible computational overhead (
  •  

Digital intervention <i>mylovia</i> improves sexual functioning in women with sexual dysfunction in randomized controlled trial

npj Digital Medicine, Published online: 03 February 2026; doi:10.1038/s41746-026-02385-z

Digital intervention mylovia improves sexual functioning in women with sexual dysfunction in randomized controlled trial
  •  

Integrative proteogenomics maps multifactorial aetiology, progression and therapeutic vulnerabilities in gastric cancer

Gut. 2026 Jan 30:gutjnl-2025-337247. doi: 10.1136/gutjnl-2025-337247. Online ahead of print.

ABSTRACT

BACKGROUND: Gastric cancer, with disproportionately higher incidence in East Asia, arises from complex host-microbiome-environment interactions beyond Helicobacter pylori (HP) infection. However, the molecular architecture linking environmental carcinogens, microbial succession and host response remains unclear.

OBJECTIVE: To delineate multifactorial aetiologies and clinically actionable subtypes/biomarkers of gastric cancer through integrative proteogenomic, microbial and environmental exposure profiling.

DESIGN: We established a multiomics atlas of paired tumour, adjacent mucosa tissues and blood from 154 treatment-naïve Taiwanese patients, integrating whole-exome sequencing, RNA-seq, proteome and phosphoproteome profiling with carcinogen signatures, HP status, microbiome composition and refined anatomical mapping. Cell-based functional assays tested carcinogen effects. Microbial subtype was assessed in an independent cohort.

RESULTS: A polycyclic-aromatic-hydrocarbon signature, dibenz[a,h]acridine, emerged as a high-risk exposure promoting invasion, immune suppression and poor survival, significantly exceeding nitrosamine-linked risk in this cohort. Multilayer integration defined three initiation ecologies: HP-driven inflammatory, non-HP microbiome-enriched immune-silent and HP-free microbially depleted states. Among HP-negative tumours, a Streptococcus-enriched subtype associated with tight-junction (CLDN18.2/ZO-1/OCLN) disruption and epithelial-mesenchymal transition, whereas a subset of clinically aggressive cases retained CLDN18.2-high epithelial-stable subtype for therapeutic accessibility. An independent cohort revealed gastric juice-derived Streptococcus anginosus abundance inversely correlated with tight-junction proteins. Anatomical mapping reveals location-specific, sex-specific, subtype-specific oncogenic networks and kinase activity, including CDK4 activation in clinical biomarker-negative tumours. Decision-tree models combining exposure and proteome-immune states refined recurrence and survival prediction beyond stage.

CONCLUSION: This proteogenomic framework defines exposure-informed and microbiome-informed gastric cancer subtypes, providing a molecular schema for patient stratification, prevention and actionable therapeutic vulnerabilities.

PMID:41617485 | DOI:10.1136/gutjnl-2025-337247

  •  

Liquid biopsy biomarkers for accurate detection of malignant pulmonary nodules: a meta-analytic approach

Discov Oncol. 2026 Jan 29;17(1):178. doi: 10.1007/s12672-025-03646-1.

ABSTRACT

Pulmonary nodules are a common radiological finding that can be classified as either benign or Malignant, with significant clinical implications. The early detection of malignant nodules is critically important for improving the prognosis of lung cancer, which remains the leading cause of cancer-related mortality worldwide. Traditional imaging techniques have Limitations in accurately classifying pulmonary nodules. Liquid biopsy, a minimally invasive method that evaluates circulating components in the Blood, presents promising diagnostic potential in this context. This study aims to evaluate the diagnostic capacity of multiple liquid biopsy biomarkers for early and accurate differentiation between benign and Malignant pulmonary nodules. Accordingly, we conducted a comprehensive study involving a meta-analysis, selecting 16 eligible studies that utilised liquid biopsy to assess various circulating biomarkers in the diagnostic yield. The most significant results were linked to circulating free DNA (cfDNA). However, other components, including circulating tumour cells (CTCs), microRNAs/pfeRNAs, extracellular vesicles (EVs), serological markers, and imaging techniques, also provided valuable information. Similarly, integrating multi-omics data with machine learning models has been shown to enhance the ability to differentiate between benign and malignant pulmonary nodules, thereby supporting early diagnosis and improved management for patients with lung cancer.

PMID:41612093 | PMC:PMC12855667 | DOI:10.1007/s12672-025-03646-1

  •  

The Relationship Between Physician Self-Disclosure and Patient Acquisition in Digital Health Markets: Cross-Sectional Study

Background: Online health communities have evolved into digital marketplaces where physicians have to compete for patients. Existing research examines physician-patient dynamics through a patient-centric lens, treating physicians as passive recipients of ratings and reviews, while the strategic role of physician self-disclosure remains unexamined. This gap constrains a comprehensive understanding of how physicians can actively shape patient decisions, making the investigation of strategic self-disclosure imperative. Objective: This study aims to investigate the relationship between physician self-disclosure breadth (scope of information) and depth (detailed expertise) and patient decision-making, as well as whether regional digital health care level (DHL) moderates these relationships. Methods: We conducted a cross-sectional analysis of observational data to test these relationships. Data were collected from China’s online health care platform Haodf from September to December 2024. Self-disclosure breadth (including clinical performance, academic experience, and social reputation), self-disclosure depth (including expertise coverage, richness, and granularity), and patient decision-making (total visits) were captured through manual content coding and quantitative measurement. We used structured content analysis to extract the disclosure components, informational scope, and descriptive details of each profile. Then, using validated operational formulas, we calculated the composite indices for disclosure breadth and depth based on the coded dimensions. The study generated 1798 final physician samples with complete data across 14 focal variables. The hypotheses were tested using an ordinary least squares regression model, and 4 robustness checks were conducted, including variable substitution and different resampling techniques. Results: In the primary ordinary least squares regression models, self-disclosure breadth was significantly and positively associated with patient visits (β=0.255, 95% CI 0.054-0.456; P=.01), as was self-disclosure depth (β=0.098, 95% CI 0.030-0.167; P=.005). The breadth×DHL interaction was positive and significant (β=0.261, 95% CI 0.061-0.461; P=.01). Similarly, the depth×DHL interaction was positive and significant (β=0.070, 95% CI 0.002-0.138; P=.045). It should be noted that the association for self-disclosure breadth was stronger than that of self-disclosure depth. DHL strengthened the relationship between the disclosure strategies with patient visits. This contextual amplification indicates that DHL serves as a critical boundary condition, determining the degree to which physician self-disclosure strategies translate into patient acquisition outcomes. Conclusions: This study reconceptualizes physicians as strategic agents shaping patient decision-making through purposeful self-disclosure. Different from existing studies treating physicians as passive recipients of ratings and reviews, our research demonstrates that physicians can strategically shape patient acquisition through self-disclosure breadth and depth. This study brings new insights to digital health markets by demonstrating that self-disclosure operates as a viable patient acquisition mechanism, wherein the DHL acts as a critical boundary condition. The findings have real-world implications: (1) physicians can leverage evidence-based disclosure strategies, (2) platforms should implement context-adaptive features, and (3) policymakers should prioritize digital infrastructure investments to enhance physicians' competitive capabilities and patient decision-making quality.
  •  

Liquid biopsy in cancer diagnosis and prognosis: a paradigm shift in precision oncology

Front Mol Biosci. 2026 Jan 12;12:1708518. doi: 10.3389/fmolb.2025.1708518. eCollection 2025.

ABSTRACT

Liquid biopsy has emerged as a transformative tool in precision oncology, offering a minimally invasive approach for cancer detection, monitoring, and treatment guidance. Unlike traditional tissue biopsies, which are invasive and limited by tumor accessibility and sampling bias, liquid biopsy enables real-time tumor assessment through the analysis of circulating biomarkers in blood and other biofluids. This review provides a comprehensive overview of recent advances in liquid biopsy, with a focus on circulating tumor cells (CTCs), circulating tumor DNA (ctDNA), non-coding RNAs, extracellular vesicles (exosomes), and secreted proteins. These biomarkers offer valuable insights into tumor biology, supporting applications in early diagnosis, prognosis, treatment response monitoring, and minimal residual disease detection across various cancer types. We also discuss state-of-the-art methodologies, including next-generation sequencing, digital PCR, microfluidics, proteomics, and emerging artificial intelligence-based approaches that enhance the sensitivity, specificity, and scalability of liquid biopsy assays. Clinical studies demonstrate the potential of liquid biopsy for tailoring targeted therapies, predicting resistance mechanisms, and identifying tumor recurrence earlier than conventional methods. Furthermore, FDA-approved assays and ongoing phase III and IV clinical trials highlight its growing integration into routine clinical practice. Beyond technical innovations, this review examines the global landscape of liquid biopsy, emphasizing opportunities and challenges for implementation across diverse healthcare settings. Disparities in access, particularly between high-income and low- and middle-income countries, underscore the need for strategies that ensure equitable adoption of liquid biopsy technologies worldwide. In summary, liquid biopsy represents a paradigm shift in oncology, bridging innovations in cancer diagnostics with clinical applications. By enabling dynamic, personalized, and less invasive cancer management, it holds great promise for improving patient outcomes and advancing precision medicine.

PMID:41602544 | PMC:PMC12832364 | DOI:10.3389/fmolb.2025.1708518

  •  

Agentic Digital Twins: A Taxonomy of Capabilities for Understanding Possible Futures

arXiv:2601.18799v1 Announce Type: cross Abstract: As digital twins (DTs) evolve to become more agentic through the integration of artificial intelligence (AI), they acquire capabilities that extend beyond dynamic representation of their target systems. This paper presents a taxonomy of agentic DTs organised around three fundamental dimensions: the locus of agency (external, internal, distributed), the tightness of coupling (loose, tight, constitutive), and model evolution (static, adaptive, reconstructive). From the resulting 27-configuration space, we identify nine illustrative configurations grouped into three clusters: "The Present" (existing tools and emerging steering systems), "The Threshold" (where emergent properties appear and coupling becomes constitutive), and "The Frontier" (where systems gain reconstructive capabilities). Our analysis explores how agentic DTs exercise performative power--not merely representing physical systems but actively participating in constituting them. Using traffic navigation systems as examples, we show how even passive tools can exhibit emergent performativity, while advanced configurations risk performative lock-in. Drawing on performative prediction theory, we trace a progression from passive tools through active steering to ontological reconstruction, examining how constitutive coupling enables systems to create self-validating realities. Understanding these configurations is essential for navigating the transformation from DTs as mirror worlds to DTs as architects of new ontologies.
  •  

Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?

arXiv:2508.10530v2 Announce Type: replace Abstract: The alignment of language models~(LMs) with human preferences is critical for building reliable AI systems. The problem is typically framed as optimizing an LM policy to maximize the expected reward that reflects human preferences. Recently, Direct Preference Optimization~(DPO) was proposed as a LM alignment method that directly optimize the policy from static preference data, and further improved by incorporating on-policy sampling~(i.e., preference candidates generated during the training loop) for better LM alignment. However, we show on-policy data is not always optimal, with systematic effectiveness difference emerging between static and on-policy preference candidates. For example, on-policy data can result in a $3\times$ effectiveness compared with static data for Llama-3, and a $0.4\times$ effectiveness for Zephyr. To explain the phenomenon, we propose the alignment stage assumption, which divides the alignment process into two distinct stages: the preference injection stage, which benefits from diverse data, and the preference fine-tuning stage, which favors high-quality data. Through theoretical and empirical analysis, we characterize these stages and propose an effective algorithm to identify the boundaries between them. We perform experiments on $5$ models~(Llama, Zephyr, Phi-2, Qwen, Pythia) and $2$ alignment methods~(DPO, SLiC-HF) to show the generalizability of alignment stage assumption and the effectiveness of the boundary measurement algorithm.
  •  

Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery

arXiv:2601.12542v2 Announce Type: replace Abstract: Artificial intelligence systems for scientific discovery have demonstrated remarkable potential, yet existing approaches remain largely proprietary and operate in batch-processing modes requiring hours per research cycle, precluding real-time researcher guidance. This paper introduces Deep Research, a multi-agent system enabling interactive scientific investigation with turnaround times measured in minutes. The architecture comprises specialized agents for planning, data analysis, literature search, and novelty detection, unified through a persistent world state that maintains context across iterative research cycles. Two operational modes support different workflows: semi-autonomous mode with selective human checkpoints, and fully autonomous mode for extended investigations. Evaluation on the BixBench computational biology benchmark demonstrated state-of-the-art performance, achieving 48.8% accuracy on open response and 64.4% on multiple-choice evaluation, exceeding existing baselines by 14 to 26 percentage points. Analysis of architectural constraints, including open access literature limitations and challenges inherent to automated novelty assessment, informs practical deployment considerations for AI-assisted scientific workflows.
  •  

General Binding Affinity Guidance for Diffusion Models in Structure-Based Drug Design

arXiv:2406.16821v2 Announce Type: replace-cross Abstract: Structure-based drug design (SBDD) aims to generate ligands that bind strongly and specifically to target protein pockets. Recent diffusion models have advanced SBDD by capturing the distributions of atomic positions and types, yet they often underemphasize binding affinity control during generation. To address this limitation, we introduce \textbf{\textnormal{\textbf{BADGER}}}, a general \textbf{binding-affinity guidance framework for diffusion models in SBDD}. \textnormal{\textbf{BADGER} }incorporates binding affinity awareness through two complementary strategies: (1) \textit{classifier guidance}, which applies gradient-based affinity signals during sampling in a plug-and-play fashion, and (2) \textit{classifier-free guidance}, which integrates affinity conditioning directly into diffusion model training. Together, these approaches enable controllable ligand generation guided by binding affinity. \textnormal{\textbf{BADGER} } can be added to any diffusion model and achieves up to a \textbf{60\% improvement in ligand--protein binding affinity} of sampled molecules over prior methods. Furthermore, we extend the framework to \textbf{multi-constraint diffusion guidance}, jointly optimizing for binding affinity, drug-likeness (QED), and synthetic accessibility (SA) to design realistic and synthesizable drug candidates.
  •  

AI-generated data contamination erodes pathological variability and diagnostic reliability

arXiv:2601.12946v3 Announce Type: replace-cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.
  •  

Products, Performance, and Technological Development of Ambulatory Oxygen Therapy Devices: Scoping Review

Background: Ambulatory oxygen therapy is prescribed for patients with chronic lung diseases who experience exertional hypoxemia. However, available devices may not adequately meet user requirements, and their performance characteristics are heterogeneous. Objective: This study aims to identify devices available for delivery of ambulatory oxygen therapy, the technologies that they use to generate oxygen, the performance characteristics of each device, and the development status. Methods: We used medical and engineering databases to identify peer-reviewed papers (eg, MEDLINE, IEEE). Gray literature was used to identify additional descriptions of ambulatory oxygen devices in military medicine, space exploration, or patents. The last search was conducted in September 2025. Documents that described a device that can deliver oxygen in an ambulatory context (defined as weighing less than 10 kg) and were written in English were included. Search results were screened for inclusion by 2 independent reviewers. Data were synthesized by descriptively mapping the performance of each product, the technology used, and the development status of emerging technologies. Results: From 9702 records identified, a total of 166 met eligibility criteria (106 scientific publications and 60 gray literature). We identified 33 portable oxygen concentrators (POCs; 29 commercially available), 10 oxygen cylinders, and 6 portable liquid oxygen (LOX) devices. The POC products showed a trade-off between portability and oxygen delivery capacity (maximum flow rate ranging from 2.0 to 6.0 L/min; device weight ranging from 1.0 to 9.1 kg). Pressure swing adsorption with zeolite was the most common oxygen generation technology in POCs on the market. The mean maximum continuous operating time of POCs was 3.8 hours. Two prototype POCs (maximum flow rate of 4-6 L/min and device weight of 8-9 kg) were developed for space exploration using modified adsorbents. LOX devices were the lightest and had the longest continuous operating time. Innovations in delivery included the downsizing of a POC by using nanozeolite as an adsorbent and pulse oximeter oxygen saturation (SpO2)–targeted automatic titration of oxygen delivery based on the user’s SpO2. Conclusions: This scoping review is the first study to integrate medical, engineering, and gray literature on ambulatory oxygen devices and their development. Although prior literature has narratively explained the products and technologies, no previous research has systematically investigated them. This review showed that POCs available to consumers may not meet the needs of patients in terms of flow rate, portability, and operating time. LOX devices offered superior performance but are limited by high costs. Limitations of this review include the difficulty of comparing product performance across oxygen delivery settings and that the records were largely obtained from English-language sources. Innovation in ambulatory oxygen technology has been limited over the past decade, highlighting urgent need for research and development of new lightweight devices with higher oxygen delivery. Clinical Trial: OSF Registries 10.17605/OSF.IO/QS7FX; https://osf.io/qs7fx
  •  
❌