❌

Normal view

Japanese AI Agent System on Human Papillomavirus Vaccination: System Design

arXiv:2601.10718v1 Announce Type: new Abstract: Human papillomavirus (HPV) vaccine hesitancy poses significant public health challenges, particularly in Japan where proactive vaccination recommendations were suspended from 2013 to 2021. The resulting information gap is exacerbated by misinformation on social media, and traditional ways cannot simultaneously address individual queries while monitoring population-level discourse. This study aimed to develop a dual-purpose AI agent system that provides verified HPV vaccine information through a conversational interface while generating analytical reports for medical institutions based on user interactions and social media. We implemented a system comprising: a vector database integrating academic papers, government sources, news media, and social media; a Retrieval-Augmented Generation chatbot using ReAct agent architecture with multi-tool orchestration across five knowledge sources; and an automated report generation system with modules for news analysis, research synthesis, social media sentiment analysis, and user interaction pattern identification. Performance was assessed using a 0-5 scoring scale. For single-turn evaluation, the chatbot achieved mean scores of 4.83 for relevance, 4.89 for routing, 4.50 for reference quality, 4.90 for correctness, and 4.88 for professional identity (overall 4.80). Multi-turn evaluation yielded higher scores: context retention 4.94, topic coherence 5.00, and overall 4.98. The report generation system achieved completeness 4.00-5.00, correctness 4.00-5.00, and helpfulness 3.67-5.00, with reference validity 5.00 across all periods. This study demonstrates the feasibility of an integrated AI agent system for bidirectional HPV vaccine communication. The architecture enables verified information delivery with source attribution while providing systematic public discourse analysis, with a transferable framework for adaptation to other medical contexts.

AnyECG: Evolved ECG Foundation Model for Holistic Health Profiling

arXiv:2601.10748v1 Announce Type: cross Abstract: Background: Artificial intelligence enabled electrocardiography (AI-ECG) has demonstrated the ability to detect diverse pathologies, but most existing models focus on single disease identification, neglecting comorbidities and future risk prediction. Although ECGFounder expanded cardiac disease coverage, a holistic health profiling model remains needed. Methods: We constructed a large multicenter dataset comprising 13.3 million ECGs from 2.98 million patients. Using transfer learning, ECGFounder was fine-tuned to develop AnyECG, a foundation model for holistic health profiling. Performance was evaluated using external validation cohorts and a 10-year longitudinal cohort for current diagnosis, future risk prediction, and comorbidity identification. Results: AnyECG demonstrated systemic predictive capability across 1172 conditions, achieving an AUROC greater than 0.7 for 306 diseases. The model revealed novel disease associations, robust comorbidity patterns, and future disease risks. Representative examples included high diagnostic performance for hyperparathyroidism (AUROC 0.941), type 2 diabetes (0.803), Crohn disease (0.817), lymphoid leukemia (0.856), and chronic obstructive pulmonary disease (0.773). Conclusion: The AnyECG foundation model provides substantial evidence that AI-ECG can serve as a systemic tool for concurrent disease detection and long-term risk prediction.

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

arXiv:2601.11505v1 Announce Type: cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D management datasets. Current datasets differ substantially in structure and are time-consuming to access and process, which impedes data integration and reduces the comparability and generalizability of algorithmic developments. This work aims to establish a unified and accessible data resource for T1D algorithm development. Multiple publicly available T1D datasets were consolidated into a unified resource, termed the MetaboNet dataset. Inclusion required the availability of both continuous glucose monitoring (CGM) data and corresponding insulin pump dosing records. Additionally, auxiliary information such as reported carbohydrate intake and physical activity was retained when present. The MetaboNet dataset comprises 3135 subjects and 1228 patient-years of overlapping CGM and insulin data, making it substantially larger than existing standalone benchmark datasets. The resource is distributed as a fully public subset available for immediate download at https://metabo-net.org/ , and with a Data Use Agreement (DUA)-restricted subset accessible through their respective application processes. For the datasets in the latter subset, processing pipelines are provided to automatically convert the data into the standardized MetaboNet format. A consolidated public dataset for T1D research is presented, and the access pathways for both its unrestricted and DUA-governed components are described. The resulting dataset covers a broad range of glycemic profiles and demographics and thus can yield more generalizable algorithmic performance than individual datasets.

Wearable device derived electrocardiographic age and its association with atrial fibrillation

npj Digital Medicine, Published online: 17 January 2026; doi:10.1038/s41746-026-02344-8

Wearable device derived electrocardiographic age and its association with atrial fibrillation

From OpenAI’s offices to a deal with Eli Lilly β€” how Chai Discovery became one of the flashiest names in AI drug development

17 January 2026 at 04:14
The startup has partnered with Eli Lilly and enjoys the backing of some of Silicon Valley's most influential VCs.
  • βœ‡TechCrunch
  • The AI healthcare gold rush is here Theresa Loconsolo
    AI companies are clustering around healthcare and fast.Β  In just the past week, OpenAIΒ bought health startup Torch, Anthropic launchedΒ Claude for healthcare, and Sam Altman-backedΒ MergeLabs closed a $250 million seed roundΒ at an $850 million valuation. The money and products are pouring into healthΒ and voice AI, but so are concerns about hallucination risks, inaccurate medical information, and […]
     

The AI healthcare gold rush is here

17 January 2026 at 02:57
AI companies are clustering around healthcare and fast.Β  In just the past week, OpenAIΒ bought health startup Torch, Anthropic launchedΒ Claude for healthcare, and Sam Altman-backedΒ MergeLabs closed a $250 million seed roundΒ at an $850 million valuation. The money and products are pouring into healthΒ and voice AI, but so are concerns about hallucination risks, inaccurate medical information, and […]

Cellular neighborhoods in cancer

Nat Cancer. 2026 Jan 16. doi: 10.1038/s43018-025-01107-w. Online ahead of print.

ABSTRACT

The concept of cellular neighborhoods, defined as recurring structures within the tissue with characteristic cell compositions and interactions, has transformed our understanding of the complexity and dynamics of tumor ecosystems. Recent advances in spatial omics and computational modeling have enabled high-resolution mapping of these neighborhoods, providing unprecedented insights into their roles in shaping tumor heterogeneity, evolution and therapeutic responses. Despite these advances, a unified framework for interpreting cellular neighborhoods remains lacking. This Perspective synthesizes emerging concepts and insights, focusing on the definition and classification of cellular neighborhoods in cancer, computational methods for identifying and comparing them, and their clinical relevance.

PMID:41545713 | DOI:10.1038/s43018-025-01107-w

Contaminating plasmid sequences and disrupted vector genomes in the liver following adeno-associated virus gene therapy

Nature Medicine, Published online: 16 January 2026; doi:10.1038/s41591-025-04073-z

Analyses of liver biopsies from a child with spinal muscular atrophy treated with adeno-associated virus gene therapy who developed hepatitis reveal contaminating manufacturing plasmids and disrupted vector genomes, possibly resulting from recombination events.

Clinical proteomics in cardiovascular medicine: Current capabilities, limitations, and future directions

Atherosclerosis. 2026 Jan 8;413:120637. doi: 10.1016/j.atherosclerosis.2026.120637. Online ahead of print.

ABSTRACT

BACKGROUND AND AIMS: Commercial high-throughput proteomics platforms, such as Olink and SomaLogic, enable large-scale epidemiological studies with integrated multi-omics measurements. While these proteomics approaches have been widely applied in biobanks, issues of data quality remain underappreciated. In this review, we discuss these limitations and outline a way forward for realizing the clinical translation of proteomics as a comprehensive 'liquid health check'.

METHODS: We reviewed the recent literature for artificial intelligence (AI) and multi-omics, particularly proteomics in atherosclerotic cardiovascular disease (ASCVD).

RESULTS: AI-driven multi-omics analyses have the potential to advance our understanding of multifactorial causes of ASCVD, including aging. Emerging concepts such as "ageotypes" suggest the potential for personalized intervention to slow aging processes. Commercial proteomics platforms have accelerated biomarker discovery in ASCVD, but challenges remain in clinical translation. Limited correlation between Olink and SomaLogic necessitates orthogonal validation of findings. Platform-specific issues, such as epitope effects and cross-reactivity, can yield divergent protein quantitative trait loci for the same protein, complicating causal inference. While tissue proteomics provides complementary insights to plasma proteomics, reliance on autopsy samples raises concerns about protein degradation and measurement reliability. Increasingly, single-cell and spatial proteomics are being explored to better capture plaque heterogeneity, complementing bulk proteomics in larger cohorts.

CONCLUSION: Beyond risk prediction, proteomics offers opportunities to elucidate disease mechanisms and enable drug repurposing. To realize the clinical potential of plasma proteomics, absolute or reliably recalibratable relative quantification will be required to guide patient care. Ultimately, the clinical value of proteomics will be determined by the quality rather than the quantity of protein measurements.

PMID:41539063 | DOI:10.1016/j.atherosclerosis.2026.120637

β€œI Want to Spend My Time Living”—Experiences With a Digital Outpatient Service With a Mobile App for Tailored Care Among Adults With Long-Term Health Service Needs: Qualitative Study Using Thematic Analysis

Background: Digital health services are increasingly used in hospital-based outpatient care, offering remote monitoring, patient-reported outcomes, information sharing, and asynchronous communication. While expected to improve self-management, timeliness, and efficiency, the success of digital health interventions relies on patients’ health literacy and digital health literacy. While some research has addressed potential associations between digital health interventions and patients’ health outcomes, research on patients’ experiences remains limited. Objective: The aim of this study was to explore and gain in-depth knowledge about the experiences of patients with chronic or long-term conditions enrolled in a 6-month digital outpatient care intervention for tailored care and health literacy. Methods: We conducted an exploratory qualitative interview study with 17 strategically recruited adult patients with cancer, interstitial lung disease, epilepsy, or complicated pain who used a digital outpatient service for 6 months. Individual telephone interviews were conducted using a semistructured guide, transcribed verbatim, and analyzed with thematic analysis to generate codes and themes. Participants had a median age of 62 years (minimum-maximum 36-83 years), with 8 females and 9 males. Results: The thematic analysis led to 1 main theme β€œDigital outpatient care as a flexible service supporting patients’ self-management,” informed by 3 subthemes β€œThe ongoing nature of managing a chronic condition and how the digital service meet the patients’ desire for autonomy in their care,” β€œDigital tools flexibly address the patients’ unique needs, but reliability depends on patient interaction,” and β€œDigital services enhance the patients’ sense of safety through easy access to a relation with competent healthcare workers.” The themes highlight patients’ appreciation for greater flexibility in their care and their desire to self-manage with the support of easily accessible health care workers, ultimately supporting their health literacy. Patients recognized the importance of actively engaging with the digital solution to fully benefit from its opportunities and emphasized the critical role of health care workers in fostering their sense of security. Conclusions: Digital outpatient care was experienced as flexible and supportive for patients with long-term conditions. The increased possibility of interacting with health care workers was welcomed by the patients, and the combination of flexibility, self-monitoring, and addressing concerns regarding their self-management may increase the patients experience of autonomy. As health literacy likely plays a role in patients’ ability to effectively engage with digital tools and self-manage their conditions, future research should explore how varying levels of health literacy influence these outcomes. In addition, research should address whether such digital outpatient clinics are positive for a wider range of patients, associated health outcomes, and any positive effects on a health system level. Trial Registration: ClinicalTrials.gov NCT05068869; https://clinicaltrials.gov/ct2/show/NCT05068869
  • βœ‡MIT Technology Review
  • Exclusive eBook: How AGI Became a Consequential Conspiracy Theory MIT Technology Review
    In this exclusive subscriber-only eBook, you’ll learn about how the idea that machines will be as smart asβ€”or smarter thanβ€”humans has hijacked an entire industry.by Will Douglas Heaven October 30, 2025 ACCESS EBOOK Table of Contents: How Silicon Valley got AGI-pilled The great AGI conspiracy How AGI hijacked an industry The great AGI conspiracy, concluded Related Stories: How AGI became the most consequential conspiracy theory of our time The New Conspiracy Age
     

Exclusive eBook: How AGI Became a Consequential Conspiracy Theory

16 January 2026 at 01:16

In this exclusive subscriber-only eBook, you’ll learn about how the idea that machines will be as smart asβ€”or smarter thanβ€”humans has hijacked an entire industry.

by Will Douglas Heaven October 30, 2025

Table of Contents:

  • How Silicon Valley got AGI-pilled
  • The great AGI conspiracy
  • How AGI hijacked an industry
  • The great AGI conspiracy, concluded

Related Stories:

Access all subscriber-only eBooks:

Molecular features of early- vs. late-onset gastric cancer: a systematic review and meta-analysis

BMC Cancer. 2026 Jan 14. doi: 10.1186/s12885-026-15567-5. Online ahead of print.

ABSTRACT

BACKGROUND: Early-onset gastric cancer (EOGC), diagnosed before age 50, is characterized by distinct clinicopathological features, though its molecular landscape remains poorly defined.

METHODS: A systematic literature search of PubMed, Embase, and Web of Science identified studies comparing molecular characteristics of EOGC and late-onset gastric cancer (LOGC). Meta-analyses assessed differences in The Cancer Genome Atlas (TCGA) molecular subtypes, gene mutations, therapeutic biomarkers, and serum tumor markers. Odds ratios (ORs) with 95% confidence intervals (CIs) were calculated; heterogeneity was assessed using the I2 statistic.

RESULTS: EOGC was associated with a higher prevalence of the genomically stable (GS) subtype (OR = 1.71, 95% CI: 1.37-2.12) and a lower prevalence of the chromosomal instability (CIN) subtype (OR = 0.62, 95% CI: 0.50-0.77). CDH1 mutations were more frequent in EOGC (OR = 3.44, 95% CI: 2.85-4.16), while HER2 expression (OR = 0.54, 95% CI: 0.43-0.67), dMMR/MSI-H status (OR = 0.25, 95% CI: 0.12-0.53), and p53 expression (OR = 0.56, 95% CI: 0.39-0.82) were significantly lower. Serum markers including CEA and CA19-9 were also less frequently elevated in EOGC.

CONCLUSION: EOGC represents a biologically distinct subset of gastric cancer with unique genomic and immunological features. These findings support age-specific diagnostic approaches and emphasize the value of multiomic strategies to uncover the mechanisms driving early-onset disease.

PMID:41535782 | DOI:10.1186/s12885-026-15567-5

ART: Action-based Reasoning Task Benchmarking for Medical AI Agents

arXiv:2601.08988v1 Announce Type: new Abstract: Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal aggregation, and conditional logic. We introduce ART, an Action-based Reasoning clinical Task benchmark for medical AI agents, which mines real-world EHR data to create challenging tasks targeting known reasoning weaknesses. Through analysis of existing benchmarks, we identify three dominant error categories: retrieval failures, aggregation errors, and conditional logic misjudgments. Our four-stage pipeline -- scenario identification, task generation, quality audit, and evaluation -- produces diverse, clinically validated tasks grounded in real patient data. Evaluating GPT-4o-mini and Claude 3.5 Sonnet on 600 tasks shows near-perfect retrieval after prompt refinement, but substantial gaps in aggregation (28--64%) and threshold reasoning (32--38%). By exposing failure modes in action-oriented EHR reasoning, ART advances toward more reliable clinical agents, an essential step for AI systems that reduce cognitive load and administrative burden, supporting workforce capacity in high-demand care settings

Human-AI Co-design for Clinical Prediction Models

arXiv:2601.09072v1 Announce Type: new Abstract: Developing safe, effective, and practically useful clinical prediction models (CPMs) traditionally requires iterative collaboration between clinical experts, data scientists, and informaticists. This process refines the often small but critical details of the model building process, such as which features/patients to include and how clinical categories should be defined. However, this traditional collaboration process is extremely time- and resource-intensive, resulting in only a small fraction of CPMs reaching clinical practice. This challenge intensifies when teams attempt to incorporate unstructured clinical notes, which can contain an enormous number of concepts. To address this challenge, we introduce HACHI, an iterative human-in-the-loop framework that uses AI agents to accelerate the development of fully interpretable CPMs by enabling the exploration of concepts in clinical notes. HACHI alternates between (i) an AI agent rapidly exploring and evaluating candidate concepts in clinical notes and (ii) clinical and domain experts providing feedback to improve the CPM learning process. HACHI defines concepts as simple yes-no questions that are used in linear models, allowing the clinical AI team to transparently review, refine, and validate the CPM learned in each round. In two real-world prediction tasks (acute kidney injury and traumatic brain injury), HACHI outperforms existing approaches, surfaces new clinically relevant concepts not included in commonly-used CPMs, and improves model generalizability across clinical sites and time periods. Furthermore, HACHI reveals the critical role of the clinical AI team, such as directing the AI agent to explore concepts that it had not previously considered, adjusting the granularity of concepts it considers, changing the objective function to better align with the clinical objectives, and identifying issues of data bias and leakage.

PrivacyReasoner: Can LLM Emulate a Human-like Privacy Mind?

arXiv:2601.09152v1 Announce Type: new Abstract: This paper introduces PRA, an AI-agent design for simulating how individual users form privacy concerns in response to real-world news. Moving beyond population-level sentiment analysis, PRA integrates privacy and cognitive theories to simulate user-specific privacy reasoning grounded in personal comment histories and contextual cues. The agent reconstructs each user's "privacy mind", dynamically activates relevant privacy memory through a contextual filter that emulates bounded rationality, and generates synthetic comments reflecting how that user would likely respond to new privacy scenarios. A complementary LLM-as-a-Judge evaluator, calibrated against an established privacy concern taxonomy, quantifies the faithfulness of generated reasoning. Experiments on real-world Hacker News discussions show that \PRA outperforms baseline agents in privacy concern prediction and captures transferable reasoning patterns across domains including AI, e-commerce, and healthcare.

Companion Agents: A Table-Information Mining Paradigm for Text-to-SQL

15 January 2026 at 13:00
arXiv:2601.08838v1 Announce Type: cross Abstract: Large-scale Text-to-SQL benchmarks such as BIRD typically assume complete and accurate database annotations as well as readily available external knowledge, which fails to reflect common industrial settings where annotations are missing, incomplete, or erroneous. This mismatch substantially limits the real-world applicability of state-of-the-art (SOTA) Text-to-SQL systems. To bridge this gap, we explore a database-centric approach that leverages intrinsic, fine-grained information residing in relational databases to construct missing evidence and improve Text-to-SQL accuracy under annotation-scarce conditions. Our key hypothesis is that when a query requires multi-step reasoning over extensive table information, existing methods often struggle to reliably identify and utilize the truly relevant knowledge. We therefore propose to "cache" query-relevant knowledge on the database side in advance, so that it can be selectively activated at inference time. Based on this idea, we introduce Companion Agents (CA), a new Text-to-SQL paradigm that incorporates a group of agents accompanying database schemas to proactively mine and consolidate hidden inter-table relations, value-domain distributions, statistical regularities, and latent semantic cues before query generation. Experiments on BIRD under the fully missing evidence setting show that CA recovers +4.49 / +4.37 / +14.13 execution accuracy points on RSL-SQL / CHESS / DAIL-SQL, respectively, with larger gains on the Challenging subset +9.65 / +7.58 / +16.71. These improvements stem from CA's automatic database-side mining and evidence construction, suggesting a practical path toward industrial-grade Text-to-SQL deployment without reliance on human-curated evidence.

Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents

arXiv:2601.08841v1 Announce Type: cross Abstract: The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this study, we explore how structured knowledge, specifically, subject-predicate-object triples, can enhance the clustering and classification of scientific papers. We propose a modular pipeline that combines unsupervised clustering and supervised classification over multiple document representations: raw abstracts, extracted triples, and hybrid formats that integrate both. Using a filtered arXiv corpus, we extract relational triples from abstracts and construct four text representations, which we embed using four state-of-the-art transformer models: MiniLM, MPNet, SciBERT, and SPECTER. We evaluate the resulting embeddings with KMeans, GMM, and HDBSCAN for unsupervised clustering, and fine-tune classification models for arXiv subject prediction. Our results show that full abstract text yields the most coherent clusters, but that hybrid representations incorporating triples consistently improve classification performance, reaching up to 92.6% accuracy and 0.925 macro-F1. We also find that lightweight sentence encoders (MiniLM, MPNet) outperform domain-specific models (SciBERT, SPECTER) in clustering, while SciBERT excels in structured-input classification. These findings highlight the complementary benefits of combining unstructured text with structured knowledge, offering new insights into knowledge-infused representations for semantic organization of scientific documents.

AI Deployment Authorisation: A Global Standard for Machine-Readable Governance of High-Risk Artificial Intelligence

arXiv:2601.08869v1 Announce Type: cross Abstract: Modern artificial intelligence governance lacks a formal, enforceable mechanism for determining whether a given AI system is legally permitted to operate in a specific domain and jurisdiction. Existing tools such as model cards, audits, and benchmark evaluations provide descriptive information about model behavior and training data but do not produce binding deployment decisions with legal or financial force. This paper introduces the AI Deployment Authorisation Score (ADAS), a machine-readable regulatory framework that evaluates AI systems across five legally and economically grounded dimensions: risk, alignment, externality, control, and auditability. ADAS produces a cryptographically verifiable deployment certificate that regulators, insurers, and infrastructure operators can consume as a license to operate, using public-key verification and transparency mechanisms adapted from secure software supply chain and certificate transparency systems. The paper presents the formal specification, decision logic, evidence model, and policy architecture of ADAS and demonstrates how it operationalizes the European Union Artificial Intelligence Act, United States critical infrastructure governance, and insurance underwriting requirements by compiling statutory and regulatory obligations into machine-executable deployment gates. We argue that deployment-level authorization, rather than model-level evaluation, constitutes the missing institutional layer required for safe, lawful, and economically scalable artificial intelligence.

From Symbolic to Natural-Language Relations: Rethinking Knowledge Graph Construction in the Era of Large Language Models

arXiv:2601.09069v1 Announce Type: cross Abstract: Knowledge graphs (KGs) have commonly been constructed using predefined symbolic relation schemas, typically implemented as categorical relation labels. This design has notable shortcomings: real-world relations are often contextual, nuanced, and sometimes uncertain, and compressing it into discrete relation labels abstracts away critical semantic detail. Nevertheless, symbolic-relation KGs remain widely used because they have been operationally effective and broadly compatible with pre-LLM downstream models and algorithms, in which KG knowledge could be retrieved or encoded into quantified features and embeddings at scale. The emergence of LLMs has reshaped how knowledge is created and consumed. LLMs support scalable synthesis of domain facts directly in concise natural language, and prompting-based inference favors context-rich free-form text over quantified representations. This position paper argues that these changes call for rethinking the representation of relations themselves rather than merely using LLMs to populate conventional schemas more efficiently. We therefore advocate moving from symbolic to natural-language relation descriptions, and we propose hybrid design principles that preserve a minimal structural backbone while enabling more flexible and context-sensitive relational representations.
❌