Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems
arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed bec
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards Execution-Grounded Automated AI Research
arXiv:2601.14525v1 Announce Type: cross Abstract: Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding may help, but it is unclear whether automated execution is feasible and whether LLMs can learn from the execution feedback. To investigate these, we first build an automated executor to implement ideas and launch large-scale parallel GPU experiments to verify their effec
Towards Execution-Grounded Automated AI Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
arXiv:2601.15161v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are particularly challenging as they often manifest as subtle clinical errors that evade detection by generic metrics, while expert-authored fine-grained rubrics remain costly to construct and difficult to scale. In this paper, we propose a retrieval-augmented multi-age
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
arXiv:2307.13658v3 Announce Type: replace-cross Abstract: We propose that future AI transparency and accountability regulations are based on an open global standard for exchanging information about AI systems, which allows co-existence of potentially conflicting local regulations. Then, we discuss key components of a lightweight and effective AI transparency and/or accountability regulation. To prevent overregulation, the proposed approach encourages collaboration between regulators and industr
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
arXiv:2509.19774v2 Announce Type: replace-cross Abstract: Electrocardiography (ECG) is the clinical gold standard for cardiovascular disease (CVD) assessment, yet continuous monitoring is constrained by the need for dedicated hardware and trained personnel. Photoplethysmography (PPG) is ubiquitous in wearable devices and readily scalable, but it lacks electrophysiological specificity, limiting diagnostic reliability. While generative methods aim to translate PPG into clinically useful ECG signa
PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection
-
npj Digital Medicine
-
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5Large language models improve transferability of electronic health record-based predictions across countries and coding systems
Large language models improve transferability of electronic health record-based predictions across countries and coding systems
npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5
Large language models improve transferability of electronic health record-based predictions across countries and coding systems-
Cell
-
Multimodal AI generates virtual population for tumor microenvironment modeling
GigaTIME leverages multimodal AI to generate virtual multiplex immunofluorescence (mIF) profiles from standard H&E slides, enabling comprehensive tumor immune microenvironment modeling across a large (>14,000) and diverse patient population. This virtual approach unlocks new opportunities for large-scale clinical discoveries that were previously hindered by the scarcity of mIF data.
Multimodal AI generates virtual population for tumor microenvironment modeling
-
(Multiomics OR Omics) AND (Pancreatic)
-
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.ABSTRACTPancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approxima
Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability
Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.
ABSTRACT
Pancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approximately 1-2% of cases. These tumors display unique molecular characteristics, including a high prevalence of wild-type KRAS and TP53, elevated tumor mutational burden, and recurrent kinase fusions, which together confer enhanced immunogenicity and increased sensitivity to immune checkpoint inhibitors (ICIs). In addition to their therapeutic relevance, dMMR/MSI-H status has important diagnostic implications for the identification of Lynch syndrome-associated pancreatic cancers, informing genetic counseling and familial risk assessment. This review summarizes current understanding of the molecular basis of mismatch repair deficiency and microsatellite instability in PC, evaluates available diagnostic approaches such as immunohistochemistry, polymerase chain reaction, and next-generation sequencing, and discusses the prognostic and predictive significance of dMMR/MSI-H status. Emerging clinical evidence supporting the use of ICIs in selected patients across neoadjuvant, adjuvant, and advanced disease settings is also reviewed, along with challenges related to assay discordance, tumor heterogeneity, and immunotherapy resistance. Finally, future directions are highlighted, emphasizing the need for standardized testing algorithms, integration of multi-omics and spatial profiling technologies, and prospective clinical studies to optimize precision treatment strategies for this rare but clinically meaningful subtype of pancreatic cancer.
PMID:41563663 | DOI:10.1007/s12094-025-04214-3
-
Omics in Hepatocellular
-
United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes
iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.ABSTRACTHepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes
United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes
iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.
ABSTRACT
Hepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes predictive of prognosis and immune response. We identified two cancer subtypes (CSs) that are associated with prognosis, with the second subtype (CS2) showing the most favorable prognostic outcomes. Subsequently, 9 key genes were screened for HCC model scoring, stratifying patients into low-risk (good prognosis, responsive to immunotherapy) and high-risk (poor outcome, not responsive to immunotherapy) groups. The high-risk group may be effective against the mTOR inhibitor AZD8055. Comprehensive multi-omics data and multiple ML algorithms offer key insights into HCC occurrence and evolution, with model scores guiding patient prognosis and treatment clinically.
PMID:41561382 | PMC:PMC12814435 | DOI:10.1016/j.isci.2025.114328
-
cs.AI, q-bio.NC updates on arXiv.org
-
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
arXiv:2601.13122v1 Announce Type: new Abstract: Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vuln
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
-
cs.AI, q-bio.NC updates on arXiv.org
-
DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research
arXiv:2601.11560v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) encode vast, heterogeneous information spanning literature, genes, pathways, drugs, diseases, and clinical trials, but leveraging them collectively for scientific discovery remains difficult. Their structural differences, continual evolution, and limited cross-resource alignment require substantial manual integration, limiting the depth and scale of knowledge exploration. We introduce DeepEvidence, an AI-agent f
DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research
-
cs.AI, q-bio.NC updates on arXiv.org
-
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
arXiv:2601.12471v1 Announce Type: cross Abstract: Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integra
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Cloud-based Multi-Agentic Workflow for Science
arXiv:2601.12607v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex decisions limits their utility. LLM-based agents bridge this gap due to their ability to call external resources and tools and thus are now rapidly gaining popularity. However, coming up with a workflow that can balance the models, cloud providers, and external resources i
A Cloud-based Multi-Agentic Workflow for Science
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
arXiv:2601.12805v1 Announce Type: cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap
SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
arXiv:2601.12910v1 Announce Type: cross Abstract: We present SciCoQA, a dataset for detecting discrepancies between scientific publications and their codebases to ensure faithful implementations. We construct SciCoQA from GitHub issues and reproducibility papers, and to scale our dataset, we propose a synthetic data generation method for constructing paper-code discrepancies. We analyze the paper-code discrepancies in detail and propose discrepancy types and categories to better understand the
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI-generated data contamination erodes pathological variability and diagnostic reliability
arXiv:2601.12946v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological varia
AI-generated data contamination erodes pathological variability and diagnostic reliability
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multi-objective fluorescent molecule design with a data-physics dual-driven generative framework
arXiv:2601.13564v1 Announce Type: cross Abstract: Designing fluorescent small molecules with tailored optical and physicochemical properties requires navigating vast, underexplored chemical space while satisfying multiple objectives and constraints. Conventional generate-score-screen approaches become impractical under such realistic design specifications, owing to their low search efficiency, unreliable generalizability of machine-learning prediction, and the prohibitive cost of quantum chemic
Multi-objective fluorescent molecule design with a data-physics dual-driven generative framework
-
cs.AI, q-bio.NC updates on arXiv.org
-
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
arXiv:2601.13659v1 Announce Type: cross Abstract: Multimodal Sentiment Analysis integrates Linguistic, Visual, and Acoustic. Mainstream approaches based on modality-invariant and modality-specific factorization or on complex fusion still rely on spatiotemporal mixed modeling. This ignores spatiotemporal heterogeneity, leading to spatiotemporal information asymmetry and thus limited performance. Hence, we propose TSDA, Temporal-Spatial Decouple before Act, which explicitly decouples each modalit
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
Who Should Have Surgery? A Comparative Study of GenAI vs Supervised ML for CRS Surgical Outcome Prediction
arXiv:2601.13710v1 Announce Type: cross Abstract: Artificial intelligence has reshaped medical imaging, yet the use of AI on clinical data for prospective decision support remains limited. We study pre-operative prediction of clinically meaningful improvement in chronic rhinosinusitis (CRS), defining success as a more than 8.9-point reduction in SNOT-22 at 6 months (MCID). In a prospectively collected cohort where all patients underwent surgery, we ask whether models using only pre-operative cl
Who Should Have Surgery? A Comparative Study of GenAI vs Supervised ML for CRS Surgical Outcome Prediction
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Survey of AI Scientists
arXiv:2510.23045v5 Announce Type: replace Abstract: Artificial intelligence is undergoing a profound transition from a computational instrument to an autonomous originator of scientific knowledge. This emerging paradigm, the AI scientist, is architected to emulate the complete scientific workflow-from initial hypothesis generation to the final synthesis of publishable findings-thereby promising to fundamentally reshape the pace and scale of discovery. However, the rapid and unstructured prolife