❌

Normal view

AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education: Observational Crossover Cohort Study

Background: Virtual patient (VP) simulations can be used to practice clinical reasoning (CR) in controlled learning environments. Traditional computer-based VP platforms often lack the authenticity and interactivity required for effective CR training. Artificial intelligence (AI)–enhanced social robotic VPs can enhance realism and engagement; however, quantitative evidence comparing them with conventional VP platforms remains limited. Objective: We compared medical students’ experience of an AI-enhanced social robotic versus a conventional computer-based VP platform regarding the extent to which the design characteristics of the respective platform facilitate CR skill training. Methods: This observational crossover cohort study involved 178 sixth-semester medical students at Karolinska Institutet, Stockholm, Sweden (response rate: 42.3%; 178 of 421 invited students; Spring 2024-Spring 2025), who experienced both a large language model–enhanced social robotic VP platform supporting dialogue (social artificial intelligence–enhanced robotic interface [SARI]) and a conventional computer-based VP platform (virtual interactive case [VIC]) during their clinical rotation within rheumatology. Platform order was determined by clinical rotation scheduling. VP design was evaluated using a validated questionnaire across 5 domains: authenticity, professional approach, coaching quality, learning effects, and overall judgment. Students’ CR training preferences were assessed using categorical responses and a Visual Analogue Scale, where a lower score favored SARI and a score of 5 indicated equal preference between platforms. Results: SARI outperformed VIC across all 5 VP design domains. Students rated SARI higher for authenticity (median 4.0, IQR 3.5-4.5 vs 3.0, IQR 2.5-3.5; P<.001 professional approach iqr vs>P<.001 coaching quality iqr vs>P<.001 learning effect iqr vs>P<.001 and overall judgment vs iqr>P<.001 students strongly preferred sari for cr training vs odds ratio ci>P<.001 with visual analogue scale scores confirming this preference iqr>P<.001 preferences were consistent across most subgroups prior vp experience and platform order in the difference was not significant that is students with vs or ci>P=.11) and students first introduced to VIC (55% vs 45%; OR 1.5; 95% CI 0.7-2.9; P=.33). Conclusions: Our findings provide the first quantitative evidence that AI-enhanced social robotic VPs offer superior design characteristics than conventional computer-based platforms for CR training in medical education. These results support the use of AI-driven social robots for VP simulations to better prepare medical students for real clinical encounters, and warrant future research on objective CR skill outcomes and long-term transfer to clinical practice. Unlike previous qualitative studies examining each platform separately, this study provides the first quantitative comparison of design characteristics between AI-enhanced social robotic and conventional computer-based VPs.

Harnessing Single-Cell RNA-Seq for Computational Drug Repurposing in Cancer Immunotherapy

27 November 2025 at 19:00

Pharmaceuticals (Basel). 2025 Nov 20;18(11):1769. doi: 10.3390/ph18111769.

ABSTRACT

Immune checkpoint inhibitors (ICIs) have revolutionized cancer treatment and show notable success in some cancer types such as non-small cell lung cancer, melanoma and colorectal cancers, while they demonstrate relatively low response rate in others, such as esophageal cancers. Due to the heterogeneous nature of the tumor microenvironment and patient-to-patient variability, there remains a need to improve ICI response rates. Combining ICIs with therapies that can overcome resistance is a promising strategy. Compared to de novo drug development, drug repurposing offers a faster and more cost-effective approach to identifying such combination candidates. A variety of computational drug repurposing tools leverage genomics and/or transcriptomic data. As single-cell RNA sequencing (scRNA-seq) technology becomes available, it enables precise targeting of cancer-driving cellular components. In this review, we highlight current computational drug repurposing tools utilizing scRNA-seq data and demonstrate the application of two such tools, scDrug and scDrugPrio, on an esophageal squamous cell carcinoma dataset to identify potential drug candidates for combination with ICI therapy to enhance treatment response. scDrug focuses on predicting tumor cell-specific cytotoxicity, while scDrugPrio prioritizes drugs by reversing gene signatures associated with ICI non-responsiveness across diverse tumor microenvironment cell types. Together, this review underscores the importance of a multi-faceted approach in computational drug repurposing and highlights its potential for identifying drugs that enhance ICI treatment. Future work can expand the application of these strategies to multi-omics and spatial transcriptomics datasets, as well as personalized patient samples, to further refine drug repurposing involving ICI therapy.

PMID:41305010 | PMC:PMC12655618 | DOI:10.3390/ph18111769

Harnessing Single-Cell RNA-Seq for Computational Drug Repurposing in Cancer Immunotherapy

Pharmaceuticals (Basel). 2025 Nov 20;18(11):1769. doi: 10.3390/ph18111769.

ABSTRACT

Immune checkpoint inhibitors (ICIs) have revolutionized cancer treatment and show notable success in some cancer types such as non-small cell lung cancer, melanoma and colorectal cancers, while they demonstrate relatively low response rate in others, such as esophageal cancers. Due to the heterogeneous nature of the tumor microenvironment and patient-to-patient variability, there remains a need to improve ICI response rates. Combining ICIs with therapies that can overcome resistance is a promising strategy. Compared to de novo drug development, drug repurposing offers a faster and more cost-effective approach to identifying such combination candidates. A variety of computational drug repurposing tools leverage genomics and/or transcriptomic data. As single-cell RNA sequencing (scRNA-seq) technology becomes available, it enables precise targeting of cancer-driving cellular components. In this review, we highlight current computational drug repurposing tools utilizing scRNA-seq data and demonstrate the application of two such tools, scDrug and scDrugPrio, on an esophageal squamous cell carcinoma dataset to identify potential drug candidates for combination with ICI therapy to enhance treatment response. scDrug focuses on predicting tumor cell-specific cytotoxicity, while scDrugPrio prioritizes drugs by reversing gene signatures associated with ICI non-responsiveness across diverse tumor microenvironment cell types. Together, this review underscores the importance of a multi-faceted approach in computational drug repurposing and highlights its potential for identifying drugs that enhance ICI treatment. Future work can expand the application of these strategies to multi-omics and spatial transcriptomics datasets, as well as personalized patient samples, to further refine drug repurposing involving ICI therapy.

PMID:41305010 | DOI:10.3390/ph18111769

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

arXiv:2506.14652v2 Announce Type: replace-cross Abstract: In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception -- in addition to a more expansive understanding of (1) methodological rigor -- should include aspects related to (2) what background knowledge informs what to work on (epistemic rigor); (3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); (4) how clearly articulated the theoretical constructs under use are (conceptual rigor); (5) what is reported and how (reporting rigor); and (6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also provide useful language and a framework for much-needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders.

Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues

Nat Cell Biol. 2025 Nov 26. doi: 10.1038/s41556-025-01811-w. Online ahead of print.

ABSTRACT

Spatial omics technologies have transformed biomedical research by enabling high-resolution molecular profiling while preserving the native tissue architecture. These advances provide unprecedented insights into tissue structure and function. However, the high cost and time-intensive nature of spatial omics experiments necessitate careful experimental design, particularly in selecting regions of interest (ROIs) from large tissue sections. Currently, ROI selection is performed manually, which introduces subjectivity, inconsistency and a lack of reproducibility. Previous studies have shown strong correlations between spatial molecular patterns and histological features, suggesting that readily available and cost-effective histology images can be leveraged to guide spatial omics experiments. Here we present Smart Spatial omics (S2-omics), an end-to-end workflow that automatically selects ROIs from histology images with the goal of maximizing molecular information content in the ROIs. Through comprehensive evaluations across multiple spatial omics platforms and tissue types, we demonstrate that S2-omics enables systematic and reproducible ROI selection and enhances the robustness and impact of downstream biological discovery.

PMID:41298871 | DOI:10.1038/s41556-025-01811-w

Human Experts' Evaluation of Generative AI for Contextualizing STEAM Education in the Global South

arXiv:2511.19482v2 Announce Type: replace-cross Abstract: This study investigates how human experts evaluate the capacity of Generative AI (GenAI) to contextualize STEAM education in the Global South, with a focus on Ghana. Using a convergent mixed-methods design, four STEAM specialists assessed GenAI-generated lesson plans created with a customized Culturally Responsive Lesson Planner (CRLP) and compared them to standardized lesson plans from the Ghana National Council for Curriculum and Assessment (NaCCA). Quantitative ratings were based on a validated 25-item Culturally Responsive Pedagogy Rubric measuring bias awareness, cultural representation, contextual relevance, linguistic responsiveness, and teacher agency. Qualitative reflections provided additional insight into how GenAI handles cultural and pedagogical appropriateness. Findings show that GenAI, when paired with the CRLP tool, can support contextualized STEAM instruction by linking abstract curriculum standards to learners' cultural knowledge, community practices, and everyday experiences. Experts rated GenAI-assisted lessons as more culturally grounded and pedagogically responsive than NaCCA plans, integrating Indigenous knowledge, bilingual elements, and locally relevant examples. However, GenAI struggled to represent Ghana's cultural pluralism, often offering surface-level references to language, history, and identity. These weaknesses were most evident in Mathematics and Computing, where cultural nuance was limited. The results highlight the need for continued teacher mediation, community involvement, and culturally attuned refinement of AI outputs. Future work should include classroom trials, expanded expert participation, and model fine-tuning using Indigenous language corpora to strengthen cultural fidelity in Global South contexts.

Precision Oncology: Current Landscape, Emerging Trends, Challenges, and Future Perspectives

26 November 2025 at 19:00

Cells. 2025 Nov 17;14(22):1804. doi: 10.3390/cells14221804.

ABSTRACT

Precision oncology is broadly defined as cancer prevention, diagnosis, and treatment specifically tailored to the patient based on his/her genetics and molecular profile. In simple terms, the goal of precision medicine is to deliver the right cancer treatment to the right patient, at the right dose, at the right time. Precision oncology is the most studied and widely applied subarea of precision medicine. Now, precision oncology has expanded to include modern technology (big data, single-cell spatial multiomics, molecular imaging, liquid biopsy, CRISPR gene editing, stem cells, organoids), a deeper understanding of cancer biology (driver cancer genes, single nucleotide polymorphism, cancer initiation, intratumor heterogeneity, tumor microenvironment ecosystem, pan-cancer), cancer stratification (subtyping of traditionally defined cancer types and pan-cancer re-classification based on shared properties across traditionally defined cancer types), clinical applications (cancer prevention, early detection, diagnosis, targeted therapy, minimal residual disease monitoring, managing drug resistance), lifestyle changes (physical activity, smoking, alcohol consumption, sunscreen), cost management, public policy, and more. Despite being the most developed area in precision medicine, precision oncology is still in its early stages and faces multiple challenges that need to be overcome for its successful implementation. In this review, we examine the history, development, and future directions of precision oncology by focusing on emerging technology, novel concepts and principles, molecular cancer stratification, and clinical applications.

PMID:41294857 | PMC:PMC12651332 | DOI:10.3390/cells14221804

  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Hybrid Neuro-Symbolic Models for Ethical AI in Risk-Sensitive Domains Chaitanya Kumar Kolli
    arXiv:2511.17644v1 Announce Type: new Abstract: Artificial intelligence deployed in risk-sensitive domains such as healthcare, finance, and security must not only achieve predictive accuracy but also ensure transparency, ethical alignment, and compliance with regulatory expectations. Hybrid neuro symbolic models combine the pattern-recognition strengths of neural networks with the interpretability and logical rigor of symbolic reasoning, making them well-suited for these contexts. This paper su
     

Hybrid Neuro-Symbolic Models for Ethical AI in Risk-Sensitive Domains

arXiv:2511.17644v1 Announce Type: new Abstract: Artificial intelligence deployed in risk-sensitive domains such as healthcare, finance, and security must not only achieve predictive accuracy but also ensure transparency, ethical alignment, and compliance with regulatory expectations. Hybrid neuro symbolic models combine the pattern-recognition strengths of neural networks with the interpretability and logical rigor of symbolic reasoning, making them well-suited for these contexts. This paper surveys hybrid architectures, ethical design considerations, and deployment patterns that balance accuracy with accountability. We highlight techniques for integrating knowledge graphs with deep inference, embedding fairness-aware rules, and generating human-readable explanations. Through case studies in healthcare decision support, financial risk management, and autonomous infrastructure, we show how hybrid systems can deliver reliable and auditable AI. Finally, we outline evaluation protocols and future directions for scaling neuro symbolic frameworks in complex, high stakes environments.

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

arXiv:2511.18298v1 Announce Type: new Abstract: The exponential growth of scientific knowledge has created significant barriers to cross-disciplinary knowledge discovery, synthesis and research collaboration. In response to this challenge, we present BioSage, a novel compound AI architecture that integrates LLMs with RAG, orchestrated specialized agents and tools to enable discoveries across AI, data science, biomedical, and biosecurity domains. Our system features several specialized agents including the retrieval agent with query planning and response synthesis that enable knowledge retrieval across domains with citation-backed responses, cross-disciplinary translation agents that align specialized terminology and methodologies, and reasoning agents that synthesize domain-specific insights with transparency, traceability and usability. We demonstrate the effectiveness of our BioSage system through a rigorous evaluation on scientific benchmarks (LitQA2, GPQA, WMDP, HLE-Bio) and introduce a new cross-modal benchmark for biology and AI, showing that our BioSage agents outperform vanilla and RAG approaches by 13\%-21\% powered by Llama 3.1. 70B and GPT-4o models. We perform causal investigations into compound AI system behavior and report significant performance improvements by adding RAG and agents over the vanilla models. Unlike other systems, our solution is driven by user-centric design principles and orchestrates specialized user-agent interaction workflows supporting scientific activities including but not limited to summarization, research debate and brainstorming. Our ongoing work focuses on multimodal retrieval and reasoning over charts, tables, and structured scientific data, along with developing comprehensive multimodal benchmarks for cross-disciplinary discovery. Our compound AI solution demonstrates significant potential for accelerating scientific advancement by reducing barriers between traditionally siloed domains.

Predicting Healthcare Provider Engagement in SMS Campaigns

arXiv:2511.17658v1 Announce Type: cross Abstract: As digital communication grows in importance when connecting with healthcare providers, traditional behavioral and content message features are imbued with renewed significance. If one is to meaningfully connect with them, it is crucial to understand what drives them to engage and respond. In this study, the authors analyzed several million text messages sent through the Impiricus platform to learn which factors influenced whether or not a doctor clicked on a link in a message. Several key insights came to light through the use of logistic regression, random forest, and neural network models, the details of which the authors discuss in this paper.

Clinician-Directed Large Language Model Software Generation for Therapeutic Interventions in Physical Rehabilitation

arXiv:2511.18274v1 Announce Type: cross Abstract: Digital health interventions are increasingly used in physical and occupational therapy to deliver home exercise programs via sensor equipped devices such as smartphones, enabling remote monitoring of adherence and performance. However, digital interventions are typically programmed as software before clinical encounters as libraries of parametrized exercise modules targeting broad patient populations. At the point of care, clinicians can only select modules and adjust a narrow set of parameters like repetitions, so patient specific needs that emerge during encounters, such as distinct movement limitations, and home environments, are rarely reflected in the software. We evaluated a digital intervention paradigm that uses large language models (LLMs) to translate clinicians' exercise prescriptions into intervention software. In a prospective single arm feasibility study with 20 licensed physical and occupational therapists and a standardized patient, clinicians created 40 individualized upper extremity programs (398 instructions) that were automatically translated into executable software. Our results show a 45% increase in the proportion of personalized prescriptions that can be implemented as software compared with a template based benchmark, with unanimous consensus among therapists on ease of use. The LLM generated software correctly delivered 99.78% (397/398) of instructions as prescribed and monitored performance with 88.4% (352/398) accuracy, with 90% (18/20) of therapists judged it safe to interact with patients, and 75% (15/20) expressed willingness to adopt it. To our knowledge, this is the first prospective evaluation of clinician directed intervention software generation with LLMs in healthcare, demonstrating feasibility and motivating larger trials to assess clinical effectiveness and safety in real patient populations.

Clinician-in-the-Loop Smart Home System to Detect Urinary Tract Infection Flare-Ups via Uncertainty-Aware Decision Support

arXiv:2511.18334v1 Announce Type: cross Abstract: Urinary tract infection (UTI) flare-ups pose a significant health risk for older adults with chronic conditions. These infections often go unnoticed until they become severe, making early detection through innovative smart home technologies crucial. Traditional machine learning (ML) approaches relying on simple binary classification for UTI detection offer limited utility to nurses and practitioners as they lack insight into prediction uncertainty, hindering informed clinical decision-making. This paper presents a clinician-in-the-loop (CIL) smart home system that leverages ambient sensor data to extract meaningful behavioral markers, train robust predictive ML models, and calibrate them to enable uncertainty-aware decision support. The system incorporates a statistically valid uncertainty quantification method called Conformal-Calibrated Interval (CCI), which quantifies uncertainty and abstains from making predictions ("I don't know") when the ML model's confidence is low. Evaluated on real-world data from eight smart homes, our method outperforms baseline methods in recall and other classification metrics while maintaining the lowest abstention proportion and interval width. A survey of 42 nurses confirms that our system's outputs are valuable for guiding clinical decision-making, underscoring their practical utility in improving informed decisions and effectively managing UTIs and other condition flare-ups in older adults.

OpenGloss: A Synthetic Encyclopedic Dictionary and Semantic Knowledge Graph

arXiv:2511.18622v1 Announce Type: cross Abstract: We present OpenGloss, a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource. OpenGloss contains 537K senses across 150K lexemes, on par with WordNet 3.1 and Open English WordNet, while providing more than four times as many sense definitions. These lexemes include 9.1M semantic edges, 1M usage examples, 3M collocations, and 60M words of encyclopedic content. Generated through a multi-agent procedural generation pipeline with schema-validated LLM outputs and automated quality assurance, the entire resource was produced in under one week for under $1,000. This demonstrates that structured generation can create comprehensive lexical resources at cost and time scales impractical for manual curation, enabling rapid iteration as foundation models improve. The resource addresses gaps in pedagogical applications by providing integrated content -- definitions, examples, collocations, encyclopedias, etymology -- that supports both vocabulary learning and natural language processing tasks. As a synthetically generated resource, OpenGloss reflects both the capabilities and limitations of current foundation models. The dataset is publicly available on Hugging Face under CC-BY 4.0, enabling researchers and educators to build upon and adapt this resource.

No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases

arXiv:2511.18635v1 Announce Type: cross Abstract: Large Language Models (LLMs) inherit societal biases from their training data, potentially leading to harmful or unfair outputs. While various techniques aim to mitigate these biases, their effects are often evaluated only along the dimension of the bias being targeted. This work investigates the cross-category consequences of targeted bias mitigation. We study four bias mitigation techniques applied across ten models from seven model families, and we explore racial, religious, profession- and gender-related biases. We measure the impact of debiasing on model coherence and stereotypical preference using the StereoSet benchmark. Our results consistently show that while targeted mitigation can sometimes reduce bias in the intended dimension, it frequently leads to unintended and often negative consequences in others, such as increasing model bias and decreasing general coherence. These findings underscore the critical need for robust, multi-dimensional evaluation tools when examining and developing bias mitigation strategies to avoid inadvertently shifting or worsening bias along untargeted axes.

Health system learning achieves generalist neuroimaging models

arXiv:2511.18640v1 Announce Type: cross Abstract: Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroimaging, in particular, is underrepresented in the public domain due to identifiable facial features within MRI and CT scans, fundamentally restricting model performance in clinical medicine. Here, we show that frontier models underperform on neuroimaging tasks and that learning directly from uncurated data generated during routine clinical care at health systems, a paradigm we call health system learning, yields high-performance, generalist neuroimaging models. We introduce NeuroVFM, a visual foundation model trained on 5.24 million clinical MRI and CT volumes using a scalable volumetric joint-embedding predictive architecture. NeuroVFM learns comprehensive representations of brain anatomy and pathology, achieving state-of-the-art performance across multiple clinical tasks, including radiologic diagnosis and report generation. The model exhibits emergent neuroanatomic understanding and interpretable visual grounding of diagnostic findings. When paired with open-source language models through lightweight visual instruction tuning, NeuroVFM generates radiology reports that surpass frontier models in accuracy, clinical triage, and expert preference. Through clinically grounded visual understanding, NeuroVFM reduces hallucinated findings and critical errors, offering safer clinical decision support. These results establish health system learning as a paradigm for building generalist medical AI and provide a scalable framework for clinical foundation models.

Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering

arXiv:2511.19220v1 Announce Type: cross Abstract: Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs demonstrate genuine visual grounding when answering Italian medical questions by testing four state-of-the-art models: Claude Sonnet 4.5, GPT-4o, GPT-5-mini, and Gemini 2.0 flash exp. Using 60 questions from the EuropeMedQA Italian dataset that explicitly require image interpretation, we substitute correct medical images with blank placeholders to test whether models truly integrate visual and textual information. Our results reveal striking variability in visual dependency: GPT-4o shows the strongest visual grounding with a 27.9pp accuracy drop (83.2% [74.6%, 91.7%] to 55.3% [44.1%, 66.6%]), while GPT-5-mini, Gemini, and Claude maintain high accuracy with modest drops of 8.5pp, 2.4pp, and 5.6pp respectively. Analysis of model-generated reasoning reveals confident explanations for fabricated visual interpretations across all models, suggesting varying degrees of reliance on textual shortcuts versus genuine visual analysis. These findings highlight critical differences in model robustness and the need for rigorous evaluation before clinical deployment.

Large Language Model-based Data Science Agent: A Survey

arXiv:2508.02744v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) has driven novel applications across diverse domains, with LLM-based agents emerging as a crucial area of exploration. This survey presents a comprehensive analysis of LLM-based agents designed for data science tasks, summarizing insights from recent studies. From the agent perspective, we discuss the key design principles, covering agent roles, execution, knowledge, and reflection methods. From the data science perspective, we identify key processes for LLM-based agents, including data preprocessing, model development, evaluation, visualization, etc. Our work offers two key contributions: (1) a comprehensive review of recent developments in applying LLMbased agents to data science tasks; (2) a dual-perspective framework that connects general agent design principles with the practical workflows in data science.
❌