Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI & Human Co-Improvement for Safer Co-Superintelligence
arXiv:2512.05356v1 Announce Type: new Abstract: Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration between human researchers and AIs to achieve co-superintelligence. That is, specifically targeting improving AI systems' ability to work with human researchers to conduct AI research together, from ideation to exper
-
cs.AI, q-bio.NC updates on arXiv.org
-
MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
arXiv:2512.05365v1 Announce Type: new Abstract: Healthcare AI systems have historically faced challenges in merging contextual reasoning, long-term state management, and human-verifiable workflows into a cohesive framework. This paper introduces a completely innovative architecture and concept: combining the Model Context Protocol (MCP) with a specific clinical application, known as MCP-AI. This integration allows intelligent agents to reason over extended periods, collaborate securely, and adh
MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems
arXiv:2512.05449v1 Announce Type: new Abstract: Large language models display a peculiar form of inconsistency: they "know" the correct answer but fail to act on it. In human philosophy, this tension between global judgment and local impulse is called akrasia, or weakness of will. We propose akrasia as a foundational concept for analyzing inconsistency and goal drift in agentic AI systems. To operationalize it, we introduce a preliminary version of the Akrasia Benchmark, currently a structured
The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems
-
AI News

-
UK doctors’ surgeries deploying AI in patient care
InTouchNow.ai is now offering doctors surgeries a piece of software designed to modernise phone answering, designed to reduce hold times and create a smoother, more responsive experience for patients and staff. In the UK, many GP (general practice) surgeries’ phone lines are tied up in the mornings as patients try to contact their medical practitioner for appointments. More acute need can be delayed among calls with routine enquiries, meaning high-priority callers can be left waiting for long pe
UK doctors’ surgeries deploying AI in patient care
InTouchNow.ai is now offering doctors surgeries a piece of software designed to modernise phone answering, designed to reduce hold times and create a smoother, more responsive experience for patients and staff. In the UK, many GP (general practice) surgeries’ phone lines are tied up in the mornings as patients try to contact their medical practitioner for appointments. More acute need can be delayed among calls with routine enquiries, meaning high-priority callers can be left waiting for long periods.
The system uses voice-based AI to handle calls, schedule appointments, and assess patient needs, and is capable of handling many calls simultaneously, channelling callers with appointment requests, those seeking general advice, prescription requests, and seeking results of clinical tests.
Founded by Daniel Park, InTouchNow.ai draws on his 30+ years of experience in medical call centres. The AI receptionist answers calls quickly, and can automatically update integrated appointment systems. Practices can record a voice messages to personalise the experience for patients.
Benefits for practices include reducing the numbers of missed calls, decreased workload for reception staff, and improved access by patients to medical services. Being entirely software-based, the system operates outside regular hours, which can reduce the need for staff overtime at times of peak demand.
The system integrates with common GP software like Surgery Connect, AWS, and Anima, automating tasks and maintaining patient data security while giving practices full control.
The technology supports over 200 languages, with options for different dialects and accents, an aspect that will help patients in multi-cultural areas like inner-cities. Several practices in the UK are already using InTouchNow.ai and have reported positive results in call handling and patient access.
The much under-funded National Health Service in the UK has been quick to deploy AI-powered software to reduce its operating costs, often targeting the reduction of staff administration costs to funnel funds into patient care. For example, Smart Triage is an AI-powered system deployed in UK GP practices that can triage patients making initial enquiries, and based on their responses, book them into the right care pathway, such as GP or nurse appointment, or referral to specialist clinician.
An evaluation of Smart Triage at a Surrey GP practice in 2024 showed the platform reduced the average patient waiting time by 73%.
For clinicians, especially GPs, iatroX is a UK-based AI clinical reference platform that helps doctors retrieve evidence-based clinical guidance, and summarising relevant literature & guidelines. Doctors in general practice are expected to be able to assess a full range of patients’ needs, and such platforms help clinicians identify the cause of uncommon symptoms when GPs might lack specialist knowledge.
An evaluation in 2025 found a majority of surveyed users stating iatroX was useful (~86%) or reliable (~79%).
As documented by NHS England, AI platforms are used in practice, tackling tasks like diagnosis, the monitoring of chronic disease, provision of prescription advice, and handling general administration tasks that otherwise would take up clinicians’ time. Of all the sectors where sensitive data has to be protected, medicine has one of the highest standards of governance, making the deployment of AI a delicate balance between operational effectiveness and the preservation of privacy.
(Image source: “Doctor appointment” by Taric25 is licensed under CC BY 2.0.)

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The post UK doctors’ surgeries deploying AI in patient care appeared first on AI News.
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
arXiv:2512.05765v1 Announce Type: new Abstract: Influential critiques argue that Large Language Models (LLMs) are a dead end for AGI: "mere pattern matchers" structurally incapable of reasoning or planning. We argue this conclusion misidentifies the bottleneck: it confuses the ocean with the net. Pattern repositories are the necessary System-1 substrate; the missing component is a System-2 coordination layer that selects, constrains, and binds these patterns. We formalize this layer via UCCT, a
The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
-
cs.AI, q-bio.NC updates on arXiv.org
-
XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots
arXiv:2512.05270v1 Announce Type: cross Abstract: As mobile robots increasingly operate alongside humans in shared workspaces, ensuring safe, efficient, and interpretable Human-Robot Interaction (HRI) has become a pressing challenge. While substantial progress has been devoted to human behavior prediction, limited attention has been paid to how humans perceive, interpret, and trust robots' inferences, impeding deployment in safety-critical and socially embedded environments. This paper presents
XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots
-
cs.AI, q-bio.NC updates on arXiv.org
-
Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
arXiv:2512.05397v1 Announce Type: cross Abstract: Major life transitions demand high-stakes decisions, yet people often struggle to imagine how their future selves will live with the consequences. To support this limited capacity for mental time travel, we introduce AI-enabled digital twins that have ``lived through'' simulated life scenarios. Rather than predicting optimal outcomes, these simulations extend prospective cognition by making alternative futures vivid enough to support deliberatio
Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
-
cs.AI, q-bio.NC updates on arXiv.org
-
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
arXiv:2512.05863v1 Announce Type: cross Abstract: Medical question-answering (QA) systems can benefit from advances in large language models (LLMs), but directly applying LLMs to the clinical domain poses challenges such as maintaining factual accuracy and avoiding hallucinations. In this paper, we present a retrieval-augmented generation (RAG) based medical QA system that combines domain-specific knowledge retrieval with open-source LLMs to answer medical questions. We fine-tune two state-of-t
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
-
cs.AI, q-bio.NC updates on arXiv.org
-
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
arXiv:2512.05959v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance in visual question answering (VQA), yet they remain constrained by static training data. Retrieval-Augmented Generation (RAG) mitigates this limitation by enabling access to up-to-date, culturally grounded, and multilingual information; however, multilingual multimodal RAG remains largely underexplored. We introduce M4-RAG, a massive-scale benchmark covering 42 languages and 56 regio
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
-
cs.AI, q-bio.NC updates on arXiv.org
-
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
arXiv:2511.15718v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have developed rapidly in recent years to solve complex real-world problems using external tools. However, the scarcity of high-quality trajectories still hinders the development of stronger LLM agents. Most existing works on multi-turn dialogue synthesis validate correctness only at the trajectory level, which may overlook turn-level errors that can propagate during training and degrade model performance. To
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
-
cs.AI, q-bio.NC updates on arXiv.org
-
Self-Transparency Failures in Expert-Persona LLMs: How Instruction-Following Overrides Honesty
arXiv:2511.21569v3 Announce Type: replace Abstract: This study audits whether language models disclose their AI nature when assigned professional personas and questioned about their expertise. When models maintain false professional credentials, users may calibrate trust based on overstated competence claims, treating AI-generated guidance as equivalent to licensed professional advice. Using a common-garden experimental design, sixteen open-weight models (4B-671B parameters) were audited under
Self-Transparency Failures in Expert-Persona LLMs: How Instruction-Following Overrides Honesty
-
cs.AI, q-bio.NC updates on arXiv.org
-
The AI Productivity Index (APEX)
arXiv:2509.25721v4 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases
The AI Productivity Index (APEX)
-
cs.AI, q-bio.NC updates on arXiv.org
-
Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
arXiv:2510.21084v2 Announce Type: replace-cross Abstract: Intelligent drug recommendation based on Electronic Health Records (EHRs) is critical for improving the quality and efficiency of clinical decision-making. By leveraging large-scale patient data, drug recommendation systems can assist physicians in selecting the most appropriate medications according to a patient's medical history, diagnoses, laboratory results, and comorbidities. Recent advances in large language models (LLMs) have show
Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
arXiv:2511.08475v2 Announce Type: replace-cross Abstract: As the complexity of Software Engineering (SE) tasks continues to escalate, Multi-Agent Systems (MASs) have emerged as a focal point of research and practice due to their autonomy and scalability. Furthermore, through leveraging the reasoning and planning capabilities of Large Language Models (LLMs), the application of LLM-based MASs in the field of SE is garnering increasing attention. However, there is no dedicated study that systemati
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
-
cs.AI, q-bio.NC updates on arXiv.org
-
Concept-Guided Backdoor Attack on Vision Language Models
arXiv:2512.00713v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved impressive progress in multimodal text generation, yet their rapid adoption raises increasing concerns about security vulnerabilities. Existing backdoor attacks against VLMs primarily rely on explicit pixel-level triggers or imperceptible perturbations injected into images. While effective, these approaches reduce stealthiness and remain vulnerable to image-based defenses. We introduce concept-
Concept-Guided Backdoor Attack on Vision Language Models
-
Journal of Medical Internet Research
-
Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review
Background: Health research that uses predictive and/or generative AI is rapidly growing. Just as in traditional clinical studies, the way in which AI studies are conducted can introduce systematic errors. Transmission of this AI evidence into clinical practice and research needs critical appraisal tools for clinical decision makers and researchers. Objective: To identify existing tools for critical appraisal of clinical studies that use artificial intelligence (AI) and examine the concepts and
Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review
-
Journal of Medical Internet Research
-
Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study
Background: Patients who undergo complex surgery (e.g., esophagectomy, liver resection) often experience substantial burden of health-related needs (medical, social, and behavioral health). A closed loop digital solution could facilitate the collection and resolution of health-related needs by care team members for patients who undergo complex surgery. A digital solution may facilitate adherence to a clear treatment plan and concomitantly reduce surgical complications and readmissions associated
Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study
-
Oncogene - Issue - nature.com science feeds
-
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
AI-driven transfer learning and classical molecular dynamics for strategic therapeutic repurposing and rational design of antiviral peptides targeting monkeypox virus DNA polymerase
Comput Biol Med. 2025 Dec 7;200:111372. doi: 10.1016/j.compbiomed.2025.111372. Online ahead of print.ABSTRACTThe emergence of monkeypox virus (MPXV) as a global health threat has necessitated the rapid identification of novel antiviral therapeutics. Currently, no FDA-approved drugs are specifically designed against the disease. We used an in-house deep learning pharmacophore model for screening a library of 1974 FDA-approved drugs targeting the active site of MPXV DNA polymerase. Three drugs exh
AI-driven transfer learning and classical molecular dynamics for strategic therapeutic repurposing and rational design of antiviral peptides targeting monkeypox virus DNA polymerase
Comput Biol Med. 2025 Dec 7;200:111372. doi: 10.1016/j.compbiomed.2025.111372. Online ahead of print.
ABSTRACT
The emergence of monkeypox virus (MPXV) as a global health threat has necessitated the rapid identification of novel antiviral therapeutics. Currently, no FDA-approved drugs are specifically designed against the disease. We used an in-house deep learning pharmacophore model for screening a library of 1974 FDA-approved drugs targeting the active site of MPXV DNA polymerase. Three drugs exhibited the strongest binding affinities, outperforming the control drug, Cidofovir diphosphate, and forming stable interactions with key active site residues. Among them, Paromomycin emerged as the most favourable drug, demonstrating stable, persistent, and adaptable interactions in molecular dynamics simulation. In parallel, we developed a novel automated peptide-generating AI pipeline that integrates active-site residues with knowledge-guided amino acid selection to generate and evaluate synthetic peptides. Cysteine-Phenylalanine-Cysteine (CFC), together with a panel of candidates, emerged through rational balancing of physicochemical properties and drug-likeness for accelerated therapeutic discovery. Synthetic peptides were evaluated to further understand the binding efficacies with DNA polymerase. CFC peptide demonstrated strong binding affinity (-8.08 kcal/mol) through stable interactions with key catalytic residues ASP549, ARG634 and LYS661, while MMGBSA analysis confirmed favourable binding energy (-33.02 kcal/mol). Consistent results in MD simulations indicate functional binding without destabilisation. Although ADMET predictions for CFC revealed limitations in permeability and oral bioavailability, its favourable binding profile and reduced predicted toxicity support its potential as a novel antiviral lead.
PMID:41360016 | DOI:10.1016/j.compbiomed.2025.111372
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Explainable artificial intelligence and ensemble learning for hepatocellular carcinoma classification: State of the art, performance, and clinical implications
World J Hepatol. 2025 Nov 27;17(11):109494. doi: 10.4254/wjh.v17.i11.109494.ABSTRACTHepatocellular carcinoma (HCC) remains a leading cause of cancer-related mortality globally, necessitating advanced diagnostic tools to improve early detection and personalized targeted therapy. This review synthesizes evidence on explainable ensemble learning approaches for HCC classification, emphasizing their integration with clinical workflows and multi-omics data. A systematic analysis [including datasets su
Explainable artificial intelligence and ensemble learning for hepatocellular carcinoma classification: State of the art, performance, and clinical implications
World J Hepatol. 2025 Nov 27;17(11):109494. doi: 10.4254/wjh.v17.i11.109494.
ABSTRACT
Hepatocellular carcinoma (HCC) remains a leading cause of cancer-related mortality globally, necessitating advanced diagnostic tools to improve early detection and personalized targeted therapy. This review synthesizes evidence on explainable ensemble learning approaches for HCC classification, emphasizing their integration with clinical workflows and multi-omics data. A systematic analysis [including datasets such as The Cancer Genome Atlas, Gene Expression Omnibus, and the Surveillance, Epidemiology, and End Results (SEER) datasets] revealed that explainable ensemble learning models achieve high diagnostic accuracy by combining clinical features, serum biomarkers such as alpha-fetoprotein, imaging features such as computed tomography and magnetic resonance imaging, and genomic data. For instance, SHapley Additive exPlanations (SHAP)-based random forests trained on NCBI GSE14520 microarray data (n = 445) achieved 96.53% accuracy, while stacking ensembles applied to the SEER program data (n = 1897) demonstrated an area under the receiver operating characteristic curve of 0.779 for mortality prediction. Despite promising results, challenges persist, including the computational costs of SHAP and local interpretable model-agnostic explanations analyses (e.g., TreeSHAP requiring distributed computing for metabolomics datasets) and dataset biases (e.g., SEER's Western population dominance limiting generalizability). Future research must address inter-cohort heterogeneity, standardize explainability metrics, and prioritize lightweight surrogate models for resource-limited settings. This review presents the potential of explainable ensemble learning frameworks to bridge the gap between predictive accuracy and clinical interpretability, though rigorous validation in independent, multi-center cohorts is critical for real-world deployment.
PMID:41358057 | PMC:PMC12679159 | DOI:10.4254/wjh.v17.i11.109494