Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI & Human Co-Improvement for Safer Co-Superintelligence
arXiv:2512.05356v1 Announce Type: new Abstract: Self-improvement is a goal currently exciting the field of AI, but is fraught with danger, and may take time to fully achieve. We advocate that a more achievable and better goal for humanity is to maximize co-improvement: collaboration between human researchers and AIs to achieve co-superintelligence. That is, specifically targeting improving AI systems' ability to work with human researchers to conduct AI research together, from ideation to exper
-
cs.AI, q-bio.NC updates on arXiv.org
-
MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
arXiv:2512.05365v1 Announce Type: new Abstract: Healthcare AI systems have historically faced challenges in merging contextual reasoning, long-term state management, and human-verifiable workflows into a cohesive framework. This paper introduces a completely innovative architecture and concept: combining the Model Context Protocol (MCP) with a specific clinical application, known as MCP-AI. This integration allows intelligent agents to reason over extended periods, collaborate securely, and adh
MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems
arXiv:2512.05449v1 Announce Type: new Abstract: Large language models display a peculiar form of inconsistency: they "know" the correct answer but fail to act on it. In human philosophy, this tension between global judgment and local impulse is called akrasia, or weakness of will. We propose akrasia as a foundational concept for analyzing inconsistency and goal drift in agentic AI systems. To operationalize it, we introduce a preliminary version of the Akrasia Benchmark, currently a structured
The Seeds of Scheming: Weakness of Will in the Building Blocks of Agentic Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
arXiv:2512.05765v1 Announce Type: new Abstract: Influential critiques argue that Large Language Models (LLMs) are a dead end for AGI: "mere pattern matchers" structurally incapable of reasoning or planning. We argue this conclusion misidentifies the bottleneck: it confuses the ocean with the net. Pattern repositories are the necessary System-1 substrate; the missing component is a System-2 coordination layer that selects, constrains, and binds these patterns. We formalize this layer via UCCT, a
The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
-
cs.AI, q-bio.NC updates on arXiv.org
-
XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots
arXiv:2512.05270v1 Announce Type: cross Abstract: As mobile robots increasingly operate alongside humans in shared workspaces, ensuring safe, efficient, and interpretable Human-Robot Interaction (HRI) has become a pressing challenge. While substantial progress has been devoted to human behavior prediction, limited attention has been paid to how humans perceive, interpret, and trust robots' inferences, impeding deployment in safety-critical and socially embedded environments. This paper presents
XR-DT: Extended Reality-Enhanced Digital Twin for Agentic Mobile Robots
-
cs.AI, q-bio.NC updates on arXiv.org
-
Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
arXiv:2512.05397v1 Announce Type: cross Abstract: Major life transitions demand high-stakes decisions, yet people often struggle to imagine how their future selves will live with the consequences. To support this limited capacity for mental time travel, we introduce AI-enabled digital twins that have ``lived through'' simulated life scenarios. Rather than predicting optimal outcomes, these simulations extend prospective cognition by making alternative futures vivid enough to support deliberatio
Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
-
cs.AI, q-bio.NC updates on arXiv.org
-
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
arXiv:2512.05863v1 Announce Type: cross Abstract: Medical question-answering (QA) systems can benefit from advances in large language models (LLMs), but directly applying LLMs to the clinical domain poses challenges such as maintaining factual accuracy and avoiding hallucinations. In this paper, we present a retrieval-augmented generation (RAG) based medical QA system that combines domain-specific knowledge retrieval with open-source LLMs to answer medical questions. We fine-tune two state-of-t
Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework
-
cs.AI, q-bio.NC updates on arXiv.org
-
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
arXiv:2512.05959v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance in visual question answering (VQA), yet they remain constrained by static training data. Retrieval-Augmented Generation (RAG) mitigates this limitation by enabling access to up-to-date, culturally grounded, and multilingual information; however, multilingual multimodal RAG remains largely underexplored. We introduce M4-RAG, a massive-scale benchmark covering 42 languages and 56 regio
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
-
cs.AI, q-bio.NC updates on arXiv.org
-
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
arXiv:2511.15718v2 Announce Type: replace Abstract: Large Language Model (LLM) agents have developed rapidly in recent years to solve complex real-world problems using external tools. However, the scarcity of high-quality trajectories still hinders the development of stronger LLM agents. Most existing works on multi-turn dialogue synthesis validate correctness only at the trajectory level, which may overlook turn-level errors that can propagate during training and degrade model performance. To
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
-
cs.AI, q-bio.NC updates on arXiv.org
-
Self-Transparency Failures in Expert-Persona LLMs: How Instruction-Following Overrides Honesty
arXiv:2511.21569v3 Announce Type: replace Abstract: This study audits whether language models disclose their AI nature when assigned professional personas and questioned about their expertise. When models maintain false professional credentials, users may calibrate trust based on overstated competence claims, treating AI-generated guidance as equivalent to licensed professional advice. Using a common-garden experimental design, sixteen open-weight models (4B-671B parameters) were audited under
Self-Transparency Failures in Expert-Persona LLMs: How Instruction-Following Overrides Honesty
-
cs.AI, q-bio.NC updates on arXiv.org
-
The AI Productivity Index (APEX)
arXiv:2509.25721v4 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases
The AI Productivity Index (APEX)
-
cs.AI, q-bio.NC updates on arXiv.org
-
Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
arXiv:2510.21084v2 Announce Type: replace-cross Abstract: Intelligent drug recommendation based on Electronic Health Records (EHRs) is critical for improving the quality and efficiency of clinical decision-making. By leveraging large-scale patient data, drug recommendation systems can assist physicians in selecting the most appropriate medications according to a patient's medical history, diagnoses, laboratory results, and comorbidities. Recent advances in large language models (LLMs) have show
Chinese Discharge Drug Recommendation in Metabolic Diseases with Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
arXiv:2511.08475v2 Announce Type: replace-cross Abstract: As the complexity of Software Engineering (SE) tasks continues to escalate, Multi-Agent Systems (MASs) have emerged as a focal point of research and practice due to their autonomy and scalability. Furthermore, through leveraging the reasoning and planning capabilities of Large Language Models (LLMs), the application of LLM-based MASs in the field of SE is garnering increasing attention. However, there is no dedicated study that systemati
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
-
cs.AI, q-bio.NC updates on arXiv.org
-
Concept-Guided Backdoor Attack on Vision Language Models
arXiv:2512.00713v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved impressive progress in multimodal text generation, yet their rapid adoption raises increasing concerns about security vulnerabilities. Existing backdoor attacks against VLMs primarily rely on explicit pixel-level triggers or imperceptible perturbations injected into images. While effective, these approaches reduce stealthiness and remain vulnerable to image-based defenses. We introduce concept-
Concept-Guided Backdoor Attack on Vision Language Models
-
Journal of Medical Internet Research
-
Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review
Background: Health research that uses predictive and/or generative AI is rapidly growing. Just as in traditional clinical studies, the way in which AI studies are conducted can introduce systematic errors. Transmission of this AI evidence into clinical practice and research needs critical appraisal tools for clinical decision makers and researchers. Objective: To identify existing tools for critical appraisal of clinical studies that use artificial intelligence (AI) and examine the concepts and
Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review
-
Journal of Medical Internet Research
-
Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study
Background: Patients who undergo complex surgery (e.g., esophagectomy, liver resection) often experience substantial burden of health-related needs (medical, social, and behavioral health). A closed loop digital solution could facilitate the collection and resolution of health-related needs by care team members for patients who undergo complex surgery. A digital solution may facilitate adherence to a clear treatment plan and concomitantly reduce surgical complications and readmissions associated
Exploring a Digital Health Solution to Collect and Manage Health-Related Needs for Patients Who Undergo Complex Surgery: Mixed Methods Study
-
Oncogene - Issue - nature.com science feeds
-
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy
Oncogene, Published online: 08 December 2025; doi:10.1038/s41388-025-03650-3
Molecular stratification of esophageal adenocarcinoma: implications for prognosis and treatment strategy-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
AI-driven transfer learning and classical molecular dynamics for strategic therapeutic repurposing and rational design of antiviral peptides targeting monkeypox virus DNA polymerase
Comput Biol Med. 2025 Dec 7;200:111372. doi: 10.1016/j.compbiomed.2025.111372. Online ahead of print.ABSTRACTThe emergence of monkeypox virus (MPXV) as a global health threat has necessitated the rapid identification of novel antiviral therapeutics. Currently, no FDA-approved drugs are specifically designed against the disease. We used an in-house deep learning pharmacophore model for screening a library of 1974 FDA-approved drugs targeting the active site of MPXV DNA polymerase. Three drugs exh
AI-driven transfer learning and classical molecular dynamics for strategic therapeutic repurposing and rational design of antiviral peptides targeting monkeypox virus DNA polymerase
Comput Biol Med. 2025 Dec 7;200:111372. doi: 10.1016/j.compbiomed.2025.111372. Online ahead of print.
ABSTRACT
The emergence of monkeypox virus (MPXV) as a global health threat has necessitated the rapid identification of novel antiviral therapeutics. Currently, no FDA-approved drugs are specifically designed against the disease. We used an in-house deep learning pharmacophore model for screening a library of 1974 FDA-approved drugs targeting the active site of MPXV DNA polymerase. Three drugs exhibited the strongest binding affinities, outperforming the control drug, Cidofovir diphosphate, and forming stable interactions with key active site residues. Among them, Paromomycin emerged as the most favourable drug, demonstrating stable, persistent, and adaptable interactions in molecular dynamics simulation. In parallel, we developed a novel automated peptide-generating AI pipeline that integrates active-site residues with knowledge-guided amino acid selection to generate and evaluate synthetic peptides. Cysteine-Phenylalanine-Cysteine (CFC), together with a panel of candidates, emerged through rational balancing of physicochemical properties and drug-likeness for accelerated therapeutic discovery. Synthetic peptides were evaluated to further understand the binding efficacies with DNA polymerase. CFC peptide demonstrated strong binding affinity (-8.08 kcal/mol) through stable interactions with key catalytic residues ASP549, ARG634 and LYS661, while MMGBSA analysis confirmed favourable binding energy (-33.02 kcal/mol). Consistent results in MD simulations indicate functional binding without destabilisation. Although ADMET predictions for CFC revealed limitations in permeability and oral bioavailability, its favourable binding profile and reduced predicted toxicity support its potential as a novel antiviral lead.
PMID:41360016 | DOI:10.1016/j.compbiomed.2025.111372
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Explainable artificial intelligence and ensemble learning for hepatocellular carcinoma classification: State of the art, performance, and clinical implications
World J Hepatol. 2025 Nov 27;17(11):109494. doi: 10.4254/wjh.v17.i11.109494.ABSTRACTHepatocellular carcinoma (HCC) remains a leading cause of cancer-related mortality globally, necessitating advanced diagnostic tools to improve early detection and personalized targeted therapy. This review synthesizes evidence on explainable ensemble learning approaches for HCC classification, emphasizing their integration with clinical workflows and multi-omics data. A systematic analysis [including datasets su
Explainable artificial intelligence and ensemble learning for hepatocellular carcinoma classification: State of the art, performance, and clinical implications
World J Hepatol. 2025 Nov 27;17(11):109494. doi: 10.4254/wjh.v17.i11.109494.
ABSTRACT
Hepatocellular carcinoma (HCC) remains a leading cause of cancer-related mortality globally, necessitating advanced diagnostic tools to improve early detection and personalized targeted therapy. This review synthesizes evidence on explainable ensemble learning approaches for HCC classification, emphasizing their integration with clinical workflows and multi-omics data. A systematic analysis [including datasets such as The Cancer Genome Atlas, Gene Expression Omnibus, and the Surveillance, Epidemiology, and End Results (SEER) datasets] revealed that explainable ensemble learning models achieve high diagnostic accuracy by combining clinical features, serum biomarkers such as alpha-fetoprotein, imaging features such as computed tomography and magnetic resonance imaging, and genomic data. For instance, SHapley Additive exPlanations (SHAP)-based random forests trained on NCBI GSE14520 microarray data (n = 445) achieved 96.53% accuracy, while stacking ensembles applied to the SEER program data (n = 1897) demonstrated an area under the receiver operating characteristic curve of 0.779 for mortality prediction. Despite promising results, challenges persist, including the computational costs of SHAP and local interpretable model-agnostic explanations analyses (e.g., TreeSHAP requiring distributed computing for metabolomics datasets) and dataset biases (e.g., SEER's Western population dominance limiting generalizability). Future research must address inter-cohort heterogeneity, standardize explainability metrics, and prioritize lightweight surrogate models for resource-limited settings. This review presents the potential of explainable ensemble learning frameworks to bridge the gap between predictive accuracy and clinical interpretability, though rigorous validation in independent, multi-center cohorts is critical for real-world deployment.
PMID:41358057 | PMC:PMC12679159 | DOI:10.4254/wjh.v17.i11.109494
-
MIT Technology Review

-
The State of AI: A vision of the world in 2030
Welcome back to The State of AI, a new collaboration between the Financial Times and MIT Technology Review. Every Monday, writers from both publications debate one aspect of the generative AI revolution reshaping global power. You can read the rest of the series here. In this final edition, MIT Technology Review’s senior AI editor Will Douglas Heaven talks with Tim Bradshaw, FT global tech correspondent, about where AI will go next, and what our world will look like in the next five years.
The State of AI: A vision of the world in 2030
Welcome back to The State of AI, a new collaboration between the Financial Times and MIT Technology Review. Every Monday, writers from both publications debate one aspect of the generative AI revolution reshaping global power. You can read the rest of the series here.
In this final edition, MIT Technology Review’s senior AI editor Will Douglas Heaven talks with Tim Bradshaw, FT global tech correspondent, about where AI will go next, and what our world will look like in the next five years.
(As part of this series, join MIT Technology Review’s editor in chief, Mat Honan, and editor at large, David Rotman, for an exclusive conversation with Financial Times columnist Richard Waters on how AI is reshaping the global economy. Live on Tuesday, December 9 at 1:00 p.m. ET. This is a subscriber-only event and you can sign up here.)

Will Douglas Heaven writes:
Every time I’m asked what’s coming next, I get a Luke Haines song stuck in my head: “Please don’t ask me about the future / I am not a fortune teller.” But here goes. What will things be like in 2030? My answer: same but different.
There are huge gulfs of opinion when it comes to predicting the near-future impacts of generative AI. In one camp we have the AI Futures Project, a small donation-funded research outfit led by former OpenAI researcher Daniel Kokotajlo. The nonprofit made a big splash back in April with AI 2027, a speculative account of what the world will look like two years from now.
The story follows the runaway advances of an AI firm called OpenBrain (any similarities are coincidental, etc.) all the way to a choose-your-own-adventure-style boom or doom ending. Kokotajlo and his coauthors make no bones about their expectation that in the next decade the impact of AI will exceed that of the Industrial Revolution—a 150-year period of economic and social upheaval so great that we still live in the world it wrought.
At the other end of the scale we have team Normal Technology: Arvind Narayanan and Sayash Kapoor, a pair of Princeton University researchers and coauthors of the book AI Snake Oil, who push back not only on most of AI 2027’s predictions but, more important, on its foundational worldview. That’s not how technology works, they argue.
Advances at the cutting edge may come thick and fast, but change across the wider economy, and society as a whole, moves at human speed. Widespread adoption of new technologies can be slow; acceptance slower. AI will be no different.
What should we make of these extremes? ChatGPT came out three years ago last month, but it’s still not clear just how good the latest versions of this tech are at replacing lawyers or software developers or (gulp) journalists. And new updates no longer bring the step changes in capability that they once did.
And yet this radical technology is so new it would be foolish to write it off so soon. Just think: Nobody even knows exactly how this technology works—let alone what it’s really for.
As the rate of advance in the core technology slows down, applications of that tech will become the main differentiator between AI firms. (Witness the new browser wars and the chatbot pick-and-mix already on the market.) At the same time, high-end models are becoming cheaper to run and more accessible. Expect this to be where most of the action is: New ways to use existing models will keep them fresh and distract people waiting in line for what comes next.
Meanwhile, progress continues beyond LLMs. (Don’t forget—there was AI before ChatGPT, and there will be AI after it too.) Technologies such as reinforcement learning—the powerhouse behind AlphaGo, DeepMind’s board-game-playing AI that beat a Go grand master in 2016—is set to make a comeback. There’s also a lot of buzz around world models, a type of generative AI with a stronger grip on how the physical world fits together than LLMs display.
Ultimately, I agree with team Normal Technology that rapid technological advances do not translate to economic or societal ones straight away. There’s just too much messy human stuff in the middle.
But Tim, over to you. I’m curious to hear what your tea leaves are saying.

Tim Bradshaw responds:
Will, I am more confident than you that the world will look quite different in 2030. In five years’ time, I expect the AI revolution to have proceeded apace. But who gets to benefit from those gains will create a world of AI haves and have-nots.
It seems inevitable that the AI bubble will burst sometime before the end of the decade. Whether a venture capital funding shakeout comes in six months or two years (I feel the current frenzy still has some way to run), swathes of AI app developers will disappear overnight. Some will see their work absorbed by the models upon which they depend. Others will learn the hard way that you can’t sell services that cost $1 for 50 cents without a firehose of VC funding.
How many of the foundation model companies survive is harder to call, but it already seems clear that OpenAI’s chain of interdependencies within Silicon Valley make it too big to fail. Still, a funding reckoning will force it to ratchet up pricing for its services.
When OpenAI was created in 2015, it pledged to “advance digital intelligence in the way that is most likely to benefit humanity as a whole.” That seems increasingly untenable. Sooner or later, the investors who bought in at a $500 billion price tag will push for returns. Those data centers won’t pay for themselves. By that point, many companies and individuals will have come to depend on ChatGPT or other AI services for their everyday workflows. Those able to pay will reap the productivity benefits, scooping up the excess computing power as others are priced out of the market.
Being able to layer several AI services on top of each other will provide a compounding effect. One example I heard on a recent trip to San Francisco: Ironing out the kinks in vibe coding is simply a matter of taking several passes at the same problem and then running a few more AI agents to look for bugs and security issues. That sounds incredibly GPU-intensive, implying that making AI really deliver on the current productivity promise will require customers to pay far more than most do today.
The same holds true in physical AI. I fully expect robotaxis to be commonplace in every major city by the end of the decade, and I even expect to see humanoid robots in many homes. But while Waymo’s Uber-like prices in San Francisco and the kinds of low-cost robots produced by China’s Unitree give the impression today that these will soon be affordable for all, the compute cost involved in making them useful and ubiquitous seems destined to turn them into luxuries for the well-off, at least in the near term.
The rest of us, meanwhile, will be left with an internet full of slop and unable to afford AI tools that actually work.
Perhaps some breakthrough in computational efficiency will avert this fate. But the current AI boom means Silicon Valley’s AI companies lack the incentives to make leaner models or experiment with radically different kinds of chips. That only raises the likelihood that the next wave of AI innovation will come from outside the US, be that China, India, or somewhere even farther afield.
Silicon Valley’s AI boom will surely end before 2030, but the race for global influence over the technology’s development—and the political arguments about how its benefits are distributed—seem set to continue well into the next decade.
Will replies:
I am with you that the cost of this technology is going to lead to a world of haves and have-nots. Even today, $200+ a month buys power users of ChatGPT or Gemini a very different experience from that of people on the free tier. That capability gap is certain to increase as model makers seek to recoup costs.
We’re going to see massive global disparities too. In the Global North, adoption has been off the charts. A recent report from Microsoft’s AI Economy Institute notes that AI is the fastest-spreading technology in human history: “In less than three years, more than 1.2 billion people have used AI tools, a rate of adoption faster than the internet, the personal computer, or even the smartphone.” And yet AI is useless without ready access to electricity and the internet; swathes of the world still have neither.
I still remain skeptical that we will see anything like the revolution that many insiders promise (and investors pray for) by 2030. When Microsoft talks about adoption here, it’s counting casual users rather than measuring long-term technological diffusion, which takes time. Meanwhile, casual users get bored and move on.
How about this: If I live with a domestic robot in five years’ time, you can send your laundry to my house in a robotaxi any day of the week.
JK! As if I could afford one.
Further reading
What is AI? It sounds like a stupid question, but it’s one that’s never been more urgent. In this deep dive, Will unpacks decades of spin and speculation to get to the heart of our collective technodream.
AGI—the idea that machines will be as smart as humans—has hijacked an entire industry (and possibly the US economy). For MIT Technology Review’s recent New Conspiracy Age package, Will takes a provocative look at how AGI is like a conspiracy.
The FT examined the economics of self-driving cars this summer, asking who will foot the multi-billion-dollar bill to buy enough robotaxis to serve a big city like London or New York.
A plausible counter-argument to Tim’s thesis on AI inequalities is that freely available open-source (or more accurately, “open weight”) models will keep pulling down prices. The US may want frontier models to be built on US chips but it is already losing the global south to Chinese software.