Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Evaluation Gap in Medicine, AI and LLMs: Navigating Elusive Ground Truth & Uncertainty via a Probabilistic Paradigm
arXiv:2601.05500v1 Announce Type: new Abstract: Benchmarking the relative capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth answers from experts. This ambiguity is particularly consequential in medicine where uncertainty is pervasive. In this paper, we introduce a probabilistic paradigm to theoretically explain how high certainty in ground truth answers is almost always necessary for e
-
cs.AI, q-bio.NC updates on arXiv.org
-
Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making
arXiv:2601.05529v1 Announce Type: new Abstract: One mistake by an AI system in a safety-critical setting can cost lives. As Large Language Models (LLMs) become integral to robotics decision-making, the physical dimension of risk grows; a single wrong instruction can directly endanger human safety. This paper addresses the urgent need to systematically evaluate LLM performance in scenarios where even minor errors are catastrophic. Through a qualitative evaluation of a fire evacuation scenario, w
Safety Not Found (404): Hidden Risks of LLM-Based Robotics Decision Making
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes
arXiv:2601.05293v1 Announce Type: cross Abstract: Agentic AI marks an important transition from single-step generative models to systems capable of reasoning, planning, acting, and adapting over long-lasting tasks. By integrating memory, tool use, and iterative decision cycles, these systems enable continuous, autonomous workflows in real-world environments. This survey examines the implications of agentic AI for cybersecurity. On the defensive side, agentic capabilities enable continuous monit
A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes
-
cs.AI, q-bio.NC updates on arXiv.org
-
Streamlining evidence based clinical recommendations with large language models
arXiv:2505.10282v2 Announce Type: replace-cross Abstract: Clinical evidence underpins informed healthcare decisions, yet integrating it into real-time practice remains challenging due to intensive workloads, complex procedures, and time constraints. This study presents Quicker, an LLM-powered system that automates evidence synthesis and generates clinical recommendations following standard guideline development workflows. Quicker delivers an end-to-end pipeline from clinical questions to recomm
Streamlining evidence based clinical recommendations with large language models
-
cs.AI, q-bio.NC updates on arXiv.org
-
CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records
arXiv:2507.22533v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for improving clinical decision support and reducing physician burnout by synthesizing complex, longitudinal cancer Electronic Health Records (EHRs). However, their implementation in this critical field faces three primary challenges: the inability to effectively process the extensive length and fragmented nature of patient records for accurate temporal analysis; a heightened risk of
CliCARE: Grounding Large Language Models in Clinical Guidelines for Decision Support over Longitudinal Cancer Electronic Health Records
-
cs.AI, q-bio.NC updates on arXiv.org
-
Benchmarking LLM-based Agents for Single-cell Omics Analysis
arXiv:2508.13201v2 Announce Type: replace-cross Abstract: The surge in multimodal single-cell omics data exposes limitations in traditional, manually defined analysis workflows. AI agents offer a paradigm shift, enabling adaptive planning, executable code generation, traceable decisions, and real-time knowledge fusion. However, the lack of a comprehensive benchmark critically hinders progress. We introduce a novel benchmarking evaluation system to rigorously assess agent capabilities in single-
Benchmarking LLM-based Agents for Single-cell Omics Analysis
-
FDA Press Releases RSS Feed
-
FDA Increases Flexibility on Requirements for Cell and Gene Therapies to Advance Innovation
The U.S. Food and Drug Administration today announced it is sharing information about the agency’s flexible approach to overseeing chemistry, manufacturing and control (CMC) requirements for cell and gene therapies (CGT).
FDA Increases Flexibility on Requirements for Cell and Gene Therapies to Advance Innovation
-
TechCrunch
-
Google removes AI Overviews for certain medical queries
This follows an investigation by the Guardian that found Google AI Overviews offering misleading information in response to some health-related queries.
Google removes AI Overviews for certain medical queries
-
STAT

-
Opinion: The NIH has lost its scientific integrity. So we left
We are National Institutes of Health scientists and administrators with more than 50 years of collective civil service. Or, more accurately, we were NIH scientists and administrators.Read the rest…
Opinion: The NIH has lost its scientific integrity. So we left
We are National Institutes of Health scientists and administrators with more than 50 years of collective civil service.
Or, more accurately, we were NIH scientists and administrators.


© Adobe
-
MRD
-
Personalizing Treatment for Pancreatic Ductal Adenocarcinoma: The Emerging Role of Minimal Residual Disease in Perioperative Decision-Making
Cancers (Basel). 2025 Dec 27;18(1):94. doi: 10.3390/cancers18010094.ABSTRACTPancreatic ductal adenocarcinoma (PDAC) is a highly aggressive malignancy with poor long-term survival despite advances in surgical techniques, systemic therapies, and perioperative management. High rates of systemic recurrence following curative-intent resection suggest that many patients harbor minimal residual disease (MRD), microscopic tumor burden that persists postoperatively and remains undetectable by conventiona
Personalizing Treatment for Pancreatic Ductal Adenocarcinoma: The Emerging Role of Minimal Residual Disease in Perioperative Decision-Making
Cancers (Basel). 2025 Dec 27;18(1):94. doi: 10.3390/cancers18010094.
ABSTRACT
Pancreatic ductal adenocarcinoma (PDAC) is a highly aggressive malignancy with poor long-term survival despite advances in surgical techniques, systemic therapies, and perioperative management. High rates of systemic recurrence following curative-intent resection suggest that many patients harbor minimal residual disease (MRD), microscopic tumor burden that persists postoperatively and remains undetectable by conventional diagnostic tools. Recent advances in liquid biopsy technologies, particularly circulating tumor DNA (ctDNA) analysis, alongside detailed characterization of the PDAC mutational landscape, offer a promising non-invasive approach for MRD detection. Emerging evidence indicates that MRD status can serve as a sensitive prognostic biomarker, identify patients at high risk of relapse, and guide personalized perioperative therapy, including optimization of adjuvant treatment. This review summarizes current knowledge on the biology and detection of MRD in PDAC, its implications for perioperative risk stratification and treatment decision-making, and discusses future directions for integrating MRD assessment into clinical practice to enable more precise, individualized patient management.
PMID:41514607 | PMC:PMC12784771 | DOI:10.3390/cancers18010094
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Integrative Genomic and AI Approaches to Lung Cancer and Implications for Disease Prevention in Former Smokers
Int J Mol Sci. 2026 Jan 4;27(1):521. doi: 10.3390/ijms27010521.ABSTRACTTobacco smoking accounts for nearly 90% of lung cancer deaths worldwide, yet the mechanisms underlying persistent cancer risk in former smokers are not fully understood. Epidemiological evidence shows that more than 40% of lung cancers develop over 15 years after cessation, demonstrating that while some smoking-induced molecular alterations resolve rapidly, others remain as long-lasting scars that promote carcinogenesis. This
Integrative Genomic and AI Approaches to Lung Cancer and Implications for Disease Prevention in Former Smokers
Int J Mol Sci. 2026 Jan 4;27(1):521. doi: 10.3390/ijms27010521.
ABSTRACT
Tobacco smoking accounts for nearly 90% of lung cancer deaths worldwide, yet the mechanisms underlying persistent cancer risk in former smokers are not fully understood. Epidemiological evidence shows that more than 40% of lung cancers develop over 15 years after cessation, demonstrating that while some smoking-induced molecular alterations resolve rapidly, others remain as long-lasting scars that promote carcinogenesis. This review synthesizes longitudinal and cross-sectional genomic, epigenomic, and transcriptomic studies of airway and lung tissues to distinguish persistent from nonpersistent smoking-induced molecular alterations. Persistent alterations include somatic mutations in TP53 and KRAS, DNA methylation at tumor suppressor loci, dysregulated noncoding RNAs, chromosomal instability, and epigenetic age acceleration. Nonpersistent changes, such as acute inflammatory responses and detoxification pathways, generally normalize within months to several years following cessation. Multi-omics profiling reveals coordinated patterns of dysregulation consistent with field cancerization in former smokers. In addition, the integration of multi-omics data with artificial intelligence may enable composite molecular signatures for stratifying high-risk former smokers, link molecular persistence to clinical outcomes, and inform chemoprevention strategies. Collectively, these observations clarify which molecular alterations sustain long-term cancer risk despite smoking cessation and highlight opportunities for precision prevention and earlier detection in high-risk populations.
PMID:41516393 | PMC:PMC12786486 | DOI:10.3390/ijms27010521
-
AI News
-
Autonomy without accountability: The real AI risk
If you have ever taken a self-driving Uber through downtown LA, you might recognise the strange sense of uncertainty that settles in when there is no driver and no conversation, just a quiet car making assumptions about the world around it. The journey feels fine until the car misreads a shadow or slows abruptly for something harmless. In that moment you see the real issue with autonomy. It does not panic when it should, and that gap between confidence and judgement is where trust is either earn
Autonomy without accountability: The real AI risk
If you have ever taken a self-driving Uber through downtown LA, you might recognise the strange sense of uncertainty that settles in when there is no driver and no conversation, just a quiet car making assumptions about the world around it. The journey feels fine until the car misreads a shadow or slows abruptly for something harmless. In that moment you see the real issue with autonomy. It does not panic when it should, and that gap between confidence and judgement is where trust is either earned or lost. Much of today’s enterprise AI feels remarkably similar. It is competent without being confident, and efficient without being empathetic, which is why the deciding factor in every successful deployment is no longer computing power but trust.
The MLQ State of AI in Business 2025 [PDF] report puts a sharp number on this. 95% of early AI pilots fail to produce measurable ROI, not because the technology is weak but because it is mismatched to the problems organisations are trying to solve. The pattern repeats itself in industries. Leaders get uneasy when they can’t tell if the output is right, teams are unsure whether dashboards can be trusted, and customers quickly lose patience when an interaction feels automated rather than supported. Anyone who has been locked out of their bank account while the automated recovery system insists their answers are wrong knows how quickly confidence evaporates.
Klarna remains the most publicised example of large-scale automation in action. The company has now halved its workforce since 2022 and says internal AI systems are performing the work of 853 full-time roles, up from 700 earlier this year. Revenues have risen 108%, while average employee compensation has increased 60%, funded in part by those operational gains. Yet the picture is more complicated. Klarna still reported a 95 million dollar quarterly loss, and its CEO has warned that further staff reductions are likely. It shows that automation alone does not create stability. Without accountability and structure, the experience breaks down long before the AI does. As Jason Roos, CEO of CCaaS provider Cirrus, puts it, “Any transformation that unsettles confidence, inside or outside the business, carries a cost you cannot ignore. it can leave you worse off.”
We have already seen what happens when autonomy runs ahead of accountability. The UK’s Department for Work and Pensions used an algorithm that wrongly flagged around 200,000 housing-benefit claims as potentially fraudulent, even though the majority were legitimate. The problem wasn’t the technology. It was the absence of clear ownership over its decisions. When an automated system suspends the wrong account, rejects the wrong claim or creates unnecessary fear, the issue is never just “why did the model misfire?” It’s “who owns the outcome?” Without that answer, trust becomes fragile.
“The missing step is always readiness,” says Roos. “If the process, the data and the guardrails aren’t in place, autonomy doesn’t accelerate performance, it amplifies the weaknesses. Accountability has to come first. Start with the outcome, find where effort is being wasted, check your readiness and governance, and only then automate. Skip those steps and accountability disappears just as fast as the efficiency gains arrive.”
Part of the problem is an obsession with scale without the grounding that makes scale sustainable. Many organisations push toward autonomous agents that can act decisively, yet very few pause to consider what happens when those actions drift outside expected boundaries. The Edelman Trust Barometer [PDF] shows a steady decline in public trust in AI over the past five years, and a joint KPMG and University of Melbourne study found that workers prefer more human involvement in almost half the tasks examined. The findings reinforce a simple point. Trust rarely comes from pushing models harder. It comes from people taking the time to understand how decisions are made, and from governance that behaves less like a brake pedal and more like a steering wheel.
The same dynamics appear on the customer side. PwC’s trust research reveals a wide gulf between perception and reality. Most executives believe customers trust their organisation, while only a minority of customers agree. Other surveys show that transparency helps to close this gap, with large majorities of consumers wanting clear disclosure when AI is used in service experiences. Without that clarity, people do not feel reassured. They feel misled, and the relationship becomes strained. Companies that communicate openly about their AI use are not only protecting trust but also normalising the idea that technology and human support can co-exist.
Some of the confusion stems from the term “agentic AI” itself. Much of the market treats it as something unpredictable or self-directing, when in reality it is workflow automation with reasoning and recall. It is a structured way for systems to make modest decisions inside parameters designed by people. The deployments that scale safely all follow the same sequence. They start with the outcome they want to improve, then look at where unnecessary effort sits in the workflow, then assess whether their systems and teams are ready for autonomy, and only then choose the technology. Reversing that order does not speed anything up. It simply creates faster mistakes. As Roos says, AI should expand human judgement, not replace it.
All of this points toward a wider truth. Every wave of automation eventually becomes a social question rather than a purely technical one. Amazon built its dominance through operational consistency, but it also built a level of confidence that the parcel would arrive. When that confidence dips, customers move on. AI follows the same pattern. You can deploy sophisticated, self-correcting systems, but if the customer feels tricked or misled at any point, the trust breaks. Internally, the same pressures apply. The KPMG global study [PDF] highlights how quickly employees disengage when they do not understand how decisions are made or who is accountable for them. Without that clarity, adoption stalls.
As agentic systems take on more conversational roles, the emotional dimension becomes even more significant. Early reviews of autonomous chat interactions show that people now judge their experience not only by whether they were helped but also by whether the interaction felt attentive and respectful. A customer who feels dismissed rarely keeps the frustration to themselves. The emotional tone of AI is becoming a genuine operational factor, and systems that cannot meet that expectation risk becoming liabilities.
The difficult truth is that technology will continue to move faster than people’s instinctive comfort with it. Trust will always lag behind innovation. That is not an argument against progress. It is an argument for maturity. Every AI leader should be asking whether they would trust the system with their own data, whether they can explain its last decision in plain language, and who steps in when something goes wrong. If those answers are unclear, the organisation is not leading transformation. It is preparing an apology.
Roos puts it simply, “Agentic AI is not the concern. Unaccountable AI is.”
When trust goes, adoption goes, and the project that looked transformative becomes another entry in the 95% failure rate. Autonomy is not the enemy. Forgetting who is responsible is. The organisations that keep a human hand on the wheel will be the ones still in control when the self-driving hype eventually fades.
The post Autonomy without accountability: The real AI risk appeared first on AI News.
-
Nature Medicine
-
BCMA-directed mRNA CAR-T cell therapy for myasthenia gravis: exploratory biomarker analysis of a placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04170-zAnalysis of a placebo-controlled trial of a BCMA-targeting CAR-T cell therapy in patients with myasthenia gravis shows that CAR-T cell infusion selectively remodels the systemic immune environment, with elimination of BCMA-high plasma cells and activated plasmacytoid dendritic cells and changes in the autoreactive B cell repertoire.
BCMA-directed mRNA CAR-T cell therapy for myasthenia gravis: exploratory biomarker analysis of a placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04170-z
Analysis of a placebo-controlled trial of a BCMA-targeting CAR-T cell therapy in patients with myasthenia gravis shows that CAR-T cell infusion selectively remodels the systemic immune environment, with elimination of BCMA-high plasma cells and activated plasmacytoid dendritic cells and changes in the autoreactive B cell repertoire.-
Nature Medicine
-
BCMA-directed mRNA CAR T cell therapy for myasthenia gravis: a randomized, double-blind, placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04171-yIn a randomized, double-blind, placebo-controlled trial comparing autologous mRNA-engineered BCMA-targeting CAR T cell therapy versus placebo in patients with generalized myasthenia gravis, a significantly higher percentage of patients exhibited a reduction in disease activity in the treatment arm than in the placebo arm.
BCMA-directed mRNA CAR T cell therapy for myasthenia gravis: a randomized, double-blind, placebo-controlled phase 2b trial
Nature Medicine, Published online: 09 January 2026; doi:10.1038/s41591-025-04171-y
In a randomized, double-blind, placebo-controlled trial comparing autologous mRNA-engineered BCMA-targeting CAR T cell therapy versus placebo in patients with generalized myasthenia gravis, a significantly higher percentage of patients exhibited a reduction in disease activity in the treatment arm than in the placebo arm.-
cs.AI, q-bio.NC updates on arXiv.org
-
Formal Analysis of AGI Decision-Theoretic Models and the Confrontation Question
arXiv:2601.04234v1 Announce Type: new Abstract: Artificial General Intelligence (AGI) may face a confrontation question: under what conditions would a rationally self-interested AGI choose to seize power or eliminate human control (a confrontation) rather than remain cooperative? We formalize this in a Markov decision process with a stochastic human-initiated shutdown event. Building on results on convergent instrumental incentives, we show that for almost all reward functions a misaligned agen
Formal Analysis of AGI Decision-Theoretic Models and the Confrontation Question
-
cs.AI, q-bio.NC updates on arXiv.org
-
Systems Explaining Systems: A Framework for Intelligence and Consciousness
arXiv:2601.04269v1 Announce Type: new Abstract: This paper proposes a conceptual framework in which intelligence and consciousness emerge from relational structure rather than from prediction or domain-specific mechanisms. Intelligence is defined as the capacity to form and integrate causal connections between signals, actions, and internal states. Through context enrichment, systems interpret incoming information using learned relational structure that provides essential context in an efficien
Systems Explaining Systems: A Framework for Intelligence and Consciousness
-
cs.AI, q-bio.NC updates on arXiv.org
-
An ASP-based Solution to the Medical Appointment Scheduling Problem
arXiv:2601.04274v1 Announce Type: new Abstract: This paper presents an Answer Set Programming (ASP)-based framework for medical appointment scheduling, aimed at improving efficiency, reducing administrative overhead, and enhancing patient-centered care. The framework personalizes scheduling for vulnerable populations by integrating Blueprint Personas. It ensures real-time availability updates, conflict-free assignments, and seamless interoperability with existing healthcare platforms by central
An ASP-based Solution to the Medical Appointment Scheduling Problem
-
cs.AI, q-bio.NC updates on arXiv.org
-
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
arXiv:2601.04577v1 Announce Type: new Abstract: While AI innovation accelerates rapidly, the intellectual process behind breakthroughs -- how researchers identify gaps, synthesize prior work, and generate insights -- remains poorly understood. The lack of structured data on scientific reasoning hinders systematic analysis and development of AI research agents. We introduce Sci-Reasoning, the first dataset capturing the intellectual synthesis behind high-quality AI research. Using community-vali
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
-
cs.AI, q-bio.NC updates on arXiv.org
-
Autonomous Agents on Blockchains: Standards, Execution Models, and Trust Boundaries
arXiv:2601.04583v1 Announce Type: new Abstract: Advances in large language models have enabled agentic AI systems that can reason, plan, and interact with external tools to execute multi-step workflows, while public blockchains have evolved into a programmable substrate for value transfer, access control, and verifiable state transitions. Their convergence introduces a high-stakes systems challenge: designing standard, interoperable, and secure interfaces that allow agents to observe on-chain s
Autonomous Agents on Blockchains: Standards, Execution Models, and Trust Boundaries
-
cs.AI, q-bio.NC updates on arXiv.org
-
ResMAS: Resilience Optimization in LLM-based Multi-agent Systems
arXiv:2601.04694v1 Announce Type: new Abstract: Large Language Model-based Multi-Agent Systems (LLM-based MAS), where multiple LLM agents collaborate to solve complex tasks, have shown impressive performance in many areas. However, MAS are typically distributed across different devices or environments, making them vulnerable to perturbations such as agent failures. While existing works have studied the adversarial attacks and corresponding defense strategies, they mainly focus on reactively det