Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
CLINB: A Climate Intelligence Benchmark for Foundational Models
arXiv:2511.11597v1 Announce Type: new Abstract: Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate s
-
cs.AI, q-bio.NC updates on arXiv.org
-
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
arXiv:2412.15925v1 Announce Type: cross Abstract: Problem: Pancreas radiological imaging is challenging due to the small size, blurred boundaries, and variability of shape and position of the organ among patients. Goal: In this work we present MiniGPT-Pancreas, a Multimodal Large Language Model (MLLM), as an interactive chatbot to support clinicians in pancreas cancer diagnosis by integrating visual and textual information. Methods: MiniGPT-v2, a general-purpose MLLM, was fine-tuned in a cascad
MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection
-
cs.AI, q-bio.NC updates on arXiv.org
-
Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
arXiv:2511.11590v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly embedded in NHS workflows, but its probabilistic and adaptive behaviour conflicts with the deterministic assumptions underpinning existing clinical-safety standards. DCB0129 and DCB0160 provide strong governance for conventional software yet do not define how AI-specific transparency, interpretability, or model drift should be evidenced within Safety Cases, Hazard Logs, or post-market monitoring. This
Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
arXiv:2511.13373v1 Announce Type: cross Abstract: Large Language Models (LLMs) face significant challenges in distributed healthcare, including consolidating specialized domain knowledge across institutions while maintaining privacy, reducing computational overhead, and preventing catastrophic forgetting during model updates.This paper presents a systematic evaluation of six parameter-space merging techniques applied to two architecturally compatible medical LLMs derived from the Mistral-7B bas
A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
arXiv:2511.13525v1 Announce Type: cross Abstract: Fairness in artificial intelligence (AI) has become a growing concern due to discriminatory outcomes in AI-based decision-making systems. While various methods have been proposed to mitigate bias, most rely on complete demographic information, an assumption often impractical due to legal constraints and the risk of reinforcing discrimination. This survey examines fairness in AI when demographics are incomplete, addressing the gap between traditi
AI Fairness Beyond Complete Demographics: Current Achievements and Future Directions
-
npj Digital Medicine
-
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
npj Digital Medicine, Published online: 17 November 2025; doi:10.1038/s41746-025-02025-yA large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions
npj Digital Medicine, Published online: 17 November 2025; doi:10.1038/s41746-025-02025-y
A large language model-based approach to quantifying the effects of social determinants in liver transplant decisions-
Journal of Medical Internet Research
-
Methods for Analytical Validation of Novel Digital Clinical Measures: Implementation Feasibility Evaluation Using Real-World Datasets
Background: Sensor-based digital health technologies (sDHTs) are increasingly used to support scientific and clinical decision-making. The digital measures (DMs) they generate offer significant potential to accelerate the drug development timeline, decrease clinical trial costs, and improve access to care. However, choosing appropriate statistical methodology when conducting analytical validation (AV) of a DM is complicated, particularly for novel DMs, for which appropriate, established referenc
Methods for Analytical Validation of Novel Digital Clinical Measures: Implementation Feasibility Evaluation Using Real-World Datasets
-
cs.AI, q-bio.NC updates on arXiv.org
-
Case Study: Transformer-Based Solution for the Automatic Digitization of Gas Plants
arXiv:2511.08609v1 Announce Type: cross Abstract: The energy transition is a key theme of the last decades to determine a future of eco-sustainability, and an area of such importance cannot disregard digitization, innovation and the new technological tools available. This is the context in which the Generative Artificial Intelligence models described in this paper are positioned, developed by Engineering Ingegneria Informatica SpA in order to automate the plant structures acquisition of SNAM en
Case Study: Transformer-Based Solution for the Automatic Digitization of Gas Plants
-
cs.AI, q-bio.NC updates on arXiv.org
-
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
arXiv:2511.08721v1 Announce Type: cross Abstract: In behavioral sciences, experiments such as the ultimatum game are conducted to assess preferences for fairness or self-interest of study participants. In the dictator game, a simplified version of the ultimatum game where only one of two players makes a single decision, the dictator unilaterally decides how to split a fixed sum of money between themselves and the other player. Although recent studies have explored behavioral patterns of AI agen
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
-
cs.AI, q-bio.NC updates on arXiv.org
-
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
arXiv:2510.00229v4 Announce Type: replace Abstract: The deployment of Large Language Models (LLMs) as agentic orchestrators has revolutionized task automation, but the need for privacy-preserving, cost-effective solutions demands on-device inference capabilities. However, local LLMs consistently underperform compared to frontier models in tool calling scenarios, struggling with both tool selection from large tool sets and accurate argument generation for complex parameter structures. We introdu
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future Trends
arXiv:2410.15281v4 Announce Type: replace-cross Abstract: With the broader adoption and highly successful development of Large Language Models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. In this paper, we first
LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future Trends
-
cs.AI, q-bio.NC updates on arXiv.org
-
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
arXiv:2508.01159v2 Announce Type: replace-cross Abstract: This study evaluates the capacity of large language models (LLMs) to generate structured clinical consultation templates for electronic consultation. Using 145 expert-crafted templates developed and routinely used by Stanford's eConsult team, we assess frontier models -- including o3, GPT-4o, Kimi K2, Claude 4 Sonnet, Llama 3 70B, and Gemini 2.5 Pro -- for their ability to produce clinically coherent, concise, and prioritized clinical qu
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
-
Omics In Lung
-
Early Detection of Lung Cancer: A Review of Innovative Milestones and Techniques
J Clin Med. 2025 Nov 3;14(21):7812. doi: 10.3390/jcm14217812.ABSTRACTLung cancer is the most frequently diagnosed cancer and the leading cause of cancer death worldwide. Early detection of lung cancer can lead to identification of the cancer at its initial treatable stages and improves survival. Low-dose CT scan (LDCT) is currently the gold standard for lung cancer screening in high-risk individuals. Despite the observed stage migration and consistently demonstrated disease-specific overall surv
Early Detection of Lung Cancer: A Review of Innovative Milestones and Techniques
J Clin Med. 2025 Nov 3;14(21):7812. doi: 10.3390/jcm14217812.
ABSTRACT
Lung cancer is the most frequently diagnosed cancer and the leading cause of cancer death worldwide. Early detection of lung cancer can lead to identification of the cancer at its initial treatable stages and improves survival. Low-dose CT scan (LDCT) is currently the gold standard for lung cancer screening in high-risk individuals. Despite the observed stage migration and consistently demonstrated disease-specific overall survival benefit, LDCT has inherent limitations, including false-positive results, radiation exposure, and low compliance. Recently, new techniques have been investigated for early detection of lung cancer. Several studies have shown that liquid biopsy biomarkers such as circulating cell-free DNA (cfDNA), microRNA molecules (miRNA), circulating tumor cells (CTCs), tumor-derived exosomes (TDEs), and tumor-educated platelets (TEPs), as well as volatile organic compounds (VOCs), have the power to distinguish lung cancer patients from healthy subjects, offering potential for minimally invasive and non-invasive means of early cancer detection. Furthermore, recent studies have shown that the integration of artificial intelligence (AI) with clinical, imaging, and laboratory data has provided significant advancements and can offer potential solutions to some challenges related to early detection of lung cancer. Adopting AI-based multimodality strategies, such as multi-omics liquid biopsy and/or VOCs' detection, with LDCT augmented by advanced AI, could revolutionize early lung cancer screening by improving accuracy, efficiency, and personalization, especially when combined with patient clinical data. However, challenges remain in validating, standardizing, and integrating these approaches into clinical practice. In this review, we described these innovative milestones and methods, as well as their advantages and limitations in screening and early diagnosis of lung cancer.
PMID:41227214 | PMC:PMC12609116 | DOI:10.3390/jcm14217812
-
Cell
-
Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution
Stereo-seq V2 facilitates single-cell-resolution spatial RNA mapping in FFPE samples through random primer capture, uncovering ncRNAs, host-pathogen transcriptome profiling, and spatial immune repertoires in situ.
Stereo-seq V2: Spatial mapping of total RNA on FFPE sections with high resolution
-
npj Digital Medicine
-
Equipping mathematical models for hospital dynamics using information theory
npj Digital Medicine, Published online: 12 November 2025; doi:10.1038/s41746-025-02013-2Equipping mathematical models for hospital dynamics using information theory
Equipping mathematical models for hospital dynamics using information theory
npj Digital Medicine, Published online: 12 November 2025; doi:10.1038/s41746-025-02013-2
Equipping mathematical models for hospital dynamics using information theory-
MIT Technology Review
-
Google is still aiming for its “moonshot” 2030 energy goals
Last week, we hosted EmTech MIT, MIT Technology Review’s annual flagship conference in Cambridge, Massachusetts. Over the course of three days of main-stage sessions, I learned about innovations in AI, biotech, and robotics. But as you might imagine, some of this climate reporter’s favorite moments came in the climate sessions. I was listening especially closely to my colleague James Temple’s discussion with Lucia Tian, head of advanced energy technologies at Google. They spoke about t
Google is still aiming for its “moonshot” 2030 energy goals
Last week, we hosted EmTech MIT, MIT Technology Review’s annual flagship conference in Cambridge, Massachusetts. Over the course of three days of main-stage sessions, I learned about innovations in AI, biotech, and robotics.
But as you might imagine, some of this climate reporter’s favorite moments came in the climate sessions. I was listening especially closely to my colleague James Temple’s discussion with Lucia Tian, head of advanced energy technologies at Google.
They spoke about the tech giant’s growing energy demand and what sort of technologies the company is looking to to help meet it. In case you weren’t able to join us, let’s dig into that session and consider how the company is thinking about energy in the face of AI’s rapid rise.
I’ve been closely following Google’s work in energy this year. Like the rest of the tech industry, the company is seeing ballooning electricity demand in its data centers. That could get in the way of a major goal that Google has been talking about for years.
See, back in 2020, the company announced an ambitious target: by 2030, it aimed to run on carbon-free energy 24-7. Basically, that means Google would purchase enough renewable energy on the grids where it operates to meet its entire electricity demand, and the purchases would match up so the electricity would have to be generated when the company was actually using energy. (For more on the nuances of Big Tech’s renewable-energy pledges, check out James’s piece from last year.)
Google’s is an ambitious goal, and on stage, Tian said that the company is still aiming for it but acknowledged that it’s looking tough with the rise of AI.
“It was always a moonshot,” she said. “It’s something very, very hard to achieve, and it’s only harder in the face of this growth. But our perspective is, if we don’t move in that direction, we’ll never get there.”
Google’s total electricity demand more than doubled from 2020 to 2024, according to its latest Environmental Report. As for that goal of 24-7 carbon-free energy? The company is basically treading water. While it was at 67% for its data centers in 2020, last year it came in at 66%.
Not going backwards is something of an accomplishment, given the rapid growth in electricity demand. But it still leaves the company some distance away from its finish line.
To close the gap, Google has been signing what feels like constant deals in the energy space. Two recent announcements that Tian talked about on stage were a project involving carbon capture and storage at a natural-gas plant in Illinois and plans to reopen a shuttered nuclear power plant in Iowa.
Let’s start with carbon capture. Google signed an agreement to purchase most of the electricity from a new natural-gas plant, which will capture and store about 90% of its carbon dioxide emissions.
That announcement was controversial, with critics arguing that carbon capture keeps fossil-fuel infrastructure online longer and still releases greenhouse gases and other pollutants into the atmosphere.
One question that James raised on stage: Why build a new natural-gas plant rather than add equipment to an already existing facility? Tacking on equipment to an operational plant would mean cutting emissions from the status quo, rather than adding entirely new fossil-fuel infrastructure.
The company did consider many existing plants, Tian said. But, as she put it, “Retrofits aren’t going to make sense everywhere.” Space can be limited at existing plants, for example, and many may not have the right geology to store carbon dioxide underground.
“We wanted to lead with a project that could prove this technology at scale,” Tian said. This site has an operational Class VI well, the type used for permanent sequestration, she added, and it also doesn’t require a big pipeline buildout.
Tian also touched on the company’s recent announcement that it’s collaborating with NextEra Energy to reopen Duane Arnold Energy Center, a nuclear power plant in Iowa. The company will purchase electricity from that plant, which is scheduled to reopen in 2029.
As I covered in a story earlier this year, Duane Arnold was basically the final option in the US for companies looking to reopen shuttered nuclear power plants. “Just a few years back, we were still closing down nuclear plants in this country,” Tian said on stage.
While each reopening will look a little different, Tian highlighted the groups working to restart the Palisades plant in Michigan, which was the first reopening to be announced, last spring. “They’re the real heroes of the story,” she said.
I’m always interested to get a peek behind the curtain at how Big Tech is thinking about energy. I’m skeptical but certainly interested to see how Google’s, and the rest of the industry’s, goals shape up over the next few years.
This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here.
-
npj Digital Medicine
-
Reply to “When do large language models cross the line: “reasoning” red teaming in healthcare”
npj Digital Medicine, Published online: 12 November 2025; doi:10.1038/s41746-025-02103-1Reply to “When do large language models cross the line: “reasoning” red teaming in healthcare”
Reply to “When do large language models cross the line: “reasoning” red teaming in healthcare”
npj Digital Medicine, Published online: 12 November 2025; doi:10.1038/s41746-025-02103-1
Reply to “When do large language models cross the line: “reasoning” red teaming in healthcare”-
cs.AI, q-bio.NC updates on arXiv.org
-
Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
arXiv:2511.07669v1 Announce Type: new Abstract: Current large language models (LLMs) excel in verifiable domains where outputs can be checked before action but prove less reliable for high-stakes strategic decisions with uncertain outcomes. This gap, driven by mutually reinforcing cognitive biases in both humans and artificial intelligence (AI) systems, threatens the defensibility of valuations and sustainability of investments in the sector. This report describes a framework emerging from sy
Making LLMs Reliable When It Matters Most: A Five-Layer Architecture for High-Stakes Decisions
-
cs.AI, q-bio.NC updates on arXiv.org
-
Clinical Uncertainty Impacts Machine Learning Evaluations
arXiv:2509.22242v2 Announce Type: replace Abstract: Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging benchmarks, accounting for the confidence in binary labels significantly impacts model rankings. We therefore argue that machine-learning evaluations should explicitly account for annotation uncertainty using prob
Clinical Uncertainty Impacts Machine Learning Evaluations
-
cs.AI, q-bio.NC updates on arXiv.org
-
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
arXiv:2511.05385v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agentic RAG employs autonomous, multi-round retrieval and reasoning to resolve queries. Although recent agentic RAG has improved via reinforcement learning, they often incur substantial token overhead from search and reasoning processes. This trade-off prioritizes accuracy over efficiency. To address this issue,