Nature, Published online: 31 December 2025; doi:10.1038/s41586-025-09819-wAn integrated systems engineering framework based on life-cycle inventories is used to quantify the global eco-footprint of wearable healthcare electronics and identify effective mitigation strategies.
An integrated systems engineering framework based on life-cycle inventories is used to quantify the global eco-footprint of wearable healthcare electronics and identify effective mitigation strategies.
Nature, Published online: 31 December 2025; doi:10.1038/d41586-025-03982-wA model quantifies the environmental footprint of wearable health-care electronics and identifies strategies to reduce their environmental toll.
Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.ABSTRACTThe global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective,
Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.
ABSTRACT
The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.
Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.ABSTRACTBackground and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was pe
Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.
ABSTRACT
Background and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was performed across multiple platforms using datasets derived from The Cancer Genome Atlas. Differential gene expression, DNA methylation, copy number variation, immune cell infiltration, drug sensitivity, and co-expression networks were systematically examined. Functional enrichment analyses of protein-protein and gene-gene interaction networks were conducted using STRING and GeneMANIA. Additionally, short interfering RNA-mediated knockdown and wound-healing assays were performed in HepG2 cells to evaluate BFSP1 function in vitro. Results: The results showed that BFSP1 mRNA expression was significantly upregulated in tissues from LIHC patients. Elevated BFSP1 levels were associated with poorer prognostic patterns, which were further supported by detailed clinicopathological subgroup analyses. Furthermore, BFSP1 expression was correlated with promoter hypomethylation and associated with patterns of tumor-infiltrating immune cells, including specific immune cell subtypes such as M1 and M2 macrophages. Integrative analyses revealed strong associations between BFSP1 and drug sensitivity, as well as a regulatory network encompassing genes involved in the cell cycle, DNA repair, and metabolic processes. Functional knockdown of BFSP1 significantly reduced HepG2 cell migration in vitro, as assessed by wound healing assay, with decreased wound closure at 24 h (11.0% vs. 16.5%) and 48 h (7.4% vs. 12.5%) compared with the control (p < 0.05, n = 6 biological replicates). Conclusions: In conclusion, these findings suggest that BFSP1 functions as a multifaceted prognostic biomarker and a potential therapeutic target for LIHC.
Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.ABSTRACTThe global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective,
Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.
ABSTRACT
The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.
Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.ABSTRACTBackground and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was pe
Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.
ABSTRACT
Background and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was performed across multiple platforms using datasets derived from The Cancer Genome Atlas. Differential gene expression, DNA methylation, copy number variation, immune cell infiltration, drug sensitivity, and co-expression networks were systematically examined. Functional enrichment analyses of protein-protein and gene-gene interaction networks were conducted using STRING and GeneMANIA. Additionally, short interfering RNA-mediated knockdown and wound-healing assays were performed in HepG2 cells to evaluate BFSP1 function in vitro. Results: The results showed that BFSP1 mRNA expression was significantly upregulated in tissues from LIHC patients. Elevated BFSP1 levels were associated with poorer prognostic patterns, which were further supported by detailed clinicopathological subgroup analyses. Furthermore, BFSP1 expression was correlated with promoter hypomethylation and associated with patterns of tumor-infiltrating immune cells, including specific immune cell subtypes such as M1 and M2 macrophages. Integrative analyses revealed strong associations between BFSP1 and drug sensitivity, as well as a regulatory network encompassing genes involved in the cell cycle, DNA repair, and metabolic processes. Functional knockdown of BFSP1 significantly reduced HepG2 cell migration in vitro, as assessed by wound healing assay, with decreased wound closure at 24 h (11.0% vs. 16.5%) and 48 h (7.4% vs. 12.5%) compared with the control (p < 0.05, n = 6 biological replicates). Conclusions: In conclusion, these findings suggest that BFSP1 functions as a multifaceted prognostic biomarker and a potential therapeutic target for LIHC.
Bioengineering (Basel). 2025 Dec 1;12(12):1315. doi: 10.3390/bioengineering12121315.ABSTRACTCancer drug screening is shifting from low-predictive, reductionist assays to human-relevant, data-integrated platforms. This review synthesizes preclinical strategies using a unified lens-Principle, Advantages, Limitations, and Clinical Application-to enable like-for-like comparison. We first appraise traditional two-dimensional (2D) monolayers and animal models, noting scalability and historical utility
Bioengineering (Basel). 2025 Dec 1;12(12):1315. doi: 10.3390/bioengineering12121315.
ABSTRACT
Cancer drug screening is shifting from low-predictive, reductionist assays to human-relevant, data-integrated platforms. This review synthesizes preclinical strategies using a unified lens-Principle, Advantages, Limitations, and Clinical Application-to enable like-for-like comparison. We first appraise traditional two-dimensional (2D) monolayers and animal models, noting scalability and historical utility alongside constrained translational fidelity. We then evaluate advanced systems-patient-derived organoids (PDOs), patient-derived xenografts (PDXs), and organ-on-a-chip-that better recapitulate architecture, microenvironmental cues, and pharmacodynamics (PD), yet face trade-offs in throughput, timelines, costs, and standardization. Functional genomic screens (CRISPR/RNAi) and large-scale pharmacogenomics are summarized as engines for mechanism-based target discovery and resistance mapping, while AI-enabled modeling supports response prediction, biomarker development, and rational combinations. Finally, we discuss trial designs (basket/umbrella), drug repurposing lessons, and regulatory momentum for new approach methodologies. Across platforms, we emphasize cross-model validation, dataset harmonization, and clinically anchored endpoints as prerequisites for real-world impact. We conclude with pragmatic guidance for matching screening modality to study goals, sample constraints, and decision timelines to accelerate precision oncology.
J Proteome Res. 2025 Dec 30. doi: 10.1021/acs.jproteome.5c00741. Online ahead of print.ABSTRACTHepatocellular carcinoma (HCC) ranks among the most common causes of cancer-related deaths globally. The high incidence of HCC is largely linked to chronic hepatitis virus infections, liver cirrhosis, and exposure to carcinogenic substances. Egypt has one of the world's highest burdens of HCC, with liver cirrhosis from chronic hepatitis C virus (HCV) infection as the primary risk factor. Malignant conv
J Proteome Res. 2025 Dec 30. doi: 10.1021/acs.jproteome.5c00741. Online ahead of print.
ABSTRACT
Hepatocellular carcinoma (HCC) ranks among the most common causes of cancer-related deaths globally. The high incidence of HCC is largely linked to chronic hepatitis virus infections, liver cirrhosis, and exposure to carcinogenic substances. Egypt has one of the world's highest burdens of HCC, with liver cirrhosis from chronic hepatitis C virus (HCV) infection as the primary risk factor. Malignant conversion of cirrhosis to HCC is often fatal in part because adequate biomarkers are not available for diagnosis of HCC in the early stage. Therefore, there is a critical need for more effective biomarkers to detect HCC at an early stage, when therapeutic intervention is more likely to be successful. Multiomics integration has emerged as a powerful strategy to uncover biomarkers and better understand the molecular underpinnings of complex diseases such as HCC. This study summarizes findings from multiple untargeted and targeted mass spectrometry-based analyses of proteins, N-linked glycans, and metabolites performed on blood samples from HCC cases and cirrhotic cohorts recruited in Egypt. Integrative analysis using machine learning methods is performed to identify a panel of multiomics features that differentiates HCC cases from the high-risk population of cirrhotic patients with liver cirrhosis.
Weβre in the midst of a global mental-Βhealth crisis. More than a billion people worldwide suffer from a mental-health condition, according to the World Health Organization. The prevalence of anxiety and depression is growing in many demographics, particularly young people, and suicide is claiming hundreds of thousands of lives globally each year.
Given the clear demand for accessible and affordable mental-health services, itβs no wonder that people have looked to artificial intelligence for
Weβre in the midst of a global mental-Βhealth crisis. More than a billion people worldwide suffer from a mental-health condition, according to the World Health Organization. The prevalence of anxiety and depression is growing in many demographics, particularly young people, and suicide is claiming hundreds of thousands of lives globally each year.
Given the clear demand for accessible and affordable mental-health services, itβs no wonder that people have looked to artificial intelligence for possible relief. Millions are already actively seeking therapy from popular chatbots like OpenAIβs ChatGPT and Anthropicβs Claude, or from specialized psychology apps like Wysa and Woebot. On a broader scale, researchers are exploring AIβs potential to monitor and collect behavioral and biometric observations using wearables and smart devices, analyze vast volumes of clinical data for new insights, and assist human mental-health professionals to help prevent burnout.Β
But so far this largely uncontrolled experiment has produced mixed results. Many people have found solace in chatbots based on large language models (LLMs), and some experts see promise in them as therapists, but other users have been sent into delusional spirals by AIβs hallucinatory whims and breathless sycophancy. Most tragically, multiple families have alleged that chatbots contributed to the suicides of their loved ones, sparking lawsuits against companies responsible for these tools. In October, OpenAI CEO Sam Altman revealed in a blog post that 0.15% of ChatGPT users βhave conversations that include explicit indicators of potential suicidal planning or intent.β Thatβs roughly a million people sharing suicidal ideations with just one of these software systems every week.
The real-world consequences of AI therapy came to a head in unexpected ways in 2025 as we waded through a critical mass of stories about human-chatbot relationships, the flimsiness of guardrails on many LLMs, and the risks of sharing profoundly personal information with products made by corporations that have economic incentives to harvest and monetize such sensitive data.Β
Several authors anticipated this inflection point. Their timely books are a reminder that while the present feels like a blur of breakthroughs, scandals, and confusion, this disorienting time is rooted in deeper histories of care, technology, and trust.Β
LLMs have often been described as βblack boxesβ because nobody knows exactly how they produce their results. The inner workings that guide their outputs are opaque because their algorithms are so complex and their training data is so vast. In mental-health circles, people often describe the human brain as a βblack box,β for analogous reasons. Psychology, psychiatry, and related fields must grapple with the impossibility of seeing clearly inside someone elseβs head, let alone pinpointing the exact causes of their distress.Β
These two types of black boxes are now interacting with each other, creating unpredictable feedback loops that may further impede clarity about the origins of peopleβs mental-Βhealth struggles and the solutions that may be possible. Anxiety about these developments has much to do with the explosive recent advances in AI, but it also revives decades-old warnings from pioneers such as the MIT computer scientist Joseph Weizenbaum, who argued against computerized therapy as early as the 1960s.Β Β
Dr. Bot: Why Doctors Can Fail Usβ and How AI Could Save Lives Charlotte Blease
YALE UNIVERSITY PRESS, 2025
Charlotte Blease, a philosopher of medicine, makes the optimistβs case in Dr. Bot: Why Doctors Can Fail Usβand How AI Could Save Lives. Her book broadly explores the possible positive impacts of AI in a range of medical fields. While she remains clear-eyed about the risks, warning that readers who are expecting βa gushing love letter to technologyβ will be disappointed, she suggests that these models can help relieve patient suffering and medical burnout alike.
βHealth systems are crumbling under patient pressure,β Blease writes. βGreater burdens on fewer doctors create the perfect petri dish for errors,β and βwith palpable shortages of doctors and increasing waiting times for patients, many of us are profoundly frustrated.β
Blease believes that AI can not only ease medical professionalsβ massive workloads but also relieve the tensions that have always existed between some patients and their caregivers. For example, people often donβt seek needed care because they are intimidated or fear judgment from medical professionals; this is especially true if they have mental-health challenges. AI could allow more people to share their concerns, she argues.Β
But sheβs aware that these putative upsides need to be weighed against major drawbacks. For instance, AI therapists can provide inconsistent and even dangerous responses to human users, according to a 2025 study, and they also raise privacy concerns, given that AI companies are currently not bound by the same confidentiality and HIPAA standards as licensed therapists.Β
While Blease is an expert in this field, her motivation for writing the book is also personal: She has two siblings with an incurable form of muscular dystrophy, one of whom waited decades for a diagnosis. During the writing of her book, she also lost her partner to cancer and her father to dementia within a devastating six-month period.Β βI witnessed first-hand the sheer brilliance of doctors and the kindness of health professionals,β she writes. βBut I also observed how things can go wrong with care.β
The Silicon Shrink: How Artificial Intelligence Made the World an Asylum Daniel Oberhaus
MIT PRESS, 2025
A similar tension animates Daniel Oberhausβs engrossing book The Silicon Shrink: How Artificial Intelligence Made the World an Asylum. Oberhaus starts from a point of tragedy: the loss of his younger sister to suicide. As Oberhaus carried out the βdistinctly twenty-first-century mourning processβ of sifting through her digital remains, he wondered if technology could have eased the burden of the psychiatric problems that had plagued her since childhood.
βIt seemed possible that all of this personal data might have held important clues that her mental health providers could have used to provide more effective treatment,β he writes. βWhat if algorithms running on my sisterβs smartphone or laptop had used that data to understand when she was in distress? Could it have led to a timely intervention that saved her life? Would she have wanted that even if it did?β
This concept of digital phenotypingβin which a personβs digital behavior could be mined for clues about distress or illnessβseems elegant in theory. But it may also become problematic if integrated into the field of psychiatric artificial intelligence (PAI), which extends well beyond chatbot therapy.
Oberhaus emphasizes that digital clues could actually exacerbate the existing challenges of modern psychiatry, a discipline that remains fundamentally uncertain about the underlying causes of mental illnesses and disorders. The advent of PAI, he says, is βthe logical equivalent of grafting physics onto astrology.β In other words, the data generated by digital phenotyping is as precise as physical measurements of planetary positions, but it is then integrated into a broader frameworkβin this case, psychiatryβthat, like astrology, is based on unreliable assumptions.Β Β
Oberhaus, who uses the phrase βswipe psychiatryβ to describe the outsourcing of clinical decisions based on behavioral data to LLMs, thinks that this approach cannot escape the fundamental issues facing psychiatry. In fact, it could worsen the problem by causing the skills and judgment of human therapists to atrophy as they grow more dependent on AI systems.Β
He also uses the asylums of the pastβin which institutionalized patients lost their right to freedom, privacy, dignity, and agency over their livesβas a touchstone for a more insidious digital captivity that may spring from PAI. LLM users are already sacrificing privacy by telling chatbots sensitive personal information that companies then mine and monetize, contributing to a new surveillance economy. Freedom and dignity are at stake when complex inner lives are transformed into data streams tailored for AI analysis.Β
AI therapists could flatten humanity into patterns of prediction, and so sacrifice the intimate, individualized care that is expected of traditional human therapists. βThe logic of PAI leads to a future where we may all find ourselves patients in an algorithmic asylum administered by digital wardens,β Oberhaus writes. βIn the algorithmic asylum there is no need for bars on the window or white padded rooms because there is no possibility of escape. The asylum is already everywhereβin your homes and offices, schools and hospitals, courtrooms and barracks. Wherever thereβs an internet connection, the asylum is waiting.β
Chatbot Therapy: A Critical Analysis of AI Mental Health Treatment Eoin Fullam
ROUTLEDGE, 2025
Eoin Fullam, a researcher who studies the intersection of technology and mental health, echoes some of the same concerns in Chatbot Therapy: A Critical Analysis of AI Mental Health Treatment. A heady academic primer, the book analyzes the assumptions underlying the automated treatments offered by AI chatbots and the way capitalist incentives could corrupt these kinds of tools.Β Β
Fullam observes that the capitalist mentality behind new technologies βoften leads to questionable, illegitimate, and illegal business practices in which the customersβ interests are secondary to strategies of market dominance.β
That doesnβt mean that therapy-bot makers βwill inevitably conduct nefarious activities contrary to the usersβ interests in the pursuit of market dominance,β Fullam writes.Β
But he notes that the success of AI therapy depends on the inseparable impulses to make money and to heal people. In this logic, exploitation and therapy feed each other: Every digital therapy session generates data, and that data fuels the system that profits as unpaid users seek care. The more effective the therapy seems, the more the cycle entrenches itself, making it harder to distinguish between care and commodification. βThe more the users benefit from the app in terms of its therapeutic or any other mental health intervention,β he writes, βthe more they undergo exploitation.βΒ
This sense of an economic and psychological ouroborosβthe snake that eats its own tailβserves as a central metaphor in Sike, the debut novel from Fred Lunzer, an author with a research background in AI.Β
Described as a βstory of boy meets girl meets AI psychotherapist,β Sike follows Adrian, a young Londoner who makes a living ghostwriting rap lyrics, in his romance with Maquie, a business professional with a knack for spotting lucrative technologies in the beta phase.Β
Sike Fred Lunzer
CELADON BOOKS, 2025
The title refers to a splashy commercial AI therapist called Sike, uploaded into smart glasses, that Adrian uses to interrogate his myriad anxieties. βWhen I signed up to Sike, we set up my dashboard, a wide black panel like an airplaneβs cockpit that showed my daily βvitals,ββ Adrian narrates. βSike can analyze the way you walk, the way you make eye contact, the stuff you talk about, the stuff you wear, how often you piss, shit, laugh, cry, kiss, lie, whine, and cough.β
In other words, Sike is the ultimate digital phenotyper, constantly and exhaustively analyzing everything in a userβs daily experiences. In a twist, Lunzer chooses to make Sike a luxury product, available only to subscribers who can foot the price tag of Β£2,000 per month.Β
Flush with cash from his contributions to a hit song, Adrian comes to rely on Sike as a trusted mediator between his inner and outer worlds. The novel explores the impacts of the app on the wellness of the well-off, following rich people who voluntarily commit themselves to a boutique version of the digital asylum described by Oberhaus.
The only real sense of danger in Sike involves a Japanese torture egg (donβt ask). The novel strangely sidesteps the broader dystopian ripples of its subject matter in favor of drunken conversations at fancy restaurants and elite dinner parties.Β
The sudden ascent of the AI therapist seems startlingly futuristic, as if it should be unfolding in some later time when the streets scrub themselves and we travel the world through pneumatic tubes.
Sikeβs creator is simply βa great guyβ in Adrianβs estimation, despite his techno-messianic vision of training the app to soothe the ills of entire nations. It always seems as if a shoe is meant to drop, but in the end, it never does, leaving the reader with a sense of non-resolution.
While Sike is set in the present day, something about the sudden ascent of the AI therapistβΒin real life as well as in fictionβseems startlingly futuristic, as if it should be unfolding in some later time when the streets scrub themselves and we travel the world through pneumatic tubes. But this convergence of mental health and artificial intelligence has been in the making for more than half a century. The beloved astronomer Carl Sagan, for example, once imagined a βnetwork of computer psychotherapeutic terminals, something like arrays of large telephone boothsβ that could address the growing demand for mental-health services.
Oberhaus notes that one of the first incarnations of a trainable neural network, known as the Perceptron, was devised not by a mathematician but by a psychologist named Frank Rosenblatt, at the Cornell Aeronautical Laboratory in 1958. The potential utility of AI in mental health was widely recognized by the 1960s, inspiring early computerized psychotherapists such as the DOCTOR script that ran on the ELIZA chatbot developed by Joseph Weizenbaum, who shows up in all three of the nonfiction books in this article.
Weizenbaum, who died in 2008, was profoundly concerned about the possibility of computerized therapy. βComputers can make psychiatric judgments,β he wrote in his 1976 book Computer Power and Human Reason. βThey can flip coins in much more sophisticated ways than can the most patient human being. The point is that they ought not to be given such tasks. They may even be able to arrive at βcorrectβ decisions in some casesβbut always and necessarily on bases no human being should be willing to accept.β
Itβs a caution worth keeping in mind. As AI therapists arrive at scale, weβre seeing them play out a familiar dynamic: Tools designed with superficially good intentions are enmeshed with systems that can exploit, surveil, and reshape human behavior. In a frenzied attempt to unlock new opportunities for patients in dire need of mental-health support, we may be locking other doors behind them.
npj Digital Medicine, Published online: 30 December 2025; doi:10.1038/s41746-025-02284-9PIC-SURE: an open-source platform for integrating clinical and genomic data
arXiv:2512.22199v1 Announce Type: new
Abstract: Retrieval-Augmented Generation RAG systems enhance large language models by grounding responses in external knowledge bases, but conventional RAG architectures operate with static corpora that cannot evolve from user interactions. We introduce Bidirectional RAG, a novel RAG architecture that enables safe corpus expansion through validated write back of high quality generated responses. Our system employs a multi stage acceptance layer combining gr
arXiv:2512.22199v1 Announce Type: new
Abstract: Retrieval-Augmented Generation RAG systems enhance large language models by grounding responses in external knowledge bases, but conventional RAG architectures operate with static corpora that cannot evolve from user interactions. We introduce Bidirectional RAG, a novel RAG architecture that enables safe corpus expansion through validated write back of high quality generated responses. Our system employs a multi stage acceptance layer combining grounding verification (NLI based entailment, attribution checking, and novelty detection to prevent hallucination pollution while enabling knowledge accumulation. Across four datasets Natural Questions, TriviaQA, HotpotQA, Stack Overflow with three random seeds 12 experiments per system, Bidirectional RAG achieves 40.58% average coverage nearly doubling Standard RAG 20.33% while adding 72% fewer documents than naive write back 140 vs 500. Our work demonstrates that self improving RAG is feasible and safe when governed by rigorous validation, offering a practical path toward RAG systems that learn from deployment.
arXiv:2512.22334v1 Announce Type: new
Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Gene
arXiv:2512.22334v1 Announce Type: new
Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.
arXiv:2512.22470v1 Announce Type: new
Abstract: The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grai
arXiv:2512.22470v1 Announce Type: new
Abstract: The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grained assessment of manipulative content in LLM outputs across seven harm categories: Legal/Power, Psychological, Emotional, Physical, Autonomy, Economic, and Societal Harm. Our framework implements a four-layer analytical pipeline comprising Multi-Granular Detection (MGD), Multi-Scale Intent Analysis (MSIAN), Threat Harmonization Protocol (THP), and Deep Contextual Risk Alignment (DCRA). The dataset contains 401 meticulously curated examples with instruction-response pairs and expert annotations. Through evaluation of state-of-the-art models including GPT-4, Claude 3.5, and LLaMA-3-70B, we observe significant performance disparities (65.2\%--89.7\%) and consistent weaknesses in detecting autonomy-undermining patterns. DarkPatterns-LLM establishes the first standardized, multi-dimensional benchmark for manipulation detection in LLMs, offering actionable diagnostics toward more trustworthy AI systems.
arXiv:2512.22568v1 Announce Type: new
Abstract: The phenomenal advances in large language models (LLMs) and other foundation models over the past few years have been based on optimizing large-scale transformer models on the surprisingly simple objective of minimizing next-token prediction loss, a form of predictive coding that is also the backbone of an increasingly popular model of brain function in neuroscience and cognitive science. However, current foundation models ignore three other impor
arXiv:2512.22568v1 Announce Type: new
Abstract: The phenomenal advances in large language models (LLMs) and other foundation models over the past few years have been based on optimizing large-scale transformer models on the surprisingly simple objective of minimizing next-token prediction loss, a form of predictive coding that is also the backbone of an increasingly popular model of brain function in neuroscience and cognitive science. However, current foundation models ignore three other important components of state-of-the-art predictive coding models: tight integration of actions with generative models, hierarchical compositional structure, and episodic memory. We propose that to achieve safe, interpretable, energy-efficient, and human-like AI, foundation models should integrate actions, at multiple scales of abstraction, with a compositional generative architecture and episodic memory. We present recent evidence from neuroscience and cognitive science on the importance of each of these components. We describe how the addition of these missing components to foundation models could help address some of their current deficiencies: hallucinations and superficial understanding of concepts due to lack of grounding, a missing sense of agency/responsibility due to lack of control, threats to safety and trustworthiness due to lack of interpretability, and energy inefficiency. We compare our proposal to current trends, such as adding chain-of-thought (CoT) reasoning and retrieval-augmented generation (RAG) to foundation models, and discuss new ways of augmenting these models with brain-inspired components. We conclude by arguing that a rekindling of the historically fruitful exchange of ideas between brain science and AI will help pave the way towards safe and interpretable human-centered AI.
arXiv:2512.23184v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choic
arXiv:2512.23184v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.
arXiv:2512.23508v1 Announce Type: new
Abstract: How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down w
arXiv:2512.23508v1 Announce Type: new
Abstract: How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.
arXiv:2512.22149v1 Announce Type: cross
Abstract: Multi-agent systems powered by large language models have emerged as a promising paradigm for solving complex reasoning tasks through collaborative intelligence. However, efficiently deploying these systems on serverless GPU platforms presents significant resource allocation challenges due to heterogeneous agent workloads, varying computational demands, and the need for cost-effective scaling. This paper presents an adaptive GPU resource allocat
arXiv:2512.22149v1 Announce Type: cross
Abstract: Multi-agent systems powered by large language models have emerged as a promising paradigm for solving complex reasoning tasks through collaborative intelligence. However, efficiently deploying these systems on serverless GPU platforms presents significant resource allocation challenges due to heterogeneous agent workloads, varying computational demands, and the need for cost-effective scaling. This paper presents an adaptive GPU resource allocation framework that achieves 85\% latency reduction compared to round-robin scheduling while maintaining comparable throughput to static allocation, using an $O(N)$ complexity algorithm for real-time adaptation. Our approach dynamically allocates GPU resources based on workload characteristics, agent priorities, and minimum resource requirements, enabling efficient utilization while maintaining quality of service. The framework addresses three key challenges: (1) heterogeneous computational demands across lightweight coordinators and heavyweight specialists, (2) dynamic workload fluctuations requiring millisecond-scale reallocation, and (3) capacity constraints in serverless environments. Through comprehensive simulations modeling realistic multi-agent workflows with four heterogeneous agents, we demonstrate that adaptive allocation outperforms static equal and round-robin strategies across latency, cost, and GPU utilization metrics. The framework provides a practical solution for deploying cost-efficient multi-agent AI systems on serverless GPU infrastructure.
arXiv:2512.22181v1 Announce Type: cross
Abstract: Artificial intelligence (AI) is transforming cancer diagnosis and treatment. The intricate nature of this disease necessitates the collaboration of diverse stakeholders with varied expertise to ensure the effectiveness of cancer research. Despite its importance, forming effective interdisciplinary research teams remains challenging. Understanding and predicting collaboration patterns can help researchers, organizations, and policymakers optimize
arXiv:2512.22181v1 Announce Type: cross
Abstract: Artificial intelligence (AI) is transforming cancer diagnosis and treatment. The intricate nature of this disease necessitates the collaboration of diverse stakeholders with varied expertise to ensure the effectiveness of cancer research. Despite its importance, forming effective interdisciplinary research teams remains challenging. Understanding and predicting collaboration patterns can help researchers, organizations, and policymakers optimize resources and foster impactful research. We examined co-authorship networks as a proxy for collaboration within AI-driven cancer research. Using 7,738 publications (2000-2017) from Scopus, we constructed 36 overlapping co-authorship networks representing new, persistent, and discontinued collaborations. We engineered both attribute-based and structure-based features and built four machine learning classifiers. Model interpretability was performed using Shapley Additive Explanations (SHAP). Random forest achieved the highest recall for all three types of examined collaborations. The discipline similarity score emerged as a crucial factor, positively affecting new and persistent patterns while negatively impacting discontinued collaborations. Additionally, high productivity and seniority were positively associated with discontinued links. Our findings can guide the formation of effective research teams, enhance interdisciplinary cooperation, and inform strategic policy decisions.
arXiv:2512.22182v1 Announce Type: cross
Abstract: The rapid evolution of Artificial intelligence in healthcare has opened avenues for enhancing various processes, including medical billing and transcription. This paper introduces an innovative approach by integrating AI with Locally Linear Embedding (LLE) to revolutionize the handling of high-dimensional medical data. This AI-enhanced LLE model is specifically tailored to improve the accuracy and efficiency of medical billing systems and transc
arXiv:2512.22182v1 Announce Type: cross
Abstract: The rapid evolution of Artificial intelligence in healthcare has opened avenues for enhancing various processes, including medical billing and transcription. This paper introduces an innovative approach by integrating AI with Locally Linear Embedding (LLE) to revolutionize the handling of high-dimensional medical data. This AI-enhanced LLE model is specifically tailored to improve the accuracy and efficiency of medical billing systems and transcription services. By automating these processes, the model aims to reduce human error and streamline operations, thereby facilitating faster and more accurate patient care documentation and financial transactions. This paper provides a comprehensive mathematical model of AI-enhanced LLE, demonstrating its application in real-world healthcare scenarios through a series of experiments. The results indicate a significant improvement in data processing accuracy and operational efficiency. This study not only underscores the potential of AI-enhanced LLE in medical data analysis but also sets a foundation for future research into broader healthcare applications.
arXiv:2512.22242v1 Announce Type: cross
Abstract: Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance acro
arXiv:2512.22242v1 Announce Type: cross
Abstract: Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p