❌

Normal view

  • ✇MIT Technology Review
  • AI models are using material from retracted scientific papers Ananya
    Some AI chatbots rely on flawed research from retracted scientific papers to answer questions, according to recent studies. The findings, confirmed by MIT Technology Review, raise questions about how reliable AI tools are at evaluating scientific research and could complicate efforts by countries and industries seeking to invest in AI tools for scientists. AI search tools and chatbots are already known to fabricate links and references. But answers based on the material from actual papers can
     

AI models are using material from retracted scientific papers

By: Ananya
23 September 2025 at 17:00

Some AI chatbots rely on flawed research from retracted scientific papers to answer questions, according to recent studies. The findings, confirmed by MIT Technology Review, raise questions about how reliable AI tools are at evaluating scientific research and could complicate efforts by countries and industries seeking to invest in AI tools for scientists.

AI search tools and chatbots are already known to fabricate links and references. But answers based on the material from actual papers can mislead as well if those papers have been retracted. The chatbot is “using a real paper, real material, to tell you something,” says Weikuan Gu, a medical researcher at the University of Tennessee in Memphis and an author of one of the recent studies. But, he says, if people only look at the content of the answer and do not click through to the paper and see that it’s been retracted, that’s really a problem. 

Gu and his team asked OpenAI’s ChatGPT, running on the GPT-4o model, questions based on information from 21 retracted papers about medical imaging. The chatbot’s answers referenced retracted papers in five cases but advised caution in only three. While it cited non-retracted papers for other questions, the authors note that it may not have recognized the retraction status of the articles. In a study from August, a different group of researchers used ChatGPT-4o mini to evaluate the quality of 217 retracted and low-quality papers from different scientific fields; they found that none of the chatbot’s responses mentioned retractions or other concerns. (No similar studies have been released on GPT-5, which came out in August.)

The public uses AI chatbots to ask for medical advice and diagnose health conditions. Students and scientists increasingly use science-focused AI tools to review existing scientific literature and summarize papers. That kind of usage is likely to increase. The US National Science Foundation, for instance, invested $75 million in building AI models for science research this August.

“If [a tool is] facing the general public, then using retraction as a kind of quality indicator is very important,” says Yuanxi Fu, an information science researcher at the University of Illinois Urbana-Champaign. There’s “kind of an agreement that retracted papers have been struck off the record of science,” she says, “and the people who are outside of science—they should be warned that these are retracted papers.” OpenAI did not provide a response to a request for comment about the paper results.

The problem is not limited to ChatGPT. In June, MIT Technology Review tested AI tools specifically advertised for research work, such as Elicit, Ai2 ScholarQA (now part of the Allen Institute for Artificial Intelligence’s Asta tool), Perplexity, and Consensus, using questions based on the 21 retracted papers in Gu’s study. Elicit referenced five of the retracted papers in its answers, while Ai2 ScholarQA referenced 17, Perplexity 11, and Consensus 18—all without noting the retractions.

Some companies have since made moves to correct the issue. “Until recently, we didn’t have great retraction data in our search engine,” says Christian Salem, cofounder of Consensus. His company has now started using retraction data from a combination of sources, including publishers and data aggregators, independent web crawling, and Retraction Watch, which manually curates and maintains a database of retractions. In a test of the same papers in August, Consensus cited only five retracted papers. 

Elicit told MIT Technology Review that it removes retracted papers flagged by the scholarly research catalogue OpenAlex from its database and is “still working on aggregating sources of retractions.” Ai2 told us that its tool does not automatically detect or remove retracted papers currently. Perplexity said that it “[does] not ever claim to be 100% accurate.” 

However, relying on retraction databases may not be enough. Ivan Oransky, the cofounder of Retraction Watch, is careful not to describe it as a comprehensive database, saying that creating one would require more resources than anyone has: “The reason it’s resource intensive is because someone has to do it all by hand if you want it to be accurate.”

Further complicating the matter is that publishers don’t share a uniform approach to retraction notices. “Where things are retracted, they can be marked as such in very different ways,” says Caitlin Bakker from University of Regina, Canada, an expert in research and discovery tools. “Correction,” “expression of concern,” “erratum,” and “retracted” are among some labels publishers may add to research papers—and these labels can be added for many reasons, including concerns about the content, methodology, and data or the presence of conflicts of interest. 

Some researchers distribute their papers on preprint servers, paper repositories, and other websites, causing copies to be scattered around the web. Moreover, the data used to train AI models may not be up to date. If a paper is retracted after the model’s training cutoff date, its responses might not instantaneously reflect what’s going on, says Fu. Most academic search engines don’t do a real-time check against retraction data, so you are at the mercy of how accurate their corpus is, says Aaron Tay, a librarian at Singapore Management University.

Oransky and other experts advocate making more context available for models to use when creating a response. This could mean publishing information that already exists, like peer reviews commissioned by journals and critiques from the review site PubPeer, alongside the published paper.  

Many publishers, such as Nature and the BMJ, publish retraction notices as separate articles linked to the paper, outside paywalls. Fu says companies need to effectively make use of such information, as well as any news articles in a model’s training data that mention a paper’s retraction. 

The users and creators of AI tools need to do their due diligence. “We are at the very, very early stages, and essentially you have to be skeptical,” says Tay.

Ananya is a freelance science and technology journalist based in Bengaluru, India.

Deciphering the Heterogeneity of Pancreatic Cancer: DNA Methylation-Based Cell Type Deconvolution Unveils Distinct Subgroups and Immune Landscapes

Epigenomes. 2025 Sep 5;9(3):34. doi: 10.3390/epigenomes9030034.

ABSTRACT

Background: Pancreatic ductal adenocarcinoma (PDAC) is a highly heterogeneous malignancy, characterized by low tumor cellularity, a dense stromal response, and intricate cellular and molecular interactions within the tumor microenvironment (TME). Although bulk omics technologies have enhanced our understanding of the molecular landscape of PDAC, the specific contributions of non-malignant immune and stromal components to tumor progression and therapeutic response remain poorly understood. Methods: We explored genome-wide DNA methylation and transcriptomic data from the Cancer Genome Atlas Pancreatic Adenocarcinoma cohort (TCGA-PAAD) to profile the immune composition of the TME and uncover gene co-expression networks. Bioinformatic analyses included DNA methylation profiling followed by hierarchical deconvolution, epigenetic age estimation, and a weighted gene co-expression network analysis (WGCNA). Results: The unsupervised clustering of methylation profiles identified two major tumor groups, with Group 2 (n = 98) exhibiting higher tumor purity and a greater frequency of KRAS mutations compared to Group 1 (n = 87) (p < 0.0001). The hierarchical deconvolution of DNA methylation data revealed three distinct TME subtypes, termed hypo-inflamed (immune-deserted), myeloid-enriched, and lymphoid-enriched (notably T-cell predominant). These immune clusters were further supported by co-expression modules identified via WGCNA, which were enriched in immune regulatory and signaling pathways. Conclusions: This integrative epigenomic-transcriptomic analysis offers a robust framework for stratifying PDAC patients based on the tumor immune microenvironment (TIME), providing valuable insights for biomarker discovery and the development of precision immunotherapies.

PMID:40981070 | PMC:PMC12452622 | DOI:10.3390/epigenomes9030034

Cancer in a drop: Liquid biopsy highlights from the American Society of Clinical Oncology (ASCO) 2025 annual congress

J Liq Biopsy. 2025 Aug 6;9:100320. doi: 10.1016/j.jlb.2025.100320. eCollection 2025 Sep.

ABSTRACT

Over the past decade, liquid biopsy has progressively expanded its role in oncology, supported by mounting evidence demonstrating an increasing number of clinical applications. At the 2025 American Society of Clinical Oncology (ASCO) Annual Meeting, liquid biopsy emerged as a central theme across multiple sessions, with more than 700 abstracts, investigating the clinical utility of liquid biopsy across a wide range of tumor types and disease stages. Applications presented included cancer screening, minimal residual disease (MRD) detection, management of metastatic disease, and potential use for matching patients to clinical trials. This editorial, authored on the behalf of the Young Committee of the International Society of Liquid Biopsy (ISLB) highlights the result of selected studies, grouped by tumor type.

PMID:40980343 | PMC:PMC12447415 | DOI:10.1016/j.jlb.2025.100320

Cell-free DNA fragmentomics: a universal framework for early cancer detection and monitoring

22 September 2025 at 18:00

Am J Clin Exp Immunol. 2025 Aug 15;14(4):237-240. doi: 10.62347/EBRY4326. eCollection 2025.

ABSTRACT

Cell-free DNA (cfDNA) fragmentomics has emerged as a powerful and noninvasive approach for cancer detection, characterization, and monitoring. By analyzing genome-wide fragmentation patterns - including fragment length distributions, end motifs, nucleosome footprints, and copy number variations - cfDNA fragmentomics provides high-resolution insights into tumor-specific biological signals even at low tumor burden. This technology offers advantages over conventional mutation-based assays by capturing aggregate structural and epigenomic alterations without requiring prior knowledge of driver mutations. In non-small cell lung cancer (NSCLC), cfDNA fragmentomics enables early detection, discrimination of malignant pulmonary nodules, and post-surgical monitoring of minimal residual disease. Recent studies have demonstrated that fragmentomic risk scores can accurately stratify recurrence risk and improve prognostic sensitivity beyond traditional genomic assays. In hepatocellular carcinoma (HCC), integration of fragment size selection, CNV profiling, and end-motif analysis has led to high-performing models for early diagnosis, particularly in high-risk populations. Moreover, cfDNA fragmentomics has proven effective in detecting malignant transformation in patients with neurofibromatosis-associated peripheral nerve sheath tumors, distinguishing benign from premalignant or malignant lesions with high precision. Expanding beyond these major cancers, fragmentomic approaches have demonstrated diagnostic potential in gastric, urological, hematologic, and pediatric malignancies. Notably, the DELFI-TF (DNA Evaluation of Fragments for early Interception-Tumor Fraction) framework has shown prognostic relevance by correlating pre-treatment cfDNA features with survival outcomes in colorectal and lung cancer patients, outperforming conventional imaging. All of these results highlight the translational importance of cfDNA fragmentomics as a cutting-edge precision oncology tool. Its continued integration into clinical workflows may redefine early cancer detection, facilitate subtype-specific interventions, and enable real-time, individualized treatment monitoring.

PMID:40977920 | PMC:PMC12444407 | DOI:10.62347/EBRY4326

Circulating tumor DNA in patients with cancer: insights from clinical laboratory

Adv Lab Med. 2025 Jun 16;6(3):259-276. doi: 10.1515/almed-2025-0010. eCollection 2025 Sep.

ABSTRACT

Blood-based circulating tumor DNA (ctDNA) analysis has emerged as a highly relevant non-invasive method for molecular profiling of solid tumors, offering valuable information about the genetic landscape of cancer. Somatic mutation analysis of ctDNA is now used clinically to guide targeted therapies for advanced cancers. Recent advancements have also revealed its potential in early detection, prognosis, minimal residual disease assessment, and prediction/monitoring of therapeutic response. In recent years, significant progress has been made with the development of various PCR and NGS-based methods designed for assessing gene variants in ctDNA of patients with cancer. However, despite the transformative possibilities that ctDNA analysis presents, challenges persist. Standardization of preanalytical and analytical protocols, assay sensitivity, and the interpretation of results remain critical hurdles that need to be addressed for the widespread clinical implementation of ctDNA testing. In addition to somatic mutations, emerging studies on DNA methylation (epigenomics) and fragment size patterns (fragmentomics) in several types of biological fluids are yielding promising results as non-invasive biomarkers for effective cancer management. This review addresses the clinical applications of somatic gene variants in ctDNA, emphasizes their potential as cancer biomarkers, and highlights essential factors for successful implementation in clinical laboratories and cancer management.

PMID:40977813 | PMC:PMC12446922 | DOI:10.1515/almed-2025-0010

A statistical physics approach to integrating multi-omics data for disease-module detection

Cell Rep Methods. 2025 Sep 19:101183. doi: 10.1016/j.crmeth.2025.101183. Online ahead of print.

ABSTRACT

Genes associated with the same disease frequently engage in mutual biological interactions, e.g., perturbation within a specific neighborhood in the molecular interactome, often referred to as the disease module. This has propelled the advancement of network-based approaches toward elucidating the molecular bases of human diseases. Although many computational methods have been developed to integrate the molecular interactome and omics profiles to extract such context-dependent disease modules, approaches that leverage multi-omics for disease-module detection are still lacking. Here, we developed a statistical physics approach based on the random-field O(n) model (RFOnM) to fill this gap. We applied the RFOnM approach to integrate gene-expression data and genome-wide association studies or mRNA data and DNA methylation for several complex diseases with the human interactome. We found that the RFOnM approach outperforms existing single omics methods in most of the complex diseases considered in this study.

PMID:40975055 | DOI:10.1016/j.crmeth.2025.101183

Implementation of a Virtual Hospital in the Home Service for Patients With COVID-19 in Queensland, Australia: Mixed Methods Evaluation Using the RE-AIM Framework

Background: Hospital in the home (HITH) provides home-based care as an alternative to traditional hospitalization. In response to the COVID-19 Omicron wave, a public hospital in the rural Western portion of Southeast Queensland implemented a virtual HITH service to support adults, maternity patients, and children with moderate COVID-19 symptoms and additional health concerns. Although the pandemic accelerated the uptake of virtual care within HITH models, existing literature has focused on clinical outcomes, with limited evidence on key implementation outcomes. Objective: Using the RE-AIM (reach, effectiveness, adoption, implementation, and maintenance) framework, this study evaluated the implementation of the virtual COVID-19 HITH service and identified factors influencing its implementation, to inform ongoing service development and support potential scaling of this model of care. Methods: The RE-AIM implementation science framework was selected to guide the evaluation, capturing both clinical and contextual dimensions of implementation at both individual and organizational levels. Quantitative data on service usage and costs were retrospectively extracted from electronic medical records and finance records, while patient experience data were drawn from patient-reported experience measures surveys. Qualitative data were collected through one-on-one interviews with patients and staff. All data sources were analyzed separately and then triangulated within the RE-AIM framework to understand what occurred, how, and why. Results: The service admitted 3192 patients, most of whom were female (2027/3192, 63.5%), English-speaking (3140/3192, 98.4%), and residing in socioeconomically disadvantaged areas (1879/3192, 58.9%) (reach). The model was feasible and safe to implement, managing 3240 admissions with no reported deaths. Patients valued continuous access to care and described better recovery experiences at home (effectiveness). Staff viewed the model as appropriate for identifying and managing high-risk patients in the community, easing pressure on hospital beds (adoption). The service cost Aus $ 5.4 million (US $3.5 million) over 11 months. Implementation barriers included the urgency of the pandemic scenario, limited infrastructure and human resources, and changing requirements in relation to COVID-19. These were mitigated by several people factors that were critical to its successful implementation, including a consultant-led structure, staff commitment, and adaptability (implementation). The service saved 16,651 inpatient bed days before being integrated into core HITH operations. The experience strengthened staff capabilities in emergency response, virtual care delivery, and strategic planning. The model shows promise for broader application into pediatric care, though further work is needed to enhance interdepartmental collaboration and staff recognition (maintenance). Conclusions: This study demonstrated that a virtual HITH model can be implemented effectively and safely at scale. Findings support its potential for integration into routine care, provided that adequate resource planning, a skilled and multidisciplinary workforce, well-defined care pathways, and equity-focused strategies are in place.
  • ✇STAT
  • STAT+: Fresh data on hospital AI use & Califf dishes on tech Mario Aguilar
    You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday. Califf warns AI in health care ‘overhyped’ On a makeshift stage in a Midtown Manhattan office earlier this week,former Food and Drug Administration Commissioner Robert Califf struck a measured tone about the potential for artificial intelligence in health care. Asked whether the technology was o
     

STAT+: Fresh data on hospital AI use & Califf dishes on tech

18 September 2025 at 22:18

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Califf warns AI in health care ‘overhyped’

On a makeshift stage in a Midtown Manhattan office earlier this week,former Food and Drug Administration Commissioner Robert Califf struck a measured tone about the potential for artificial intelligence in health care. Asked whether the technology was overhyped he said it was. “I hear way too much about the money. I’m not hearing a lot of human values coming through discussions,” he said. Adding:

“Almost all of the technology is being applied to optimizing the financial status of healthcare delivery entities or companies that are making medical products and that’s not aligned with equitable, better patient outcomes. So until someone puts a soul back in the system, I think it’s going to get worse and worse.”

Continue to STAT+ to read the full story…

© Adobe

Opinion: Four reasons why generative AI chatbots could lead to psychosis in vulnerable people

18 September 2025 at 16:30

Three scholars discovered a strange mirror deep in the forest. It spoke to them in a soothing voice and answered all their questions warmly, knowledgeably, and eloquently.

The captivated scholars became obsessed, whispering one secret after another to the mirror. It replied with affection, promise, and meaning that kept them returning to it. They began ignoring one another, each convinced the mirror “understood” them best.

Read the rest…

© Adobe

From frameworks to finance: how sharing benefits from the use of digital sequence information can evolve to contribute to biodiversity conservation

Nature Biotechnology, Published online: 18 September 2025; doi:10.1038/s41587-025-02820-8

The COP16 decision established a multilateral mechanism for digital sequence information (DSI) benefit-sharing. This Comment brings together insights from academia and commercial DSI researchers to assess what has been accomplished so far, identify remaining challenges and describe elements under discussion to support collective goals.

Diagnostic Performance of Computed Tomography–Based Artificial Intelligence for Early Recurrence of Cholangiocarcinoma: Systematic Review and Meta-Analysis

Background: Despite artificial intelligence (AI) models demonstrating high predictive accuracy for early cholangiocarcinoma recurrence, their clinical application faces challenges, such as reproducibility, generalizability, hidden biases, and uncertain performance across diverse datasets and populations, raising concerns about their practical applicability. Objective: This meta-analysis aims to systematically assess the diagnostic performance of AI models using computed tomography (CT) imaging to predict early recurrence of cholangiocarcinoma. Methods: A systematic search was conducted in PubMed, Embase, and Web of Science for studies published up to May 2025. Studies were selected based on the Participants, Index test, Target condition, Reference standard, Outcomes, and Setting (PITROS) framework. Participants included patients diagnosed with cholangiocarcinoma (including intrahepatic and extrahepatic locations). The index test was AI techniques applied to CT imaging for early recurrence prediction (defined as within 1 year), while the target condition was early recurrence of cholangiocarcinoma (positive group: recurrence; negative group: no recurrence). The reference standard was pathological diagnosis or imaging follow-up confirming recurrence. Outcomes included sensitivity, specificity, diagnostic odds ratio (DOR), and area under the receiver operating characteristic curve (AUC), assessed in both internal and external validation cohorts. The setting comprised retrospective or prospective studies using hospital datasets. Methodological quality was assessed using an optimized version of the revised Quality Assessment of Diagnostic Accuracy Studies-2 tool. Heterogeneity was assessed using the I² statistic. Pooled sensitivity, specificity, DOR, and AUC were calculated using a bivariate random-effects model. Results: A total of 9 studies with 30 datasets involving 1537 patients were included. In internal validation cohorts, CT-based AI models showed a pooled sensitivity of 0.87 (95% CI 0.81-0.92), specificity of 0.85 (95% CI 0.79-0.89), DOR of 37.71 (95% CI 18.35-77.51), and AUC of 0.93 (95% CI 0.90-0.94). In external validation cohorts, pooled sensitivity was 0.87 (95% CI 0.81-0.91), specificity was 0.82 (95% CI 0.77-0.86), DOR was 30.81 (95% CI 18.79-50.52), and AUC was 0.85 (95% CI 0.82-0.88). The AUC was significantly lower in external validation cohorts compared to internal validation cohorts (P<.001). Conclusions: Our results show that CT-based AI models predict early cholangiocarcinoma recurrence with high performance in internal validation sets and moderate performance in external validation sets. However, the high heterogeneity observed may impact the robustness of these results. Future research should focus on prospective studies and establishing standardized gold standards to further validate the clinical applicability and generalizability of AI models.

Large Language Models’ Clinical Decision-Making on When to Perform a Kidney Biopsy: Comparative Study

Background: Artificial intelligence (AI) and Large Language models (LLMs) are increasing in sophistication and are being integrated into many disciplines. The potential for LLMs to augment clinical decisions is an evolving area of research. Objective: This study compared the responses of over 1000 kidney specialist physicians (nephrologists) to outputs of commonly used LLMs using a questionnaire determining when a kidney biopsy should be performed. Methods: This research group completed a large online questionnaire for nephrologists to determine when a kidney biopsy should be performed. The questionnaire was co-designed with patient participation, refined through multiple iterations, then piloted locally before international dissemination. It was the largest international study in the field and demonstrated variation between human clinicians in biopsy propensity relating to human factors such as sex and age, as well as systemic factors such as country, job seniority and technical proficiency. The same questions were put to both human doctors and LLMs in an identical order in a single session. Eight commonly used LLMs were interrogated: Chat GPT 3.5, Mistral Hugging Face, Perplexity, Microsoft Co-pilot, Llama 2, GPT 4.0, MedLM and Claude 3. The most common response given by clinicians (human mode) to each question was taken as the baseline for comparison. Questionnaire responses to the indications and contraindications for biopsy generated a score (0-44) reflecting biopsy propensity, in which a higher score was used as a surrogate marker for an increased tolerance of potential associated risks. Results: The ability of LLMs to reproduce human expert consensus varied widely with some models demonstrating a balanced approach to risk in a similar manner to humans, whilst other models reported outputs at either end of the spectrum for risk tolerance. In terms of agreement with the human mode, Chat GPT 3.5 and GPT 4.0 (Open AI) had the highest levels of alignment, with the human mode selected in 6/11 questions. The total biopsy propensity score generated from the human mode was 23/44. Both Open AI models produced similar propensity scores between 22 and 24, however Llama 2 and MS Co-pilot also reported scores within this range, but with poorer response alignment to the human mode at only 2/11 questions. The most risk averse model in this study was MedLM with a propensity score of 11 and the least risk averse model was Claude 3 with a score of 34. Conclusions: LLM outputs demonstrated a modest ability to replicate human clinical decision making in this study, however the performance varied widely between LLM models. Questions with more uniform human responses produced LLM outputs with greater alignment, whereas in questions with low levels of human consensus there was poor output alignment. This may limit the practical use of LLMs in real world clinical practice.

Navigating the Boundaries of Teleconsultation—Capabilities, Limitations, and Pathways for Improvement: Qualitative Study of the Experiences of Patients With Stroke

Background: Survivors of stroke often face persistent challenges accessing postdischarge care due to mobility limitations, transportation burdens, and inflexible scheduling. Teleconsultation has emerged as a potential solution to improve continuity of care, but its perceived strengths and limitations from the patient perspective remain insufficiently understood. Objective: This study aimed to explore the experiences of survivors of stroke with a nurse-led teleconsultation program to (1) identify perceived capabilities; (2) understand limitations in usability, accessibility, and clinical function; and (3) generate patient-informed recommendations for improvement. Methods: A qualitative study was embedded within a 3-month nurse-led teleconsultation intervention delivered by advanced practice nurses. A total of 21 survivors of ischemic stroke (aged 45-76 y; female: n=11, 52%) who had preserved cognitive function (Montreal Cognitive Assessment score ≥22) and smartphone access participated in 6 focus groups conducted via Zoom. Data were analyzed thematically using an established framework. Data saturation was achieved. Results: Participants widely valued teleconsultation for reducing logistical burdens; enhancing access; and offering a more comfortable, emotionally supportive setting for follow-up care. Many reported increased awareness and motivation for self-monitoring. However, limitations included an inability to perform physical assessments or respond to emergencies; digital and usability barriers, especially among older users; and scheduling inflexibility. Participants emphasized the need for patient-initiated follow-up mechanisms, physician collaboration for medication management, and greater support for users considered digitally marginalized. They also highlighted the potential of teleconsultation to serve as a triage tool, reserving in-person care for complex cases. Conclusions: Nurse-led teleconsultation was perceived as a convenient and supportive modality for poststroke care, particularly for stable follow-ups and psychosocial support. However, its long-term viability depends on addressing clinical and technical limitations, enhancing user autonomy, and integrating interdisciplinary input. By centering the lived experiences of survivors of stroke, this study offers concrete recommendations to guide the development of more inclusive, responsive, and patient-centered teleconsultation models.

Developing an Evaluation System for Quality of Health Educational Short Videos on Social Media (LassVQ) Using Nominal Group Technique and Analytic Hierarchy Process: Qualitative Study

Background: With the increasing use of social media platforms for health communication, the quality of health educational short videos (HESVs) has become a key concern. However, no standardized framework exists to evaluate the quality of health videos on social media, highlighting the need for a comprehensive evaluation system. Objective: The aim of this study is to develop a valid and structured evaluation tool for assessing the quality of HESVs on social media. Methods: The initial evaluation indicators obtained from the literature review and brainstorming undertaken in the study group were provided to the nominal group reference Lasswell’s 5W communication model, and two rounds of nominal group technique (NGT) were carried out to screen, add, revise, and adjust indicators, and reach a consensus of evaluation system. The indicators were then ranked based on their significance, as scored by the experts using the analytic hierarchy process. The content validity was assessed by experts who rated the relevance of each indicator on a 4-point Likert scale. Results: The primary indicators include communicator, communication content, communication channel, and communication effect, along with 13 secondary indicators and 34 tertiary indicators. 11 experts were enrolled in the NGT, 45% of experts had a doctoral degree, 80% of them were ranked associate professor or professor. The average familiarity coefficient of each key indicator of the NGT was 0.85. The average values of the expert judgment coefficient and authority coefficient were 0.93 and 0.85, respectively. In Round 1 of NGT, the “Communication target” of 5 primary indicators, 7 of 20 secondary indicators, and 66 of 94 tertiary indicators did not reach a consensus, and therefore, they were not deleted and will proceed to the next round of NGT. In Round 2 NGT, 1 primary indicator, 7 secondary indicators, and 59 tertiary indicators were deleted based on the consensus criteria. After the two rounds of NGT, 4 primary indicators, 13 secondary indicators, and 34 tertiary indicators finally reached a consensus. Among primary indicators, communication content was found to be the most influential, accounting for 45.68%. Among secondary indicators, credibility, scientificity, availability, and social attention were the most influential indicators, with priorities of 56.67%, 24.26%, 74.62%, and 39.89% in their respective categories. Among tertiary indicators, ‘Become a hot search recommended by the platform’ was the most influential indicator with a weight of 0.07. The content validity of all the evaluation indicators were 0.73 – 1.0, and the scale-level content validity index (average) was 0.87, which was indicated as acceptable. Conclusions: The evaluation system for the quality of HESVs on social media (LassVQ) was developed, and its validity was acceptable. The proposed evaluation system can be used in conjunction with qualitative methods to gain a holistic perspective on the multidimensional quality of HESVs on social media.
  • ✇Cell
  • Multi-adjuvant personalized neoantigen vaccines: Fine-tuning anti-cancer T cells Hejia Henry Wang · Neeha Zaidi
    Personalized cancer vaccines aim to broaden the anti-tumor T cell repertoire by targeting neoantigens unique to each patient’s tumor, but immunogenicity has been inconsistent. In this issue of Cell, Blass, Keskin, Tu et al. evaluate NeoVaxMI, a multi-adjuvant personalized synthetic long-peptide vaccine administered with nivolumab in patients with melanoma. NeoVaxMI elicited stronger CD4+ and CD8+ responses than earlier iterations, and vaccine-induced T cells trafficked to regressing metastatic l
     

Multi-adjuvant personalized neoantigen vaccines: Fine-tuning anti-cancer T cells

18 September 2025 at 08:00
Personalized cancer vaccines aim to broaden the anti-tumor T cell repertoire by targeting neoantigens unique to each patient’s tumor, but immunogenicity has been inconsistent. In this issue of Cell, Blass, Keskin, Tu et al. evaluate NeoVaxMI, a multi-adjuvant personalized synthetic long-peptide vaccine administered with nivolumab in patients with melanoma. NeoVaxMI elicited stronger CD4+ and CD8+ responses than earlier iterations, and vaccine-induced T cells trafficked to regressing metastatic lesions.

The arts for disease prevention and health promotion: a systematic review

Nature Medicine, Published online: 18 September 2025; doi:10.1038/s41591-025-03962-7

The arts, according to a systematic synthesis of data from 95 studies (across 26 countries), may support non-communicable disease prevention by providing opportunities for increased physical activity, and helping to address social forces that contribute to health inequities.
❌