❌

Reading view

  •  

Opinion: My patient almost quit a clinical trial to save her job

When one of my patients was first diagnosed with sarcoidosis, her specialist gave her two options: take steroids or join a clinical trial. She opted for the trial, hoping it might lead to better treatment for herself and others.

However, her initial excitement waned as the logistical demands began to take a toll on her personal and professional life. With twice-monthly appointments, each trial day required an early two-hour drive, followed by eight hours of appointments, and concluded with a long drive home — all while managing symptoms of the disease. The trial sponsor eventually covered an overnight hotel room to ease her travel, but that doubled her time away from work and further amplified her stress.

Read the rest…

© Adobe

  •  

Molecular advances in early-stage and locally advanced non-small cell lung carcinoma: Shaping the future of precision oncology-systematic review

Sci Prog. 2025 Jul-Sep;108(3):368504251383055. doi: 10.1177/00368504251383055. Epub 2025 Sep 24.

ABSTRACT

ObjectiveTo synthesize recent molecular advances that inform diagnosis, risk-stratification, and perioperative treatment in early-stage and locally advanced non-small cell lung carcinoma (NSCLC), with emphasis on comprehensive genomic profiling, minimal residual disease (MRD) detection by circulating tumor DNA (ctDNA), and the translation of biomarkers into targeted and immunotherapy strategies.MethodsSystematic review registered in PROSPERO (CRD420251076423). Searches of PubMed, Scopus, Web of Science, and Embase (January 2015-April 2025) followed PRISMA 2020/PRISMA-S. From 4640 records, 890 duplicates were removed; 3750 titles/abstracts were screened; 150 full texts were assessed; 75 studies met inclusion criteria. Risk of bias used Newcastle-Ottawa Scale (NOS) for observational studies and Cochrane RoB 2 tool for randomized controlled trials; certainty was summarized with GRADE where applicable.ResultsActionable alterations (e.g. EGFR, ALK, KRAS, MET, RET, BRAF, NTRK) are prevalent in early-stage NSCLC and comparable to advanced disease, supporting routine comprehensive genomic profiling in curative-intent settings. Next-generation sequencing (NGS) and ctDNA enable the detection of MRD, earlier relapse prediction, and dynamic treatment monitoring. Perioperative strategies integrating targeted therapy and immunotherapy (e.g. adjuvant EGFR-TKI, neoadjuvant chemo-immunotherapy) improve pathological and disease-free outcomes in selected biomarker-defined populations. Evidence profiles generally show low-to-moderate risk of bias and moderate-to-high certainty for key outcomes related to profiling and MRD, with heterogeneity across platforms and endpoints.ConclusionsMolecular advances-particularly broad NGS and ctDNA-based MRD-are reshaping the perioperative management of early and locally advanced NSCLC, enabling precision selection for targeted and immunotherapy approaches. Standardization of testing workflows and reporting, and cost-effective implementation are priorities for equitable adoption and for future trials that combine NGS, MRD, and multi-omic/AI-driven risk stratification.

PMID:40990633 | PMC:PMC12461064 | DOI:10.1177/00368504251383055

  •  

Expanding care coordination in an integrated health system through causal machine learning

npj Digital Medicine, Published online: 24 September 2025; doi:10.1038/s41746-025-01925-3

Expanding care coordination in an integrated health system through causal machine learning
  •  

Diabetic Foot Ulcer Classification Models Using Artificial Intelligence and Machine Learning Techniques: Systematic Review

Background: Diabetes-related foot ulceration (DFU) is a common complication of diabetes, with a significant impact on survival, health care costs, and health-related quality of life. The prognosis of DFU varies widely among individuals. The International Working Group on the Diabetic Foot recently updated their guidelines on how to classify ulcers using “classical” classification and scoring systems. No system was recommended for individual prognostication, and the group considered that more detail in ulcer characterization was needed and that machine learning (ML)–based models may be the solution. Despite advances in the field, no assessment of available evidence was done. Objective: This study aimed to identify and collect available evidence assessing the ability of ML-based models to predict clinical outcomes in people with DFU. Methods: We searched the MEDLINE database (PubMed), Scopus, Web of Science, and IEEE Xplore for papers published up to July 2023. Studies were eligible if they were anterograde analytical studies that examined the prognostic abilities of ML models in predicting clinical outcomes in a population that included at least 80% of adults with DFU. The literature was screened independently by 2 investigators (MMS and DAR or EH in the first phase, and MMS and MAS in the second phase) for eligibility criteria and data extracted. The risk of bias was evaluated using the Quality In Prognosis Studies tool and the Prediction model Risk Of Bias Assessment Tool by 2 investigators (MMS and MAS) independently. A narrative synthesis was conducted. Results: We retrieved a total of 2412 references after removing duplicates, of which 167 were subjected to full-text screening. Two references were added from searching relevant studies’ lists of references. A total of 11 studies, comprising 13 papers, were included focusing on 3 outcomes: wound healing, lower extremity amputation, and mortality. Overall, 55 predictive models were created using mostly clinical characteristics, random forest as the developing method, and area under the receiver operating characteristic curve (AUROC) as a discrimination accuracy measure. AUROC varied from 0.56 to 0.94, with the majority of the models reporting an AUROC equal or superior to 0.8 but lacking 95% CIs. All studies were found to have a high risk of bias, mainly due to a lack of uniform variable definitions, outcome definitions and follow-up periods, insufficient sample sizes, and inadequate handling of missing data. Conclusions: We identified several ML-based models predicting clinical outcomes with good discriminatory ability in people with DFU. Due to the focus on development and internal validation of the models, the proposal of several models in each study without selecting the “best one,” and the use of nonexplainable techniques, the use of this type of model is clearly impaired. Future studies externally validating explainable models are needed so that ML models can become a reality in DFU care. Trial Registration: PROSPERO CRD42022308248; https://www.crd.york.ac.uk/PROSPERO/view/CRD42022308248
  •  

Article: InfoQ AI, ML and Data Engineering Trends Report - 2025

This InfoQ Trends Report offers readers a comprehensive overview of emerging trends and technologies in the areas of AI, ML, and Data Engineering. This report summarizes the InfoQ editorial team’s and external guests' view on the current trends in AI and ML technologies and what to look out for in the next 12 months.

By Srini Penchikala, Savannah Kunovsky, Anthony Alford, Daniel Dominguez, Vinod Goje
  •  

Martin Frederik, Snowflake: Data quality is key to AI-driven growth

As companies race to implement AI, many are finding that project success hinges directly on the quality of their data. This dependency is causing many ambitious initiatives to stall, never making it beyond the experimental proof-of-concept stage.

So, what’s the secret to turning these experiments into real revenue generators? AI News caught up with Martin Frederik, regional leader for the Netherlands, Belgium, and Luxembourg at data cloud giant Snowflake, to find out.

“There’s no AI strategy without a data strategy,” Frederik says simply. “AI apps, agents, and models are only as effective as the data they’re built on, and without unified, well-governed data infrastructure, even the most advanced models can fall short.”

Improving data quality is key to AI project success

It’s a familiar story for many organisations: a promising proof-of-concept impresses the team but never translates into a tool that makes the company money. According to Frederik, this often happens because leaders treat the technology as the end goal.

Headshot of Martin Frederik, regional leader for the Netherlands, Belgium, and Luxembourg at AI data cloud giant Snowflake.

“AI is not the destination – it’s the vehicle to achieving your business goals,” Frederik advises.

When projects get stuck, it’s usually down to a few common culprits: the project isn’t truly aligned with what the business needs, teams aren’t talking to each other, or the data is a mess. It’s easy to get disheartened by statistics suggesting that 80% of AI projects don’t reach production, but Frederik offers a different perspective. This isn’t necessarily a failure, he suggests, but “part of the maturation process”.

For those who get the foundation right, the payoff is very real. A recent Snowflake study found that 92% of companies are already seeing a return on their AI investments. In fact, for every £1 spent, they’re getting back £1.41 in cost savings and new revenue. The key, Frederik repeats, is having a “secure, governed and centralised platform” for your data from the very beginning.

It’s not just about tech, it’s about people

Even with the best technology, an AI strategy can fall flat if the company culture isn’t ready for it. One of the biggest challenges is getting data into the hands of everyone who needs it, not just a select few data scientists. To make AI work at scale, you have to build strong foundations in your “people, processes, and technology.”

This means breaking down the walls between departments and making quality data and AI tools accessible to everyone.

“With the right governance, AI becomes a shared resource rather than a siloed tool,” Frederik explains. When everyone works from a single source of truth, teams can stop arguing about whose numbers are correct and start making faster and smarter decisions together.

The next leap: AI that reasons for itself

The true breakthrough we’re seeing now is the emergence of AI agents that can understand and reason over all kinds of data at once regardless of structure quality; from the neat rows and columns in a spreadsheet, to the unstructured information in documents, videos, and emails. Considering that this unstructured data makes up 80-90% of a typical company’s data, this is a huge step forward.

New tools are enabling staff, no matter their technical skill level, to simply ask complex questions in plain English and get answers directly from the data.

Frederik explains that this is a move towards what he calls “goal-directed autonomy”. Until now, AI has been a helpful assistant you had to constantly direct. “You ask a question, you get an answer; you ask for code, you get a snippet,” he notes.

The next generation of AI is different. You can give an agent a complex goal, and it will figure out the necessary steps on its own, from writing code to pulling in information from other apps to deliver a complete answer. This will automate the most time-consuming parts of a data scientist’s job, like “tedious data cleaning” and “repetitive model tuning.”

The result? It frees up your brightest minds to focus on what really matters. This elevates your people “from practitioner to strategist” and allows them to drive real value for the business. That can only be a good thing.

Snowflake is a key sponsor of this year’s AI & Big Data Expo Europe and will have a range of speakers sharing their deep insights during the event. Swing by Snowflake’s booth at stand number 50 to hear more from the company about making enterprise AI easy, efficient, and trusted.

See also: Public trust deficit is a major hurdle for AI growth

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Martin Frederik, Snowflake: Data quality is key to AI-driven growth appeared first on AI News.

  •  

Long-read sequencing and chemical synthesis access previously hidden antibiotics

Nature Biotechnology, Published online: 23 September 2025; doi:10.1038/s41587-025-02861-z

Most bacteria cannot be grown in the laboratory, which means that their genetic diversity is hidden to traditional culture-based studies. We combined a new DNA extraction method, long-read sequencing, bioinformatics and chemical synthesis to access the genetic diversity of uncultured soil bacteria and abiotically decode it to discover bioactive small molecules.
  •  

Fine-Tuning Methods for Large Language Models in Clinical Medicine by Supervised Fine-Tuning and Direct Preference Optimization: Comparative Evaluation

Background: Large language model (LLM) fine tuning is the process of adjusting out-of-the-box model weights using a dataset of interest. Fine tuning can be a powerful technique to improve model performance in fields like medicine, where data access is restricted and LLMs may have poor out-of-the-box performance. Objective: In this study we investigated the benefits of fine tuning with supervised fine tuning (SFT) and direct preference optimization (DPO) across a range of LLM applications for medicine Methods: We use Llama3 7B and Mistral 7B v2 to compare the performance of SFT and DPO across four datasets for common natural language tasks in medicine. The tasks evaluated were simple classification, clinical reasoning, summarization, and clinical triage. Results: Clinical Reasoning accuracy increased 8% and 7% with DPO over SFT for Llama3 (p value 0.003) and Mistral2 (p value 0.004) respectively. Summarization quality, graded on a five point Likert scale, increased 0.13 and 0.10 for Llama3 and Mistral2 (p values
  •  

Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions: Multicenter Observational Study

Background: Large language models (LLMs) hold promise for supporting clinical tasks, particularly in data-driven and technical disciplines such as radiation oncology. While prior evaluation studies have focused on examination-style settings for evaluating LLMs, their performance in real-life clinical scenarios remains unclear. In the future, LLMs might be used as general AI assistants to answer questions arising in clinical practice. It is unclear how well a modern LLM, locally executed within the infrastructure of a hospital, would answer such questions compared with clinical experts. Objective: This study aimed to assess the performance of a locally deployed, state-of-the-art medical LLM in answering real-world clinical questions in radiation oncology compared with clinical experts. The aim was to evaluate the overall quality of answers, as well as the potential harmfulness of the answers if used for clinical decision-making. Methods: Physicians from 10 departments of European hospitals collected questions arising in the clinical practice of radiation oncology. Fifty of these questions were answered by 3 senior radiation oncology experts with at least 10 years of work experience, as well as the LLM Llama3-OpenBioLLM-70B (Ankit Pal and Malaikannan Sankarasubbu). In a blinded review, physicians rated the overall answer quality on a 5-point Likert scale (quality), assessed whether an answer might be potentially harmful if used for clinical decision-making (harmfulness), and determined if responses were from an expert or the LLM (recognizability). Comparisons between clinical experts and LLMs were then made for quality, harmfulness, and recognizability. Results: There were no significant differences between the quality of the answers between LLM and clinical experts (mean scores of 3.38 vs 3.63; median 4.00, IQR 3.00-4.00 vs median 3.67, IQR 3.33-4.00; P=.26; Wilcoxon signed rank test). The answers were deemed potentially harmful in 13% of cases for the clinical experts compared with 16% of cases for the LLM (P=.63; Fisher exact test). Physicians correctly identified whether an answer was given by a clinical expert or an LLM in 78% and 72% of cases, respectively. Conclusions: A state-of-the-art medical LLM can answer real-life questions from the clinical practice of radiation oncology similarly well as clinical experts regarding overall quality and potential harmfulness. Such LLMs can already be deployed within the local hospital environment at an affordable cost. While LLMs may not yet be ready for clinical implementation as general AI assistants, the technology continues to improve at a rapid pace. Evaluation studies based on real-life situations are important to better understand the weaknesses and limitations of LLMs in clinical practice. Such studies are also crucial to define when the technology is ready for clinical implementation. Furthermore, education for health care professionals on generative AI is needed to ensure responsible clinical implementation of this transforming technology.
  •  

AI models are using material from retracted scientific papers

Some AI chatbots rely on flawed research from retracted scientific papers to answer questions, according to recent studies. The findings, confirmed by MIT Technology Review, raise questions about how reliable AI tools are at evaluating scientific research and could complicate efforts by countries and industries seeking to invest in AI tools for scientists.

AI search tools and chatbots are already known to fabricate links and references. But answers based on the material from actual papers can mislead as well if those papers have been retracted. The chatbot is “using a real paper, real material, to tell you something,” says Weikuan Gu, a medical researcher at the University of Tennessee in Memphis and an author of one of the recent studies. But, he says, if people only look at the content of the answer and do not click through to the paper and see that it’s been retracted, that’s really a problem. 

Gu and his team asked OpenAI’s ChatGPT, running on the GPT-4o model, questions based on information from 21 retracted papers about medical imaging. The chatbot’s answers referenced retracted papers in five cases but advised caution in only three. While it cited non-retracted papers for other questions, the authors note that it may not have recognized the retraction status of the articles. In a study from August, a different group of researchers used ChatGPT-4o mini to evaluate the quality of 217 retracted and low-quality papers from different scientific fields; they found that none of the chatbot’s responses mentioned retractions or other concerns. (No similar studies have been released on GPT-5, which came out in August.)

The public uses AI chatbots to ask for medical advice and diagnose health conditions. Students and scientists increasingly use science-focused AI tools to review existing scientific literature and summarize papers. That kind of usage is likely to increase. The US National Science Foundation, for instance, invested $75 million in building AI models for science research this August.

“If [a tool is] facing the general public, then using retraction as a kind of quality indicator is very important,” says Yuanxi Fu, an information science researcher at the University of Illinois Urbana-Champaign. There’s “kind of an agreement that retracted papers have been struck off the record of science,” she says, “and the people who are outside of science—they should be warned that these are retracted papers.” OpenAI did not provide a response to a request for comment about the paper results.

The problem is not limited to ChatGPT. In June, MIT Technology Review tested AI tools specifically advertised for research work, such as Elicit, Ai2 ScholarQA (now part of the Allen Institute for Artificial Intelligence’s Asta tool), Perplexity, and Consensus, using questions based on the 21 retracted papers in Gu’s study. Elicit referenced five of the retracted papers in its answers, while Ai2 ScholarQA referenced 17, Perplexity 11, and Consensus 18—all without noting the retractions.

Some companies have since made moves to correct the issue. “Until recently, we didn’t have great retraction data in our search engine,” says Christian Salem, cofounder of Consensus. His company has now started using retraction data from a combination of sources, including publishers and data aggregators, independent web crawling, and Retraction Watch, which manually curates and maintains a database of retractions. In a test of the same papers in August, Consensus cited only five retracted papers. 

Elicit told MIT Technology Review that it removes retracted papers flagged by the scholarly research catalogue OpenAlex from its database and is “still working on aggregating sources of retractions.” Ai2 told us that its tool does not automatically detect or remove retracted papers currently. Perplexity said that it “[does] not ever claim to be 100% accurate.” 

However, relying on retraction databases may not be enough. Ivan Oransky, the cofounder of Retraction Watch, is careful not to describe it as a comprehensive database, saying that creating one would require more resources than anyone has: “The reason it’s resource intensive is because someone has to do it all by hand if you want it to be accurate.”

Further complicating the matter is that publishers don’t share a uniform approach to retraction notices. “Where things are retracted, they can be marked as such in very different ways,” says Caitlin Bakker from University of Regina, Canada, an expert in research and discovery tools. “Correction,” “expression of concern,” “erratum,” and “retracted” are among some labels publishers may add to research papers—and these labels can be added for many reasons, including concerns about the content, methodology, and data or the presence of conflicts of interest. 

Some researchers distribute their papers on preprint servers, paper repositories, and other websites, causing copies to be scattered around the web. Moreover, the data used to train AI models may not be up to date. If a paper is retracted after the model’s training cutoff date, its responses might not instantaneously reflect what’s going on, says Fu. Most academic search engines don’t do a real-time check against retraction data, so you are at the mercy of how accurate their corpus is, says Aaron Tay, a librarian at Singapore Management University.

Oransky and other experts advocate making more context available for models to use when creating a response. This could mean publishing information that already exists, like peer reviews commissioned by journals and critiques from the review site PubPeer, alongside the published paper.  

Many publishers, such as Nature and the BMJ, publish retraction notices as separate articles linked to the paper, outside paywalls. Fu says companies need to effectively make use of such information, as well as any news articles in a model’s training data that mention a paper’s retraction. 

The users and creators of AI tools need to do their due diligence. “We are at the very, very early stages, and essentially you have to be skeptical,” says Tay.

Ananya is a freelance science and technology journalist based in Bengaluru, India.

  •  

Deciphering the Heterogeneity of Pancreatic Cancer: DNA Methylation-Based Cell Type Deconvolution Unveils Distinct Subgroups and Immune Landscapes

Epigenomes. 2025 Sep 5;9(3):34. doi: 10.3390/epigenomes9030034.

ABSTRACT

Background: Pancreatic ductal adenocarcinoma (PDAC) is a highly heterogeneous malignancy, characterized by low tumor cellularity, a dense stromal response, and intricate cellular and molecular interactions within the tumor microenvironment (TME). Although bulk omics technologies have enhanced our understanding of the molecular landscape of PDAC, the specific contributions of non-malignant immune and stromal components to tumor progression and therapeutic response remain poorly understood. Methods: We explored genome-wide DNA methylation and transcriptomic data from the Cancer Genome Atlas Pancreatic Adenocarcinoma cohort (TCGA-PAAD) to profile the immune composition of the TME and uncover gene co-expression networks. Bioinformatic analyses included DNA methylation profiling followed by hierarchical deconvolution, epigenetic age estimation, and a weighted gene co-expression network analysis (WGCNA). Results: The unsupervised clustering of methylation profiles identified two major tumor groups, with Group 2 (n = 98) exhibiting higher tumor purity and a greater frequency of KRAS mutations compared to Group 1 (n = 87) (p < 0.0001). The hierarchical deconvolution of DNA methylation data revealed three distinct TME subtypes, termed hypo-inflamed (immune-deserted), myeloid-enriched, and lymphoid-enriched (notably T-cell predominant). These immune clusters were further supported by co-expression modules identified via WGCNA, which were enriched in immune regulatory and signaling pathways. Conclusions: This integrative epigenomic-transcriptomic analysis offers a robust framework for stratifying PDAC patients based on the tumor immune microenvironment (TIME), providing valuable insights for biomarker discovery and the development of precision immunotherapies.

PMID:40981070 | PMC:PMC12452622 | DOI:10.3390/epigenomes9030034

  •  
❌