❌

Reading view

Businesses still face the AI data challenge

A few years ago, the business technology world’s favourite buzzword was ‘Big Data’ – a reference to organisations’ mass collection of information that could be used to suggest previously unexplored ways of operating, and float ideas about what strategies they may best pursue.

What’s becoming increasingly apparent is that the problems companies faced in using Big Data to their advantage still remain, and it’s a new technology – AI – that’s making those problems rise once again to the surface. Without tackling the problems that beset Big Data, AI implementations will continue to fail.

So what are the issues stopping AI deliver on its promises?

The vast majority of problems stem from the data resources themselves. To understand the issue, consider the following sources of information used in a very average working day.

In a small-to-medium sized business:

  • Spreadsheets, stored on users’ laptops, in Google Sheets, Office 365 cloud.
  • The customer relationship manager (CRM) platform.
  • Email exchanges between colleagues, customers, suppliers.
  • Word documents, PDFs, web forms.
  • Messaging apps.

In an enterprise business:

  • All of the above, plus,
  • Enterprise resource planning (ERP) systems.
  • Real-time data feeds.
  • Data lakes.
  • Disparate databases behind multiple point-products.

It’s worth noting that the simple list above isn’t comprehensive, and nor is it intended to be. What it demonstrates is that in just five lines, there are around a dozen places where information can be found. What Big Data needed (perhaps still needs) and what AI projects also rest on, is somehow bringing all those elements together in such a way that a computer algorithm can make sense of it.

Marketing behemoth Gartner’s hype cycle for artificial intelligence, 2024, placed AI-Ready Data on the upward curve of the hype cycle, estimating it would be 2-5 years before it reached the ‘plateau of productivity’. Given that AI systems mine and extract data, most organisations – save those of the very largest size – don’t have the foundations on which to build, and may not have AI assistance in the endeavour for another 1-4 years.

The underlying problem for AI implementation is the same as dogged Big Data innovations as they, in the past, made their way through the hype cycle – from innovation trigger, peak of inflated expectations, trough of disillusionment, slope of enlightenment, to plateau of productivity – data comes in many forms; it can be inconsistent; perhaps it adheres to different standards; it may be inaccurate or biased; it could be highly sensitive information, or old and therefore irrelevant.

Transforming data so it’s AI-ready remains a process that’s as relevant today (perhaps more so) than it’s ever been. Those companies wanting to get a jump start could experiment with the many data treatment platforms currently available, and as is becoming the common advice, might begin with discrete projects as test-beds to assess the effectiveness of emerging technologies.

The advantage of the latest data preparation and assembly systems is that they are designed to prepare an organisation’s information resources in ways that are designed for the data to be used by AI value-creation platforms. They can offer, for example, carefully-coded guardrails that will help ensure data compliance, and protect users from accessing biased or commercially-sensitive information.

But the challenge of producing coherent, safe, and well-formulated data resources remains an ongoing issue. As organisations gain more data in their everyday operations, compiling up-to-date data resources on which to draw is a constant process. Where big data could be considered a static asset, data for AI ingestion has to be prepared and treated in as close to real-time as possible.

The situation therefore remains a three-way balance between opportunity, risk, and cost. Never before has the choice of vendor or platform been so crucial to the modern business.

(Source: “Inside the business school” by Darien and Neil is licensed under CC BY-NC 2.0.)

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and co-located with other leading technology events. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Businesses still face the AI data challenge appeared first on AI News.

  •  

Quantum cryptography and data protection for medical devices before and after they meet Q-Day

npj Digital Medicine, Published online: 21 October 2025; doi:10.1038/s41746-025-02082-3

Although still at a nascent state, quantum computing promises advances in healthcare, from drug discovery to personalised treatments. But it also threatens current cryptographic systems that protect medical data and infrastructure. The concept of “Q-Day” highlights risks such as “harvest now, decrypt later” attacks, with particular concerns for medical devices and sensitive applications in fields like femtech. Preparing for this future requires the rapid adoption of post-quantum cryptography, the coordination of time-phased and scalable “technology rollout” strategies, and revised regulatory frameworks to safeguard patient safety, privacy, and trust.
  •  

Article: Three Questions That Help You Build a Better Software Architecture

To architect effectively for an MVP, teams must answer three questions in order: Is the business idea worth pursuing? What performance and scalability are needed? How much maintainability and supportability are required? These guide Minimum Viable Architecture decisions. Empirical testing helps reject costly assumptions early and adapt architecture as the MVP evolves.

By Pierre Pureur, Kurt Bittner
  •  

STAT+: Effort to guide treatment of cancer with novel blood tests may get boost from new research

BERLIN — For years, scientists have held out hope that tests that look for the molecular fingerprints of a cancer’s presence could help clinicians determine which patients need further treatment after surgery, and which can be considered truly cured after an operation. Perhaps, the thinking goes, those in the latter category could be spared from unnecessary, expensive therapies that can carry serious side effects.

A study presented here Monday bolstered the case that such tests, which detect circulating tumor DNA, or ctDNA, that cancer cells slough off into the blood, could be effective tools. 

Researchers used the diagnostics to identify patients who appeared to have remnants of muscle-invasive bladder cancer after surgery, and directed them to treatment. Those who tested negative were allowed to live treatment-free, and were ultimately found to be at low risk of recurrence. 

Continue to STAT+ to read the full story…

© STEVE GSCHMEISSNER / SCIENCE PHOTO LIBRARY

  •  

STAT+: Is a battle brewing between Abridge and OpenEvidence?

You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.

Abridge vs. OpenEvidence, round one

New announcements from Abridge and OpenEvidence,  two of the most prominent health tech startups to emerge in recent years, highlight how despite different initial offerings, many artificial intelligence companies in health care will just end up competing with each other for physician eyeballs.

On Monday morning, Abridge, best known for its AI scribe that helps doctors automate the writing of clinical notes, announced a new product that will surface “real-time insights, prompts, and pathways” from the widely-used medical resource UpToDate based on things said in a clinical conversation and that are written in the patient’s record. 

Continue to STAT+ to read the full story…

© ADOBE, Alex Hogan/STAT

  •  

Effectiveness of a Digital Therapy on 6-Month Weight Loss in People With Obesity: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) Randomized Clinical Trial

Background: Obesity is a chronic, relapsing disease influenced by environmental, lifestyle, biological, and genetic factors, affecting over 1 billion people globally. Treatment for adults typically involves multicomponent lifestyle interventions—diet, physical activity, and behavior change—for at least 6-12 months. However, adherence is often low, and in-person sessions can be time-consuming and costly. Digital therapeutics (DTx), which enhance patient engagement and support long-term outcomes, have proven effective in managing chronic and mental health conditions. DTx offer scalable, evidence-based solutions with the potential to improve obesity management. Objective: The Digital Therapy to Promote Weight Loss in Patients With Obesity by Increasing Their Adherence to Treatment (DEMETRA) study is a prospective, multicenter, pragmatic, randomized, double-arm, single-blind, placebo-controlled trial evaluating the 6-month efficacy of an innovative, multicomponent digital intervention for obesity, which combines dietary, physical activity, and behavioral strategies in people with obesity (primary objective). Secondary objectives were assessing changes in BMI, waist circumference, blood pressure, glucose metabolism, lipid profile, adherence, and factors associated with absolute 6-month weight loss. Methods: The trial was conducted at 2 obesity centers in Italy with 246 participants aged 18-65 years (BMI 30-45 kg/m2), randomly assigned to either the Digital Therapeutics for Obesity (DTxO) app or a placebo app. DTxO offered personalized diet plans, exercise routines, and psycho-behavioral support, while the placebo app only allowed users to log data without feedback. Both groups followed a Mediterranean-style low-calorie diet with an 800 kcal/day deficit. On average, participants used the DTxO app for 42 minutes/day and the placebo app for 35 minutes, primarily for physical activity tracking. Univariable and multivariable generalized linear models were used to assess associations with 6-month absolute weight change (primary end point) and percent weight change (secondary end point). Results: Overall, 207 participants (84.1%) completed the 6-month visit. Both arms achieved a statistically significant absolute (DtxO: –3.2 kg, IQR –6.0 kg to –0.9 kg; placebo: –4.0 kg, IQR –6.9 kg to –0.5 kg; P<.001) and percent loss in body weight (DtxO: –3.0%, IQR –5.7% to –0.8%; placebo: –4.0%, IQR –8.5% to –0.5%; P<.001) after 6 months, without significant between-group differences (univariable generalized linear models: P=.34 and P=.17, respectively). Univariable regression analyses showed a significant association between adherence to app use and 6-month absolute weight loss (β=–.06, SE 0.02, P=.01) as well as percent weight loss (β=–.05, SE 0.01, P=.01). Adherent participants, defined as those with overall adherence at or above the 75th percentile of daily usage, included 35 individuals in the intervention group and 10 in the placebo group. In this subgroup, the estimated 6-month mean absolute weight change was –7.02 kg (95% CI –9.45 to –4.59) in the DTxO-adherent group and –3.50 kg (95% CI –7.01 to 0.01) in the placebo-adherent group (P=.02). The estimated 6-month mean percent change in weight was –6.31% (95% CI –8.86 to –3.76) in the DTxO-adherent group and –2.78% (95% CI –6.48 to 0.92) in the placebo-adherent group (P=.03). A significantly greater weight loss (P=.01 for study arm, either on absolute or percent change in weight from baseline) among adherent participants randomized to the DTxO app was also confirmed by analyses using mixed linear models for repeated measures. Conclusions: Although overall weight loss did not differ significantly between the DTxO and placebo groups, participants who used the DTxO app for at least 40% of the expected time achieved significantly greater weight loss. These results suggest that higher engagement with DTx can improve obesity outcomes. Further research should explore combining DTxO with pharmacological treatments or bariatric surgery. Trial Registration: ClinicalTrials.gov NCT05394779; https://clinicaltrials.gov/ct2/show/NCT05394779
  •  

Improving Large Language Model Applications in the Medical and Nursing Domains With Retrieval-Augmented Generation: Scoping Review

Background: Retrieval-augmented generation (RAG) is increasingly used to improve large language models in the medical and nursing domains. However, a comprehensive understanding of its specific architecture and applications in medical and nursing reasoning remains limited. Objective: We aimed to summarize the current state, existing limitations, and future development directions of RAG in the medical and nursing domains. Methods: The PubMed, Web of Science, IEEE Xplore, and arXiv databases were searched for relevant articles using queries that combined terms related to RAG, medical, and nursing domains, covering the period from November 1, 2022, to May 31, 2025. This review was conducted following the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines. Results: A total of 917 articles were retrieved, of which 67 met the inclusion criteria. Most studies focused on the medical domain (63/67, 94%), while only a few addressed nursing applications (4/67, 6%). The RAG frameworks included in this review were categorized into 5 functional types: text-based RAG (36/67, 54%), knowledge graph–enhanced RAG (17/67, 25%), agentic RAG (6/67, 9%), multimodal RAG (2/67, 3%), and plug-and-play RAG (6/67, 9%). On the basis of the Simon decision-making process theory, we divided the RAG workflow into 4 stages: intent recognition, knowledge retrieval, knowledge integration, and generation. Only 26 studies included explicit reasoning support, and few were aligned with real-world clinical workflows. Only 12 studies attempted to address ethical considerations related to RAG. Conclusions: We identified 4 key shifts in recent RAG development: shifting from surface-level matching toward contextualized intent recognition, from vague semantics toward logic-driven dynamic retrieval, from passive toward active knowledge retrieval, and from simple aggregation toward coherent context construction. However, most RAG systems in the medical and nursing domains have not yet introduced reasoning methods, and those that have are still predominantly reliant on data‑driven associations without causal modeling. This highlights the need to integrate causal mechanisms for more effective and domain-relevant reasoning in health care. Trial Registration: OSF Registries 10.17605/OSF.IO/WBSV5; https://osf.io/wbsv5
  •  

Toward the best generalizable performance of machine learning in modeling omic and clinical data

Lab Invest. 2025 Oct 15:104253. doi: 10.1016/j.labinv.2025.104253. Online ahead of print.

ABSTRACT

There are often performance differences between intra-dataset and cross-dataset tests in machine learning (ML) modeling. However, reducing these differences may reduce ML performances. It is thus a challenging dilemma for developing models that excel in intra-dataset testing and are generalizable to cross-dataset testing. Therefore, we aimed to understand and improve performance and generalizability of ML in intra-dataset and cross-dataset testing. We evaluated 4,200 ML models of classifying lung adenocarcinoma (LUAD) deaths using the The Cancer Genome Atlas (TCGA, n=286) and Oncogenomic-Singapore (OncoSG, n=167) datasets, and 1,680 models of classifying glioblastoma deaths using TCGA (n=151) and Clinical Proteomic Tumor Analysis Consortium (CPTAC, n=97) datasets. After examining performance distributions of these ML models, we applied a dual analytical framework, including statistical analyses and SHapley Additive exPlanations-based meta-analysis, to quantify factors' importance and trace model success back to design principles. We also developed a framework to identify the best generalizable model. Strikingly, Jarque-Bera test revealed significant deviations of model performances from normality in both cancer types and testing contexts. Simple linear models with sparse feature sets consistently dominated in LUAD experiments, whereas non-linear models dominated in glioblastoma ones, suggesting that the best modeling strategy appears cancer-type/disease dependent. Importantly, both robust Analysis of Variance (ANOVA) and Kruskal-Wallis tests consistently identified differentially expressed genes as one of the most influential factors in both cancer types. The proposed multi-criteria framework successfully identified the model that achieved both the best cross-dataset performance and similar intra-dataset performance. In summary, ML performance distributions significantly deviated from normality, which motivates using both robust parametric and non-parametric statistical tests. We quantified and provided possible exploitability on the factors associated with cross-dataset performances and generalizability of ML models in two cancer types. A multi-criteria framework was developed and validated to identify the models that are accurate and consistently robust cross datasets.

PMID:41106592 | DOI:10.1016/j.labinv.2025.104253

  •  

Multi-omics analyses inform mechanisms of immunotherapy response in pancreatic cancer

Front Immunol. 2025 Oct 2;16:1673098. doi: 10.3389/fimmu.2025.1673098. eCollection 2025.

ABSTRACT

INTRODUCTION: Pancreatic ductal adenocarcinoma (PDAC) continues to exhibit resistance to immunotherapy. In this study, we evaluated the efficacy of combining immunotherapy with chemotherapy for the treatment of advanced pancreatic cancer. Additionally, we employed a multimodal analytical approach to elucidate the immune landscape and conduct transcriptomic profiling in PDAC.

METHODS: A retrospective analysis was conducted on the clinical data of 52 patients diagnosed with advanced PDAC who underwent a combined treatment regimen of immunotherapy and chemotherapy. The study evaluated the objective response rate (ORR), disease control rate (DCR), and progression-free survival (PFS). To characterize the immune landscape in treatment-naive pancreatic ductal adenocarcinoma (PDAC) tumors and in the systemic circulation, flow cytometry, multiplex immunohistochemistry (mIHC), and whole transcriptome sequencing were employed.

RESULTS: The study reported an ORR of 32.7%, a DCR of 67.3%, and a 6-month PFS rate of 38.5%, with a median PFS of 5.5 months. Patients treated with a combination of immunotherapy and gemcitabine achieved the longest PFS. The first-line treatment cohort exhibited a significantly higher DCR (79.3% vs. 52.2%, P = 0.038) and a longer median PFS (6.6 vs. 3.5 months, P = 0.032) compared to the second-line treatment cohort. The efficacy of treatment varied depending on the drug combinations used. Flow cytometry analysis revealed a greater frequency of CD45- CD64+ cells in the peripheral blood of patients with progressive disease (PD) compared to those with a partial response (PR). Multiplex immunofluorescence (MIF) analysis indicated an increased intratumoral infiltration of CD8+ T cells and CD137+ CD8+ T cells in patients with PR. Whole transcriptome sequencing (WTSS) identified key genes involved in immune regulation, signal transduction, and digestive function. Hemopexin (HPX) and regulatory factor X-associated protein (RFXAP) were upregulated in PR patients and showed a positive correlation with survival, whereas Interleukin-6 (IL-6) expression was linked to poor prognosis.

CONCLUSIONS: These findings indicate that immunochemotherapy shows potential for the treatment of advanced PDAC. Our study elucidates the immune landscape associated with PDAC and provides critical insights for the identification of prospective therapeutic targets, which could guide the development of innovative combination immunotherapy strategies.

PMID:41112307 | PMC:PMC12528169 | DOI:10.3389/fimmu.2025.1673098

  •  

Integrative Transcriptomic and Metabolomic Analysis Reveals Aberrant Glycosylation as a Hallmark of Lung Adenocarcinoma

OMICS. 2025 Oct 16. doi: 10.1177/15578100251387518. Online ahead of print.

ABSTRACT

Lung adenocarcinoma (LUAD) remains the most common subtype of lung cancer, characterized by high heterogeneity and poor survival outcomes. Although transcriptomic and metabolomic alterations have been individually studied, integrated multi-omics analyses are needed to uncover the convergent pathways that drive tumor progression. Differentially expressed genes (DEGs) were identified from the GSE229253 transcriptomic dataset comprising LUAD tumor and adjacent normal tissues, while significantly altered metabolites were obtained from the Lung Cancer Metabolome Database. The top 10 DEGs and metabolites were analyzed using the search tool for interacting chemicals (STITCH) to construct gene-metabolite networks, and Integrated Molecular Pathway Level Analysis (IMPaLA) was employed for integrated pathway enrichment to identify overlapping molecular processes. Transcriptomic profiling revealed 973 DEGs (410 upregulated and 563 downregulated), and metabolomic analysis identified significant alterations in metabolites linked to redox balance, amino acid derivatives, and nucleotide metabolism. Integration through STITCH generated a network of 16 nodes and 9 edges, highlighting gene-metabolite associations of probable biological relevance. Joint pathway enrichment analysis using IMPaLA consistently identified glycosylation-related pathways, particularly O-linked glycosylation of mucins, as major axes of convergence between transcriptomic and metabolomic alterations in LUAD (joint p = 0.00129-0.00434). Several genes (B3GNT6, FEZF1-AS1, and LCAL1) and metabolites (isoleucylleucine, leucylleucine, and isoleucylvaline) are probable novel candidates, warranting further investigation. These findings provide systems-level evidence that aberrant glycosylation is likely a central hallmark of LUAD, underscore the potential of glycosylation pathways as biomarkers and therapeutic targets, and demonstrate the utility of cross-omics approaches to unpack the molecular complexity of lung cancer.

PMID:41103242 | DOI:10.1177/15578100251387518

  •  

Are cancer surgeries removing the body’s secret weapon against cancer?

Scientists have found that preserving lymph nodes during cancer surgery could dramatically improve how patients respond to immunotherapy. The research shows that lymph nodes are essential for training and sustaining cancer-fighting T cells. Removing them may unintentionally weaken the immune response, while keeping them intact could help unlock stronger, longer-lasting treatments.
  •  

Capacity to Invest Effort as a Predictor of Preference for Digital Mental Health Interventions Over Psychotherapy: Cross-Sectional Study Using an Ecological Digital Screening Tool

Background: Research typically shows a higher preference for professionally-led face-to-face mental health interventions over digital ones. It remains unclear in which circumstances digital self-help tools are preferred. To address this gap, it is important to examine user characteristics that may help predict when digital interventions are more desirable, ultimately guiding their design to enhance engagement and appeal. Objective: To examine how distress severity and capacity to invest effort relate to intervention preferences, using an ecological assessment of individuals who seek to receive feedback on their mental health. Methods: A comprehensive digital mental health screening tool providing automated feedback was developed and advertised on social media. The sample comprised 684 adult participants aged 18-82 who opted to complete the screening to receive feedback on their mental health state. Participants completed questionnaires measuring general psychological distress, depression, generalized anxiety and demographics. Kessler Psychological Distress Scale–6 was used as the primary measure for distress. Participants were also presented with questions measuring capacity to invest effort and preferences for a professional vs digital self-help tools and for psychotherapy vs a mobile application. The effectiveness of distress, capacity to invest effort, and background characteristics in predicting preferences (a professional vs digital self-help tools; psychotherapy vs a mobile application) was examined using hierarchical linear regressions. The distributions of dichotomized preferences were plotted against distress and capacity to invest effort for transparent visualization. Results: A hierarchical linear regression found that distress, capacity to invest, and currently being in psychotherapy significantly predicted preference for a professional vs digital self-help tools. Distress (β=.25, 95% CI .18 to .32, P<.001 and capacity to invest effort ci .16 .30 p were the strongest predictors with similar effect size. model explained of variance in preference uniquely contributing most distressed participants low preferred digital self-help tools whereas high favored a professional. results obtained when using phq-4 as an alternative distress measure. remained significant .10 .26 predicting for psychotherapy vs mobile application while was not .05 conclusions: this study highlights that interventions is driven by reduced intervention. attempts reduce mental health treatment gap through should focus on optimizing elicited users improve desirability engagement.>
  •  
❌