❌

Normal view

  • ✇MIT Technology Review
  • Building a high performance data and AI organization (2nd edition) MIT Technology Review Insights
    Four years is a lifetime when it comes to artificial intelligence. Since the first edition of this study was published in 2021, AI’s capabilities have been advancing at speed, and the advances have not slowed since generative AI’s breakthrough. For example, multimodality— the ability to process information not only as text but also as audio, video, and other unstructured formats—is becoming a common feature of AI models. AI’s capacity to reason and act autonomously has also grown, and organizati
     

Building a high performance data and AI organization (2nd edition)

Four years is a lifetime when it comes to artificial intelligence. Since the first edition of this study was published in 2021, AI’s capabilities have been advancing at speed, and the advances have not slowed since generative AI’s breakthrough. For example, multimodality— the ability to process information not only as text but also as audio, video, and other unstructured formats—is becoming a common feature of AI models. AI’s capacity to reason and act autonomously has also grown, and organizations are now starting to work with AI agents that can do just that.

Amid all the change, there remains a constant: the quality of an AI model’s outputs is only ever as good as the data
that feeds it. Data management technologies and practices have also been advancing, but the second edition of this study suggests that most organizations are not leveraging those fast enough to keep up with AI’s development. As a result of that and other hindrances, relatively few organizations are delivering the desired business results from their AI strategy. No more than 2% of senior executives we surveyed rate their organizations highly in terms of delivering results from AI.

To determine the extent to which organizational data performance has improved as generative AI and other AI advances have taken hold, MIT Technology Review Insights surveyed 800 senior data and technology executives. We also conducted in-depth interviews with 15 technology and business leaders.

Key findings from the report include the following:

• Few data teams are keeping pace with AI. Organizations are doing no better today at delivering on data strategy than in pre-generative AI days. Among those surveyed in 2025, 12% are self-assessed data “high achievers” compared with 13% in 2021. Shortages of skilled talent remain a constraint, but teams also struggle with accessing fresh data, tracing lineage, and dealing with security complexity—important requirements for AI success.

• Partly as a result, AI is not fully firing yet. There are even fewer “high achievers” when it comes to AI. Just 2% of respondents rate their organizations’ AI performance highly today in terms of delivering measurable business results. In fact, most are still struggling to scale generative AI. While two thirds have deployed it, only 7% have done so widely.

Download the report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.

Advancing Non-Small-Cell Lung Cancer Management Through Multi-Omics Integration: Insights from Genomics, Metabolomics, and Radiomics

Diagnostics (Basel). 2025 Oct 14;15(20):2586. doi: 10.3390/diagnostics15202586.

ABSTRACT

The integration of multi-omics technologies is transforming the landscape of cancer management, offering unprecedented insights into tumor biology, early diagnosis, and personalized therapy. This review provides a comprehensive overview of the current state of omics approaches, with a particular focus on the application of genomics, NMR-based metabolomics, and radiomics in non-small cell lung cancer (NSCLC). Genomics currently represents one of the most established omics technologies in oncology, as it enables the identification of genetic alterations that drive tumor initiation, progression, and therapeutic response. Interestingly, genomic analyses have revealed that many tumors harbor mutations in genes encoding metabolic enzymes, thus establishing a tight connection between genomics and tumor metabolism. In parallel, metabolomics profiling-by capturing the metabolic phenotype of tumors-has, in recent years, identified specific biomarkers associated with tumor burden, progression, and prognosis. Such findings have catalyzed growing interest in metabolomics as a complementary approach to better characterize cancer biology and discover novel diagnostic and therapeutic targets. Moreover, radiomics, through the extraction of quantitative features from standard imaging modalities, captures tumor heterogeneity and contributes predictive information on tumor biology, treatment response, and clinical outcomes. As a non-invasive and widely available technique, radiomics has the potential to support longitudinal monitoring and individualized treatment planning. Both metabolomics and radiomics, when integrated with genomic data, could support a more comprehensive understanding of NSCLC and pave the way for the development of non-invasive, predictive models and personalized therapeutic strategies. In addition, we explore the specific contributions of these technologies in enhancing clinical decision-making for lung cancer patients, with particular attention to their potential in early diagnosis, treatment selection, and real-time monitoring.

PMID:41153258 | DOI:10.3390/diagnostics15202586

Prospective proteomics for discovering biomarkers in lung adenocarcinoma: a literature review

Transl Cancer Res. 2025 Sep 30;14(9):6102-6117. doi: 10.21037/tcr-2025-1092. Epub 2025 Sep 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung adenocarcinoma (LUAD), as the main subtype of non-small cell lung cancer (NSCLC), faces clinical challenges including molecular heterogeneity, late diagnosis, and aggressive growth, leading to a low 5-year survival rate. Biomarkers are critical for early detection, accurate differentiation of benign/malignant lesions, and guiding personalized treatment strategies. Proteomic technologies using liquid biopsy show potential by analyzing protein changes and post-translational modifications (PTMs) to identify novel biomarkers and unravel cancer mechanisms. This review examines proteomic advances in LUAD, compares platform strengths, lists validated protein markers, and discusses challenges like specificity and regulations. It aims to develop a precision medicine framework by integrating multi-omics data for improved diagnosis and treatment.

METHODS: This study conducted a literature review by searching the PubMed and Web of Science databases for original articles written in English from 2002 to 2025, using the keywords "lung adenocarcinoma" OR "LUAD" AND "biomarkers" AND "proteomics" OR "SomaScan" OR "spatial proteomics" to identify the latest research findings in the field of proteomics technology and LUAD biomarkers. The included studies mainly focused on the current landscape of biomarkers in the diagnosis, treatment, and prognosis of LUAD.

KEY CONTENT AND FINDINGS: This review discusses high-throughput methods for comprehensive protein profiling in accessible biospecimens (tissues, blood, urine) to identify biomarkers for LUAD. We systematically evaluate emerging proteomic strategies, including mass spectrometry (MS), proximity extension assays (PEAs), spatial proteomics techniques, and SomaScan platforms-coupled with innovative computational frameworks have revolutionized biomarkers discovery and their translational potential in developing precision diagnostics and targeted therapies. Additionally, the review addresses challenges in integrating proteomics with genomics, transcriptomics, and metabolomics, offering new methodologies and expanding research in life sciences. As technological advancements continue, it is anticipated that more potential biomarkers will be conducted to validate the broader application in LUAD treatment, addressing early-stage disease complexities and aiding in selecting more effective treatment strategies.

CONCLUSIONS: By synthesizing cutting-edge evidence on proteome-driven LUAD biomarkers, this review elucidates actionable strategies to refine early detection protocols and mechanism-informed personalized treatment frameworks, directly advancing precision oncology initiatives for this prevalent malignancy through biomarker-guided clinical decision-making and multi-omics integration.

PMID:41158224 | PMC:PMC12554480 | DOI:10.21037/tcr-2025-1092

  • ✇STAT
  • STAT+: Natera, known for spotting cancer recurrence, wades into early detection Elaine Chen
    Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox. Good morning. It seems everyone I know has been getting sick lately — hope you are all taking care of yourselves! Onto the news today. BridgeBio notches another Phase 3 win BridgeBio said this morning that its investigational drug succeeded in a late-stage trial of patients with autosomal dominant hypocalcemia type 1, a rare genetic condition that causes low calciu
     

STAT+: Natera, known for spotting cancer recurrence, wades into early detection

29 October 2025 at 21:26

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Good morning. It seems everyone I know has been getting sick lately — hope you are all taking care of yourselves! Onto the news today.

BridgeBio notches another Phase 3 win

BridgeBio said this morning that its investigational drug succeeded in a late-stage trial of patients with autosomal dominant hypocalcemia type 1, a rare genetic condition that causes low calcium levels in the blood.

Continue to STAT+ to read the full story…

© Adobe

Improving Recruitment Into Research Studies via Electronically Collected Patient-Entered Data: Mixed Methods Study

Background: Patient recruitment remains a critical challenge in clinical research. Although the integration of electronically collected patient-entered data within clinical practices enables innovative recruitment approaches, existing methods present challenges such as increased patient burden and potential violation of autonomy. A more nuanced approach involves identifying patient attributes associated with higher propensity for research participation, enabling research teams to efficiently prioritize outreach efforts. Objective: This study aims to (1) develop patient-reported questions reflecting perceptions about research participation and (2) determine whether patient responses are predictive of interest in joining a precision medicine registry. Methods: This mixed methods study used an exploratory sequential design in 2 phases. Phase 1 involved cognitive interviews with 32 patients recruited through the Cleveland Clinic Healthcare Partners program to develop “research perception” questions. Participants evaluated 9 candidate questions that were based on a literature review of research participation factors. Three questions were selected for implementation. Phase 2 was a cross-sectional cohort study incorporating these 3 questions into routine electronic questionnaires completed by primary care patients through the patient portal. The study population included 1077 patients who completed both “research perception” and “research recruitment” questions between August 2018 and April 2019. Diagnostic accuracy was assessed using receiver operating characteristic curve analysis, and multivariable logistic regression models evaluated associations while adjusting for demographic and health factors. Results: Phase 1 revealed strong research support among participants, with 97% (31/32) agreeing that research should be part of the institution’s mission and 100% (32/32) affirming that research enhances patient care. Phase 2 included 1077 patients (mean age 48.3, SD 16.3 years; 625/1065 female, 58.68%; 661/1005 White, 65.77%), of whom 278 (25.8%) expressed interest in being contacted about the precision medicine registry. Patients expressing interest were older and had worse self-reported health, more depressive symptoms, and greater social needs. “Strongly agree” and “very important” responses to any “research perception” question were significantly associated with study interest, with adjusted odds ratios ranging from 6.36 (95% CI 2.77-14.6) to 17.6 (95% CI 5.08-61.1; P<.001). The “research perception” questions demonstrated high sensitivity (>80%) but limited specificity (24%-31%). Conclusions: Patient-reported questions assessing research participation likelihood can help identify patients more likely to enroll in clinical studies. This approach enables effective recruitment prioritization while preserving patient autonomy and reducing patient burden. High sensitivity makes these questions valuable as screening tools, although limited specificity suggests use for prioritizing rather than excluding participants. Further validation across different trial types and populations is warranted.

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

arXiv:2510.24161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.

From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning

arXiv:2510.24528v1 Announce Type: new Abstract: The capability of in-context learning (ICL) enables large language models (LLMs) to perform novel tasks without parameter updates by conditioning on a few input-output examples. However, collecting high-quality examples for new or challenging tasks can be costly and labor-intensive. In this work, we propose a cost-efficient two-stage pipeline that reduces reliance on LLMs for data labeling. Our approach first leverages readily available cross-task examples to prompt an LLM and pseudo-label a small set of target task instances. We then introduce a graph-based label propagation method that spreads label information to the remaining target examples without additional LLM queries. The resulting fully pseudo-labeled dataset is used to construct in-task demonstrations for ICL. This pipeline combines the flexibility of cross-task supervision with the scalability of LLM-free propagation. Experiments across five tasks demonstrate that our method achieves strong performance while lowering labeling costs.

Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives

arXiv:2510.24551v1 Announce Type: new Abstract: Generative Artificial Intelligence (GenAI) is taking the world by storm. It promises transformative opportunities for advancing and disrupting existing practices, including healthcare. From large language models (LLMs) for clinical note synthesis and conversational assistance to multimodal systems that integrate medical imaging, electronic health records, and genomic data for decision support, GenAI is transforming the practice of medicine and the delivery of healthcare, such as diagnosis and personalized treatments, with great potential in reducing the cognitive burden on clinicians, thereby improving overall healthcare delivery. However, GenAI deployment in healthcare requires an in-depth understanding of healthcare tasks and what can and cannot be achieved. In this paper, we propose a data-centric paradigm in the design and deployment of GenAI systems for healthcare. Specifically, we reposition the data life cycle by making the medical data ecosystem as the foundational substrate for generative healthcare systems. This ecosystem is designed to sustainably support the integration, representation, and retrieval of diverse medical data and knowledge. With effective and efficient data processing pipelines, such as semantic vector search and contextual querying, it enables GenAI-powered operations for upstream model components and downstream clinical applications. Ultimately, it not only supplies foundation models with high-quality, multimodal data for large-scale pretraining and domain-specific fine-tuning, but also serves as a knowledge retrieval backend to support task-specific inference via the agentic layer. The ecosystem enables the deployment of GenAI for high-quality and effective healthcare delivery.

Genotype-Phenotype Integration through Machine Learning and Personalized Gene Regulatory Networks for Cancer Metastasis Prediction

arXiv:2510.23620v1 Announce Type: cross Abstract: Metastasis is the leading cause of cancer-related mortality, yet most predictive models rely on shallow architectures and neglect patient-specific regulatory mechanisms. Here, we integrate classical machine learning and deep learning to predict metastatic potential across multiple cancer types. Gene expression profiles from the Cancer Cell Line Encyclopedia were combined with a transcription factor-target prior from DoRothEA, focusing on nine metastasis-associated regulators. After selecting differential genes using the Kruskal-Wallis test, ElasticNet, Random Forest, and XGBoost models were trained for benchmarking. Personalized gene regulatory networks were then constructed using PANDA and LIONESS and analyzed through a graph attention neural network (GATv2) to learn topological and expression-based representations. While XGBoost achieved the highest AUROC (0.7051), the GNN captured non-linear regulatory dependencies at the patient level. These results demonstrate that combining traditional machine learning with graph-based deep learning enables a scalable and interpretable framework for metastasis risk prediction in precision oncology.

Integrating Genomics into Multimodal EHR Foundation Models

arXiv:2510.23639v1 Announce Type: cross Abstract: This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health profiles. Leveraging the extensive and diverse data from the All of Us (AoU) Research Program, this multimodal framework aims to learn complex relationships between clinical data and genetic predispositions. The methodology extends advancements in generative AI to the EHR foundation model space, enhancing predictive capabilities and interpretability. Evaluation on AoU data demonstrates the model's predictive value for the onset of various conditions, particularly Type 2 Diabetes (T2D), and illustrates the interplay between PRS and EHR data. The work also explores transfer learning for custom classification tasks, showcasing the architecture's versatility and efficiency. This approach is pivotal for unlocking new insights into disease prediction, proactive health management, risk stratification, and personalized treatment strategies, laying the groundwork for more personalized, equitable, and actionable real-world evidence generation in healthcare.

Quanvolutional Neural Networks for Pneumonia Detection: An Efficient Quantum-Assisted Feature Extraction Paradigm

arXiv:2510.23660v1 Announce Type: cross Abstract: Pneumonia poses a significant global health challenge, demanding accurate and timely diagnosis. While deep learning, particularly Convolutional Neural Networks (CNNs), has shown promise in medical image analysis for pneumonia detection, CNNs often suffer from high computational costs, limitations in feature representation, and challenges in generalizing from smaller datasets. To address these limitations, we explore the application of Quanvolutional Neural Networks (QNNs), leveraging quantum computing for enhanced feature extraction. This paper introduces a novel hybrid quantum-classical model for pneumonia detection using the PneumoniaMNIST dataset. Our approach utilizes a quanvolutional layer with a parameterized quantum circuit (PQC) to process 2x2 image patches, employing rotational Y-gates for data encoding and entangling layers to generate non-classical feature representations. These quantum-extracted features are then fed into a classical neural network for classification. Experimental results demonstrate that the proposed QNN achieves a higher validation accuracy of 83.33 percent compared to a comparable classical CNN which achieves 73.33 percent. This enhanced convergence and sample efficiency highlight the potential of QNNs for medical image analysis, particularly in scenarios with limited labeled data. This research lays the foundation for integrating quantum computing into deep-learning-driven medical diagnostic systems, offering a computationally efficient alternative to traditional approaches.

Closing Gaps: An Imputation Analysis of ICU Vital Signs

arXiv:2510.24217v1 Announce Type: cross Abstract: As more Intensive Care Unit (ICU) data becomes available, the interest in developing clinical prediction models to improve healthcare protocols increases. However, the lack of data quality still hinders clinical prediction using Machine Learning (ML). Many vital sign measurements, such as heart rate, contain sizeable missing segments, leaving gaps in the data that could negatively impact prediction performance. Previous works have introduced numerous time-series imputation techniques. Nevertheless, more comprehensive work is needed to compare a representative set of methods for imputing ICU vital signs and determine the best practice. In reality, ad-hoc imputation techniques that could decrease prediction accuracy, like zero imputation, are still used. In this work, we compare established imputation techniques to guide researchers in improving the performance of clinical prediction models by selecting the most accurate imputation technique. We introduce an extensible and reusable benchmark with currently 15 imputation and 4 amputation methods, created for benchmarking on major ICU datasets. We hope to provide a comparative basis and facilitate further ML development to bring more models into clinical practice.

Dynamic Monitoring of Recurrent Ovarian Cancer Using Serial ctDNA: A Real-World Case Series

Curr Oncol. 2025 Oct 21;32(10):585. doi: 10.3390/curroncol32100585.

ABSTRACT

Recurrent ovarian cancer (OC) is challenging to detect early using current methods like CA-125 and imaging. Circulating tumor DNA (ctDNA) may improve disease monitoring. Here, we assess the real-world clinical utility of serial ctDNA analyses in patients with recurrent OC. We analyzed serial plasma samples (N = 23) from six patients with recurrent OC using a tumor-informed next-generation sequencing assay targeting 68 cancer-related genes developed at the University of Washington. ctDNA variant allele frequencies (VAFs) were correlated with CA-125 levels, radiographic findings, and clinical outcomes. ctDNA levels generally reflected clinical status, accurately mirroring disease progression and therapeutic response. In one patient, rising ctDNA preceded clinical recurrence by four months, despite normal CA-125 and imaging, highlighting its potential advantage. Conversely, some patients exhibited clinical progression with undetectable ctDNA, indicating limitations in assay sensitivity, biological factors, or metastatic sites (e.g., brain metastases). ctDNA and CA-125 showed complementary value in most cases, suggesting potential combined use in clinical monitoring. Our findings demonstrate that ctDNA is a promising biomarker to complement existing monitoring approaches for recurrent OC. In some cases, capable of predicting relapse and treatment response ahead of current clinical indicators. However, identified discordances underscore technical and biological challenges that warrant further investigation. Larger prospective studies are necessary to refine ctDNA's clinical utility and integration into personalized OC care.

PMID:41149505 | PMC:PMC12563156 | DOI:10.3390/curroncol32100585

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.

The Confidence Paradox: Can LLM Know When It's Wrong

arXiv:2506.23464v2 Announce Type: replace Abstract: Document Visual Question Answering (DocVQA) models often produce overconfident or ethically misaligned responses, especially under uncertainty. Existing models like LayoutLMv3, UDOP, and DONUT focus on accuracy but lack ethical calibration. We propose HonestVQA, a model-agnostic, self-supervised framework that aligns model confidence with correctness using weighted loss and contrastive learning. We introduce two new metrics Honesty Score (H-Score) and Ethical Confidence Index (ECI)-to evaluate ethical alignment. HonestVQA improves accuracy and F1 by up to 4.3% across SpDocVQA, InfographicsVQA, and SROIE datasets, while reducing overconfidence. It also generalizes well across domains, achieving 78.9% accuracy and 76.1% F1-score.

The Role of Omentin in Gastrointestinal Cancer: Diagnostic, Prognostic, and Therapeutic Perspectives

Metabolites. 2025 Sep 30;15(10):649. doi: 10.3390/metabo15100649.

ABSTRACT

Background/Objectives: Omentin, also known as intelectin-1, is a secreted adipokine with anti-inflammatory, insulin-sensitizing, and immune-modulatory functions, primarily expressed in visceral adipose tissue. While omentin has been associated with favorable metabolic outcomes, its role in cancer pathogenesis appears context-dependent and remains poorly understood. This review investigates the biological functions, expression patterns, and clinical relevance of omentin across gastrointestinal malignancies. Methods: A comprehensive review of the literature was conducted using PubMed, Scopus, and Web of Science up to August 2025 to evaluate the role of omentin in gastrointestinal cancers. Both preclinical and clinical studies evaluating omentin, its analogues and omentin-enhancing agents in gastric, colorectal, hepatic, pancreatic, and esophageal cancers were included. Results: Omentin exhibits anti-proliferative, anti-inflammatory, and anti-angiogenic effects within the tumor microenvironment in several GI malignancies. However, evidence also indicates a dual role. High intratumoral omentin expression correlates with improved prognosis in colorectal, gastric, and hepatic cancers; in contrast, elevated circulating levels-particularly in colorectal and pancreatic cancers-have been paradoxically associated with increased cancer risk and poor outcomes. Mechanistically, omentin modulates PI3K/Akt, NF-κB, AMPK, and oxidative stress pathways, and interacts with TMEM207. However, most available studies are small-scale and heterogeneous, with methodological inconsistencies and limited multi-omics integration, leaving major knowledge gaps. Conclusions: This review highlights omentin's distinct systemic and local roles across GI cancers, underscoring its translational implications. Omentin emerges as a promising but context-dependent biomarker and therapeutic target, with future research needed to address heterogeneity, standardize assays, and validate its clinical utility in large-scale prospective studies.

PMID:41149627 | PMC:PMC12566161 | DOI:10.3390/metabo15100649

A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications

arXiv:2510.16724v2 Announce Type: replace Abstract: The advent of large language models (LLMs) has transformed information access and reasoning through open-ended natural language interaction. However, LLMs remain limited by static knowledge, factual hallucinations, and the inability to retrieve real-time or domain-specific information. Retrieval-Augmented Generation (RAG) mitigates these issues by grounding model outputs in external evidence, but traditional RAG pipelines are often single turn and heuristic, lacking adaptive control over retrieval and reasoning. Recent advances in agentic search address these limitations by enabling LLMs to plan, retrieve, and reflect through multi-step interaction with search environments. Within this paradigm, reinforcement learning (RL) offers a powerful mechanism for adaptive and self-improving search behavior. This survey provides the first comprehensive overview of \emph{RL-based agentic search}, organizing the emerging field along three complementary dimensions: (i) What RL is for (functional roles), (ii) How RL is used (optimization strategies), and (iii) Where RL is applied (scope of optimization). We summarize representative methods, evaluation protocols, and applications, and discuss open challenges and future directions toward building reliable and scalable RL driven agentic search systems. We hope this survey will inspire future research on the integration of RL and agentic search. Our repository is available at https://github.com/ventr1c/Awesome-RL-based-Agentic-Search-Papers.
❌