❌

Reading view

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p
  •  

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p
  •  

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.
  •  

Targeting spermine metabolism to overcome immunotherapy resistance in pancreatic cancer

Nat Commun. 2025 Aug 22;16(1):7827. doi: 10.1038/s41467-025-63146-2.

ABSTRACT

While dysregulation of polyamine metabolism is frequently observed in cancer, it is unknown how polyamines alter the tumor microenvironment (TME) and contribute to therapeutic resistance. Analysis of polyamines in the plasma of pancreatic cancer patients reveals that spermine levels are significantly elevated and correlate with poor prognosis. Using a multi-omics approach, we identify Serpinb9 as a vulnerability in spermine metabolism in pancreatic cancer. Serpinb9, a serine protease inhibitor, directly interacts with spermine synthase (SMS), impeding its lysosome-mediated degradation and thereby augmenting spermine production and secretion. Mechanistically, the accumulation of spermine in the TME alters the metabolic landscape of immune cells, promoting CD8+ T cell dysfunction and pro-tumor polarization of macrophages, thus creating an immunosuppressive microenvironment. Small peptides that disrupt the Serpinb9-SMS interaction significantly enhance the efficacy of immune checkpoint blockade therapy. Together, our findings suggest that targeting spermine metabolism is a promising strategy to improve pancreatic cancer immunotherapy.

PMID:40846845 | PMC:PMC12373741 | DOI:10.1038/s41467-025-63146-2

  •  

Targeting MFAP5 in cancer-associated fibroblasts sensitizes pancreatic cancer to PD-L1-based immunochemotherapy via remodeling the matrix

Oncogene, Published online: 08 May 2023; doi:10.1038/s41388-023-02711-9

Targeting MFAP5 in cancer-associated fibroblasts sensitizes pancreatic cancer to PD-L1-based immunochemotherapy via remodeling the matrix
  •  
❌