❌

Reading view

PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading

arXiv:2510.22242v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a benchmark that systematically evaluates LLMs across four key research tasks: citation retrieval, content extraction, paper discovery, and claim verification. We evaluate GPT-4o, GPT-5, and Gemini-2.5-Flash under realistic usage conditions-via web interfaces where search operations are opaque to the user. Through controlled experiments, we find consistent reliability failures: citation retrieval fails in 48-98% of multi-reference queries, section-specific content extraction fails in 72-91% of cases, and topical paper discovery yields F1 scores below 0.32, missing over 60% of relevant literature. Further human analysis attributes these failures to the uncontrolled expansion of retrieved context and the tendency of LLMs to prioritize semantically relevant text over task instructions. Across basic tasks, the LLMs display distinct failure behaviors: ChatGPT often withholds responses rather than risk errors, whereas Gemini produces fluent but fabricated answers. To address these issues, we develop lightweight reliability classifiers trained on PaperAsk data to identify unreliable outputs. PaperAsk provides a reproducible and diagnostic framework for advancing the reliability evaluation of LLM-based scholarly assistance systems.
  •  

Do we need AI guardians to protect us from health information overload?

npj Digital Medicine, Published online: 27 October 2025; doi:10.1038/s41746-025-02093-0

The rise of digital health technologies has provided individuals with unprecedented access to biometric data and health insights. However, excess monitoring may contribute to fatigue, anxiety, and information overload, sometimes reducing engagement and worsening outcomes. This article explores how artificial intelligence-enabled assistants might help address this challenge by filtering, contextualizing, and personalizing health information, potentially supporting informed self-management while mitigating some unintended harms of digital health technologies.
  •  

Preclinical application of a CD155 targeting chimeric antigen receptor T cell therapy for digestive system cancers

Oncogene, Published online: 01 March 2025; doi:10.1038/s41388-025-03322-2

Preclinical application of a CD155 targeting chimeric antigen receptor T cell therapy for digestive system cancers
  •  

Tumor microenvironment and drug resistance in lung adenocarcinoma: molecular mechanisms, prognostic implications, and therapeutic strategies

Discov Oncol. 2025 Feb 25;16(1):238. doi: 10.1007/s12672-025-01981-x.

ABSTRACT

The fight against lung adenocarcinoma (LUAD) is challenged by tumor microenvironment (TME)-mediated drug resistance, which limits effective treatment. This study examines the LUAD TME and identifies four distinct subtypes through multi-omics profiling: immune-rich, immune-exhausted, stromal-dominant, and TME-desert. Each subtype has unique molecular features, tumor diversity, and links to clinical outcomes. Immune-rich subtypes respond better to immune checkpoint inhibitors, while stromal-dominant and TME-desert subtypes show resistance to treatment and poor prognosis. Molecular analysis uncovers subtype-specific mutations, chromosomal instability, and altered signaling pathways, pointing to potential therapeutic targets. In silico drug screening identifies promising treatments for resistant subtypes. These findings, validated in independent cohorts, highlight the critical role of the TME in drug resistance and treatment response, providing insights for personalized treatment strategies in LUAD.

PMID:40000527 | PMC:PMC11861463 | DOI:10.1007/s12672-025-01981-x

  •  
❌