Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications
arXiv:2510.25908v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated transformative potential in scientific research, yet their deployment in high-stakes contexts raises significant trustworthiness concerns. Here, we introduce SciTrust 2.0, a comprehensive framework for evaluating LLM trustworthiness in scientific applications across four dimensions: truthfulness, adversarial robustness, scientific safety, and scientific ethics. Our framework incorporates novel, open-e
-
cs.AI, q-bio.NC updates on arXiv.org
-
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
arXiv:2510.25992v1 Announce Type: cross Abstract: Large Language Models (LLMs) often struggle with problems that require multi-step reasoning. For small-scale open-source models, Reinforcement Learning with Verifiable Rewards (RLVR) fails when correct solutions are rarely sampled even after many attempts, while Supervised Fine-Tuning (SFT) tends to overfit long demonstrations through rigid token-by-token imitation. To address this gap, we propose Supervised Reinforcement Learning (SRL), a frame
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
arXiv:2510.26151v1 Announce Type: cross Abstract: Large annotated datasets are essential for training robust Computer-Aided Diagnosis (CAD) models for breast cancer detection or risk prediction. However, acquiring such datasets with fine-detailed annotation is both costly and time-consuming. Vision-Language Models (VLMs), such as CLIP, which are pre-trained on large image-text pairs, offer a promising solution by enhancing robustness and data efficiency in medical imaging tasks. This paper intr
MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
-
cs.AI, q-bio.NC updates on arXiv.org
-
Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
arXiv:2410.14879v4 Announce Type: replace-cross Abstract: Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate periods of raw sensor data to model momentary behaviors, (e.g., physical activity recognition), there still remains a significant gap in the translation of these sensing streams into meaningful, high-l
Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
-
Nature - Issue - nature.com science feeds
-
Multi-omic profiling reveals age-related immune dynamics in healthy adults
Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.
Multi-omic profiling reveals age-related immune dynamics in healthy adults
Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5
This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.-
cs.AI, q-bio.NC updates on arXiv.org
-
From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
arXiv:2510.24528v1 Announce Type: new Abstract: The capability of in-context learning (ICL) enables large language models (LLMs) to perform novel tasks without parameter updates by conditioning on a few input-output examples. However, collecting high-quality examples for new or challenging tasks can be costly and labor-intensive. In this work, we propose a cost-efficient two-stage pipeline that reduces reliance on LLMs for data labeling. Our approach first leverages readily available cross-task
From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Tongyi DeepResearch Technical Report
arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data syn
Tongyi DeepResearch Technical Report
-
cs.AI, q-bio.NC updates on arXiv.org
-
Impact and Implications of Generative AI for Enterprise Architects in Agile Environments: A Systematic Literature Review
arXiv:2510.22003v1 Announce Type: cross Abstract: Generative AI (GenAI) is reshaping enterprise architecture work in agile software organizations, yet evidence on its effects remains scattered. We report a systematic literature review (SLR), following established SLR protocols of Kitchenham and PRISMA, of 1,697 records, yielding 33 studies across enterprise, solution, domain, business, and IT architect roles. GenAI most consistently supports (i) design ideation and trade-off exploration; (ii) r
Impact and Implications of Generative AI for Enterprise Architects in Agile Environments: A Systematic Literature Review
-
cs.AI, q-bio.NC updates on arXiv.org
-
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
arXiv:2510.22242v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a benchmark that systematically evaluates LLMs across four key research tasks: citation retrieval, content extraction, paper discovery, and claim verification. We evaluate GPT-4o, GPT-5, and Gemini-2.5-Flash under realistic usage conditions-via web interfaces where search o
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
-
cs.AI, q-bio.NC updates on arXiv.org
-
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
arXiv:2510.22620v1 Announce Type: cross Abstract: AI agents powered by large language models (LLMs) are being deployed at scale, yet we lack a systematic understanding of how the choice of backbone LLM affects agent security. The non-deterministic sequential nature of AI agents complicates security modeling, while the integration of traditional software with AI components entangles novel LLM vulnerabilities with conventional security risks. Existing frameworks only partially address these chall
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
arXiv:2510.23241v1 Announce Type: cross Abstract: In this work, we introduce Progressive Growing of Patch Size, an automatic curriculum learning approach for 3D medical image segmentation. Our approach progressively increases the patch size during model training, resulting in an improved class balance for smaller patch sizes and accelerated convergence of the training process. We evaluate our curriculum approach in two settings: a resource-efficient mode and a performance mode, both regarding D
Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
-
cs.AI, q-bio.NC updates on arXiv.org
-
DataRater: Meta-Learned Dataset Curation
arXiv:2505.17895v2 Announce Type: replace-cross Abstract: The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics. An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training. This type of meta-learning co
DataRater: Meta-Learned Dataset Curation
-
cs.AI, q-bio.NC updates on arXiv.org
-
CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
arXiv:2505.18087v2 Announce Type: replace-cross Abstract: Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer, offering limited insight into whether models engage in clinically meaningful reasoning. To address this, we present CheXStruct and CXReasonBench, a structured pipeline and benchmark built on the public
CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays
-
cs.AI, q-bio.NC updates on arXiv.org
-
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
arXiv:2506.06522v2 Announce Type: replace-cross Abstract: Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible to the public, with limited information about their construction process. This lack of transparency has motivated the recent development of op
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
-
cs.AI, q-bio.NC updates on arXiv.org
-
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
arXiv:2507.05177v3 Announce Type: replace-cross Abstract: Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expressive responses. However, the most powerful empathetic LSLMs are increasingly closed off, leaving the crucial details about the architecture, data and development opaque to researchers. Given the critical need for transparent research into the LSLMs and empath
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
-
npj Digital Medicine
-
Do we need AI guardians to protect us from health information overload?
npj Digital Medicine, Published online: 27 October 2025; doi:10.1038/s41746-025-02093-0The rise of digital health technologies has provided individuals with unprecedented access to biometric data and health insights. However, excess monitoring may contribute to fatigue, anxiety, and information overload, sometimes reducing engagement and worsening outcomes. This article explores how artificial intelligence-enabled assistants might help address this challenge by filtering, contextualizing, and pe
Do we need AI guardians to protect us from health information overload?
npj Digital Medicine, Published online: 27 October 2025; doi:10.1038/s41746-025-02093-0
The rise of digital health technologies has provided individuals with unprecedented access to biometric data and health insights. However, excess monitoring may contribute to fatigue, anxiety, and information overload, sometimes reducing engagement and worsening outcomes. This article explores how artificial intelligence-enabled assistants might help address this challenge by filtering, contextualizing, and personalizing health information, potentially supporting informed self-management while mitigating some unintended harms of digital health technologies.-
cs.AI, q-bio.NC updates on arXiv.org
-
When Models Outthink Their Safety: Mitigating Self-Jailbreak in Large Reasoning Models with Chain-of-Guardrails
arXiv:2510.21285v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) demonstrate remarkable capabilities on complex reasoning tasks but remain vulnerable to severe safety risks, including harmful content generation and jailbreak attacks. Existing mitigation strategies rely on injecting heuristic safety signals during training, which often suppress reasoning ability and fail to resolve the safety-reasoning trade-off. To systematically investigate this issue, we analyze the reasoning tra
When Models Outthink Their Safety: Mitigating Self-Jailbreak in Large Reasoning Models with Chain-of-Guardrails
-
cs.AI, q-bio.NC updates on arXiv.org
-
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
arXiv:2510.21652v1 Announce Type: new Abstract: AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging from general-purpose "deep research" systems to specialized science-specific agents, such as AI Scientist and AIGS. Rigorous evaluation of these agents is critical for progress. Yet existing benchmarks fall short
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
-
Omics In Lung
-
Biomarkers for non-small cell lung cancer risk using multi-omics approaches: a nested case-control study
Transl Lung Cancer Res. 2025 Sep 30;14(9):3645-3658. doi: 10.21037/tlcr-2025-603. Epub 2025 Sep 25.ABSTRACTBACKGROUND: Lung cancer poses a major public health challenge, accounting for the highest cancer-related mortality worldwide. This study aimed to identify non-invasive biomarkers for the early detection of non-small cell lung cancer (NSCLC) risk.METHODS: We randomly selected 150 incident NSCLC cases during follow-up from the Korean Cancer Prevention Study-II. Controls (n=150) were matched t
Biomarkers for non-small cell lung cancer risk using multi-omics approaches: a nested case-control study
Transl Lung Cancer Res. 2025 Sep 30;14(9):3645-3658. doi: 10.21037/tlcr-2025-603. Epub 2025 Sep 25.
ABSTRACT
BACKGROUND: Lung cancer poses a major public health challenge, accounting for the highest cancer-related mortality worldwide. This study aimed to identify non-invasive biomarkers for the early detection of non-small cell lung cancer (NSCLC) risk.
METHODS: We randomly selected 150 incident NSCLC cases during follow-up from the Korean Cancer Prevention Study-II. Controls (n=150) were matched to cases by age, gender, and the time of blood collection. Non-targeted metabolite screening by ultra-high-performance liquid chromatography (UHPLC)/mass spectrometry (MS) was conducted on the pre-diagnostic biological samples. The 11 reported lung cancer-associated single-nucleotide polymorphisms (SNPs) in Koreans were extracted from DNA genotyping data of the study population. Metabolite markers related to NSCLC risk were identified through clustering using hierarchical density-based spatial clustering of applications with noise. The associations between smoking, dietary factors, and NSCLC were also examined.
RESULTS: Six discriminative serum metabolites were identified as having an association with NSCLC incidence. Notably, the relationship between specific metabolite levels and NSCLC risk differed by rs7086803 genotype. Smoking status and occupational exposures appear to influence specific metabolite profiles, while dietary vegetable intake may modulate the risk of NSCLC among smokers.
CONCLUSIONS: The meaningful biomarkers revealed in the current research could be used to enhance the predictive ability for NSCLC risk. Furthermore, we suggest that the protective role of dietary vegetables against NSCLC may be attenuated or absent in smokers.
PMID:41133005 | PMC:PMC12541849 | DOI:10.21037/tlcr-2025-603
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI PB: A Grounded Generative Agent for Personalized Investment Insights
arXiv:2510.20099v1 Announce Type: new Abstract: We present AI PB, a production-scale generative agent deployed in real retail finance. Unlike reactive chatbots that answer queries passively, AI PB proactively generates grounded, compliant, and user-specific investment insights. It integrates (i) a component-based orchestration layer that deterministically routes between internal and external LLMs based on data sensitivity, (ii) a hybrid retrieval pipeline using OpenSearch and the finance-domain