Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
arXiv:2602.12670v3 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurat
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards AI Search Paradigm
arXiv:2506.17188v2 Announce Type: replace-cross Abstract: In this paper, we introduce the AI Search Paradigm, a comprehensive blueprint for next-generation search systems capable of emulating human information processing and decision-making. The paradigm employs a modular architecture of four LLM-powered agents (Master, Planner, Executor and Writer) that dynamically adapt to the full spectrum of information needs, from simple factual queries to complex multi-stage reasoning tasks. These agents
Towards AI Search Paradigm
-
cs.AI, q-bio.NC updates on arXiv.org
-
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
arXiv:2602.07801v3 Announce Type: replace-cross Abstract: In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize-clip-answer pipeline in which the model actively identifies relevant video segments, performs dense sampling within those clips, and then produces answers. However, existing
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
-
Pulmonary nodule
-
Profiling of the mycobiome and metabolome: a comparative study of benign pulmonary nodules and lung adenocarcinoma
Front Cell Infect Microbiol. 2026 Feb 23;16:1732958. doi: 10.3389/fcimb.2026.1732958. eCollection 2026.ABSTRACTINTRODUCTION: Lung adenocarcinoma (LUAD), the most common subtype of non-small cell lung cancer, is a form of malignant pulmonary nodule that requires clinical differentiation from benign pulmonary nodules (BPN). The mechanisms underlying the development of LUAD are complex, and effective non-invasive methods for differentiating BPN from LUAD are lacking. This study aimed not only to di
Profiling of the mycobiome and metabolome: a comparative study of benign pulmonary nodules and lung adenocarcinoma
Front Cell Infect Microbiol. 2026 Feb 23;16:1732958. doi: 10.3389/fcimb.2026.1732958. eCollection 2026.
ABSTRACT
INTRODUCTION: Lung adenocarcinoma (LUAD), the most common subtype of non-small cell lung cancer, is a form of malignant pulmonary nodule that requires clinical differentiation from benign pulmonary nodules (BPN). The mechanisms underlying the development of LUAD are complex, and effective non-invasive methods for differentiating BPN from LUAD are lacking. This study aimed not only to distinguish BPN from LUAD using gut fungi and serum metabolites, but also to establish an integrated network of gut fungi-metabolite-cytokine interactions.
METHODS: Fecal and serum samples from individuals with BPN and patients with LUAD were subjected to internal transcribed spacer sequencing, ultra-performance liquid chromatography-tandem mass spectrometry, and multiplex Luminex assays to quantify gut fungi, metabolites, and cytokines, respectively.
RESULTS: A significant difference in gut fungal communities was observed between the BPN and LUAD groups. Multiple genera and species were more abundant in LUAD than in BPN. Docosapentaenoic acid n-6 (DPAn-6), indole-3-propionic acid (IPA), and interferon-γ-induced protein 10 (IP-10) were significantly elevated in the LUAD group. The integrated model established using a combination of gut fungi and metabolites demonstrated excellent performance in distinguishing BPN from LUAD. A network of interactions was established among differentially abundant gut fungi, serum metabolites, and cytokines.
CONCLUSION: Our study identifies a novel panel of fungal and metabolite biomarkers for differentiating between BPN and LUAD, and constructs a multi-omics network that provides new insights into investigating the mechanistic role of gut mycobiota dysbiosis in LUAD.
PMID:41809995 | PMC:PMC12968269 | DOI:10.3389/fcimb.2026.1732958
-
cs.AI, q-bio.NC updates on arXiv.org
-
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
arXiv:2603.07427v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, existing safety evaluations face a fundamental trade-off: manual benchmarks are costly, while LLM-based simulators are scalable but suffer from logic hallucination. We present AutoControl Arena, an automated framework for frontier AI risk evaluation built on the principle of logic-narrative decoupling. By grounding deterministic state in executable code while delegating generative dyna
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning
arXiv:2603.07733v1 Announce Type: new Abstract: This work investigated the capabilities of different models, including the Llama-3 series of models and CHATGPT, with different forms of expression in solving discrete optimization problems by testing natural language datasets. In contrast to formal datasets with a limited scope of parameters, our dataset included a variety of problem types in discrete optimization problems and featured a wide range of parameter magnitudes, including instances wit
Large Language Model for Discrete Optimization Problems: Evaluation and Step-by-step Reasoning
-
Journal of Medical Internet Research
-
eHealth Literacy and Type 2 Diabetes Prevention Among At-Risk Populations: Mechanistic Systematic Review Using Theory-Driven Thematic Analysis
Background: Type 2 diabetes (T2D) is emerging as a growing global public health crisis. Early and effective interventions can reduce T2D incidence among at-risk populations. Compared with traditional approaches, digital health technologies offer promising opportunities for prevention, with eHealth literacy (eHL) emerging as a critical determinant of digital prevention outcomes. Objective: This systematic review aims to synthesize and explain the pathways and mechanisms through which eHL supports
eHealth Literacy and Type 2 Diabetes Prevention Among At-Risk Populations: Mechanistic Systematic Review Using Theory-Driven Thematic Analysis
-
cs.AI, q-bio.NC updates on arXiv.org
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
arXiv:2602.12670v2 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurat
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
-
cs.AI, q-bio.NC updates on arXiv.org
-
MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultation
arXiv:2603.03677v1 Announce Type: cross Abstract: Large language models (LLMs) have advanced medical dialogue systems, yet psychiatric consultation poses substantially higher demands due to subjective ambiguity and comorbidity complexity: an agent must continuously extract psychopathological cues from incomplete and inconsistent patient reports in multi-turn interactions and perform rigorous differential diagnostic reasoning. However, existing methods face two fundamental challenges. First, wit
MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
arXiv:2603.03752v1 Announce Type: cross Abstract: Large language models (LLMs) demonstrate superior reasoning capabilities compared to small language models (SLMs), but incur substantially higher costs. We propose COllaborative REAsoner (COREA), a system that cascades an SLM with an LLM to achieve a balance between accuracy and cost in complex reasoning tasks. COREA first attempts to answer questions using the SLM, which outputs both an answer and a verbalized confidence score. Questions with c
Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
DisenReason: Behavior Disentanglement and Latent Reasoning for Shared-Account Sequential Recommendation
arXiv:2603.03782v1 Announce Type: cross Abstract: Shared-account usage is common on streaming and e-commerce platforms, where multiple users share one account. Existing shared-account sequential recommendation (SSR) methods often assume a fixed number of latent users per account, limiting their ability to adapt to diverse sharing patterns and reducing recommendation accuracy. Recent latent reasoning technique applied in sequential recommendation (SR) generate intermediate embeddings from the us
DisenReason: Behavior Disentanglement and Latent Reasoning for Shared-Account Sequential Recommendation
-
cs.AI, q-bio.NC updates on arXiv.org
-
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
arXiv:2602.07801v2 Announce Type: replace-cross Abstract: In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize-clip-answer pipeline in which the model actively identifies relevant video segments, performs dense sampling within those clips, and then produces answers. However, existing
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
-
cs.AI, q-bio.NC updates on arXiv.org
-
CRCC: Contrast-Based Robust Cross-Subject and Cross-Site Representation Learning for EEG
arXiv:2602.19138v1 Announce Type: new Abstract: EEG-based neural decoding models often fail to generalize across acquisition sites due to structured, site-dependent biases implicitly exploited during training. We reformulate cross-site clinical EEG learning as a bias-factorized generalization problem, in which domain shifts arise from multiple interacting sources. We identify three fundamental bias factors and propose a general training framework that mitigates their influence through data stan
CRCC: Contrast-Based Robust Cross-Subject and Cross-Site Representation Learning for EEG
-
cs.AI, q-bio.NC updates on arXiv.org
-
FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations
arXiv:2602.18437v1 Announce Type: cross Abstract: Generating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-optimize citation fidelity while overlooking relevance to the user query, which degrades answer quality and robustness in real-world settings with noisy or irrelevant retrieved content. Moreover, the prevailing single-pass paradigm struggles to deliver optimal answers in
FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations
-
cs.AI, q-bio.NC updates on arXiv.org
-
Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models
arXiv:2602.19101v1 Announce Type: cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we repo
Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations
arXiv:2602.19320v1 Announce Type: cross Abstract: Agentic memory systems enable large language model (LLM) agents to maintain state across long interactions, supporting long-horizon reasoning and personalization beyond fixed context windows. Despite rapid architectural development, the empirical foundations of these systems remain fragile: existing benchmarks are often underscaled, evaluation metrics are misaligned with semantic utility, performance varies significantly across backbone models,
Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Very Big Video Reasoning Suite
arXiv:2602.20159v1 Announce Type: cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindere
A Very Big Video Reasoning Suite
-
cs.AI, q-bio.NC updates on arXiv.org
-
MAS-on-the-Fly: Dynamic Adaptation of LLM-based Multi-Agent Systems at Test Time
arXiv:2602.13671v1 Announce Type: cross Abstract: Large Language Model (LLM)-based multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. However, existing works often rely on manual designs or "one-size-fits-all" automation, lacking dynamic adaptability after deployment. Inspired by how biological systems adapt, we introduce MASFly, a novel multi-agent framework enabling dynamic adaptation at test time. To adapt system generation, MASFly employs a retrieval-a
MAS-on-the-Fly: Dynamic Adaptation of LLM-based Multi-Agent Systems at Test Time
-
cs.AI, q-bio.NC updates on arXiv.org
-
DenseMLLM: Standard Multimodal LLMs are Intrinsic Dense Predictors
arXiv:2602.14134v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending these models to fine-grained dense prediction tasks, such as semantic segmentation and depth estimation, typically necessitates the incorporation of complex, task-specific decoders and other customizations. This architectural fragmentation increases model complexity and deviates from the generalist design of
DenseMLLM: Standard Multimodal LLMs are Intrinsic Dense Predictors
-
cs.AI, q-bio.NC updates on arXiv.org
-
AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
arXiv:2510.18428v3 Announce Type: replace Abstract: Optimization modeling underlies critical decision-making across industries, yet remains difficult to automate: natural-language problem descriptions must be translated into precise mathematical formulations and executable solver code. Existing LLM-based approaches typically rely on brittle prompting or costly retraining, both of which offer limited generalization. Recent work suggests that large models can improve via experience reuse, but how