Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Denario project: Deep knowledge AI agents for scientific discovery
arXiv:2510.26887v1 Announce Type: new Abstract: We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific
-
cs.AI, q-bio.NC updates on arXiv.org
-
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
arXiv:2510.27176v1 Announce Type: new Abstract: Can an AI autonomously design mechanisms for computer systems on par with the creativity and reasoning of human experts? We present Glia, an AI architecture for networked systems design that uses large language models (LLMs) in a human-inspired, multi-agent workflow. Each agent specializes in reasoning, experimentation, and analysis, collaborating through an evaluation framework that grounds abstract reasoning in empirical feedback. Unlike prior M
Glia: A Human-Inspired AI for Automated Systems Design and Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
arXiv:2510.27363v1 Announce Type: new Abstract: Recently, large language models (LLMs) have demonstrated remarkable problem-solving capabilities by autonomously integrating with external tools for collaborative reasoning. However, due to the inherently complex and diverse nature of multimodal information, enabling multimodal large language models (MLLMs) to flexibly and efficiently utilize external tools during reasoning remains an underexplored challenge. In this work, we introduce ToolScope,
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
-
cs.AI, q-bio.NC updates on arXiv.org
-
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
arXiv:2510.27617v1 Announce Type: new Abstract: Automation of Register Transfer Level (RTL) design can help developers meet increasing computational demands. Large Language Models (LLMs) show promise for Hardware Description Language (HDL) generation, but face challenges due to limited parametric knowledge and domain-specific constraints. While prompt engineering and fine-tuning have limitations in knowledge coverage and training costs, multi-agent architectures offer a training-free paradigm t
VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
arXiv:2510.26969v1 Announce Type: cross Abstract: We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients' visits to primary care units. A total of eight pattern
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
-
cs.AI, q-bio.NC updates on arXiv.org
-
Detecting Data Contamination in LLMs via In-Context Learning
arXiv:2510.27055v1 Announce Type: cross Abstract: We present Contamination Detection via Context (CoDeC), a practical and accurate method to detect and quantify training data contamination in large language models. CoDeC distinguishes between data memorized during training and data outside the training distribution by measuring how in-context learning affects model performance. We find that in-context examples typically boost confidence for unseen datasets but may reduce it when the dataset was
Detecting Data Contamination in LLMs via In-Context Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
arXiv:2510.27163v1 Announce Type: cross Abstract: Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the differenc
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
arXiv:2510.27196v1 Announce Type: cross Abstract: The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on mLLMs' detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. In this paper, we propose MemeArena, an agent-based arena-style eval
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
arXiv:2510.27287v1 Announce Type: cross Abstract: Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences, and efficient information retrieval, driving operational efficiency and strategic growth. However, developing and evaluating such systems is challenging due to the inherent complexity of enterprise environments, where data is fr
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
-
cs.AI, q-bio.NC updates on arXiv.org
-
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
arXiv:2510.27606v1 Announce Type: cross Abstract: Spatial understanding remains a weakness of Large Vision-Language Models (LVLMs). Existing supervised fine-tuning (SFT) and recent reinforcement learning with verifiable rewards (RLVR) pipelines depend on costly supervision, specialized tools, or constrained environments that limit scale. We introduce Spatial-SSRL, a self-supervised RL paradigm that derives verifiable signals directly from ordinary RGB or RGB-D images. Spatial-SSRL automatically
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
arXiv:2510.27629v1 Announce Type: cross Abstract: Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who mig
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Survey of AI Scientists
arXiv:2510.23045v3 Announce Type: replace Abstract: Artificial intelligence is undergoing a profound transition from a computational instrument to an autonomous originator of scientific knowledge. This emerging paradigm, the AI scientist, is architected to emulate the complete scientific workflow-from initial hypothesis generation to the final synthesis of publishable findings-thereby promising to fundamentally reshape the pace and scale of discovery. However, the rapid and unstructured prolife
A Survey of AI Scientists
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
arXiv:2410.22377v3 Announce Type: replace-cross Abstract: In recent years, spatio-temporal graph neural networks (GNNs) have attracted considerable interest in the field of time series analysis, due to their ability to capture, at once, dependencies among variables and across time points. The objective of this systematic literature review is hence to provide a comprehensive overview of the various modeling approaches and application domains of GNNs for time series classification and forecasting
A Systematic Literature Review of Spatio-Temporal Graph Neural Network Models for Time Series Forecasting and Classification
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Multi-Stage Framework with Taxonomy-Guided Reasoning for Occupation Classification Using Large Language Models
arXiv:2503.12989v3 Announce Type: replace-cross Abstract: Automatically annotating job data with standardized occupations from taxonomies, known as occupation classification, is crucial for labor market analysis. However, this task is often hindered by data scarcity and the challenges of manual annotations. While large language models (LLMs) hold promise due to their extensive world knowledge and in-context learning capabilities, their effectiveness depends on their knowledge of occupational ta
A Multi-Stage Framework with Taxonomy-Guided Reasoning for Occupation Classification Using Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
arXiv:2509.24204v2 Announce Type: replace-cross Abstract: Vision foundation models like the Segment Anything Model (SAM), pretrained on large-scale natural image datasets, often struggle in medical image segmentation due to a lack of domain-specific adaptation. In clinical practice, fine-tuning such models efficiently for medical downstream tasks with minimal resource demands, while maintaining strong performance, is challenging. To address these issues, we propose BALR-SAM, a boundary-aware lo
BALR-SAM: Boundary-Aware Low-Rank Adaptation of SAM for Resource-Efficient Medical Image Segmentation
-
MRD
-
International expert consensus on the clinical integration of circulating tumor cells in solid tumors
Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.ABSTRACTBACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an in
International expert consensus on the clinical integration of circulating tumor cells in solid tumors
Eur J Cancer. 2025 Dec 9;231:116050. doi: 10.1016/j.ejca.2025.116050. Epub 2025 Oct 20.
ABSTRACT
BACKGROUND: Circulating tumor cells (CTCs) are a versatile biomarker in solid tumors. Extensive research supports their clinical relevance and led to regulatory approval in breast, prostate, and colorectal cancers. However, clinical adoption remains limited mainly due to the lack of consensus and standardized technologies. Additionally, CTC research lacks unified direction. To address these gaps, an international expert panel was established to assess the current and future clinical utility of CTCs.
METHODS: A panel of 11 CTC experts identified key areas of controversy, informing a structured survey distributed to 55 international multidisciplinary experts. Consensus was predefined as ≥ 70 % agreement. Areas without consensus were discussed in a virtual meeting, leading to final statements on the clinical integration of CTCs.
RESULTS: Thirty-seven experts completed the survey. Consensus was reached on the clinical utility of CTCs for prognosis and treatment monitoring in metastatic breast (BC) and prostate (PC) cancers, including AR-V7 testing in metastatic castration-resistant PC for therapy selection. In other tumors, CTCs remain investigational. Experts agreed that while clinical utility is not yet established in early-stage disease, CTCs show promise in early BC, especially combined with cell-free DNA (cfDNA) for minimal residual disease detection. CellSearch® is currently the only platform with high-level evidence for clinical use, though emerging technologies are promising. Key challenges include improving detection sensitivity/specificity, standardizing workflows, generating robust data, and clinician education. Experts emphasized shifting from enumeration to phenotypic and molecular characterization, particularly for treatment guidance, and highlighted the complementary role of CTCs and cfDNA, advocating for integrated liquid biopsy approaches.
CONCLUSIONS: This consensus offers practical guidance for clinical integration of CTCs and outlines strategic research priorities to unlock their full potential in precision oncology.
PMID:41172567 | DOI:10.1016/j.ejca.2025.116050
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance
Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.ABSTRACTBACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.METHODS: PubMed, Scopus, and Web of Science w
Animal models in tuberculosis metabolomics: a systematic review of current evidence and the road to translational relevance
Front Mol Biosci. 2025 Oct 15;12:1688882. doi: 10.3389/fmolb.2025.1688882. eCollection 2025.
ABSTRACT
BACKGROUND: Animal models are important for tuberculosis (TB) research, offering controlled settings to study disease mechanisms. However, their ability to replicate TB-induced metabolic responses in humans is uncertain. This systematic review evaluated the current use of animal models in metabolomics studies aimed at characterising active pulmonary TB.
METHODS: PubMed, Scopus, and Web of Science were systematically searched for metabolomics studies of pulmonary TB in humans and animal models, following PRISMA guidelines. Eligible studies were screened, and quality was assessed using QUDOMICS and STAIR tools. Data were synthesised by species, sample matrix, experimental design, and reported differential metabolites. Differential metabolite names were compared between species and subjected to pathway analysis in MetaboAnalyst 6.0.
RESULTS: Of the 80 eligible studies, nine involved animal models, predominantly mice. These models captured only 4.7% of human TB-associated differential metabolites, with the highest overlap (3.8%) in mouse lung tissue. Despite low concordance at metabolite level, conserved disruptions were observed in amino acid, glutathione, and one-carbon metabolism pathways. Interspecies variation was evident, influenced by host species, sample matrix, infection protocol, and analytical method.
CONCLUSION: Animal models partially replicated key metabolic features of human TB, particularly at the pathway level. However, variability across studies hampers current translational interpretation. Broader model use, standardised protocols, and integrated multi-platform omics approaches are needed to improve the relevance and comparability of animal models in TB metabolomics research.
PMID:41169614 | PMC:PMC12568366 | DOI:10.3389/fmolb.2025.1688882
-
Journal of Medical Internet Research
-
Digital Health Technology Compliance With Clinical Safety Standards In the National Health Service in England: National Cross-Sectional Study
Background: To be authorized for use in the National Health Service (NHS) in England, digital health technologies (DHTs) must meet 2 mandatory clinical risk management standards, Data Coordination Board (DCB) 0129 and 0160, demonstrating that risks from design and use have been assessed and mitigated. NHS organizations must not procure a DHT without DCB0129 assurance and must not deploy one without DCB0160 assurance. Despite legal requirement, no public data exist on how many DHTs are in use in
Digital Health Technology Compliance With Clinical Safety Standards In the National Health Service in England: National Cross-Sectional Study
-
cs.AI, q-bio.NC updates on arXiv.org
-
An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0
arXiv:2510.25813v1 Announce Type: new Abstract: We present a novel framework for Industry 5.0 that simplifies the deployment of AI models on edge devices in various industrial settings. The design reduces latency and avoids external data transfer by enabling local inference and real-time processing. Our implementation is agent-based, which means that individual agents, whether human, algorithmic, or collaborative, are responsible for well-defined tasks, enabling flexibility and simplifying inte
An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0
-
cs.AI, q-bio.NC updates on arXiv.org
-
Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world
arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. Th