❌

Normal view

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.

A data-intelligence-intensive bioinformatics copilot system for large-scale omics research and scientific insights

Brief Bioinform. 2025 Jul 2;26(4):bbaf312. doi: 10.1093/bib/bbaf312.

ABSTRACT

Advancements in high-throughput sequencing technologies and artificial intelligence (AI) offer unprecedented opportunities for groundbreaking discoveries in bioinformatics research. However, the challenges of exponential growth of omics data and the rapid development of AI technologies require automated big biological data analysis capability and interdisciplinary knowledge-driven scientific insight. Here, we propose a data-intelligence-intensive bioinformatics copilot (Bio-Copilot) system that synergizes AI capabilities with human researchers to facilitate hypothesis-free exploratory research and inspire novel scientific insights in large-scale omics studies. Bio-Copilot forms high-quality intensive intelligence through close collaboration between multiple agents, driven by large language models (LLMs), and human researchers. To augment the capabilities of Bio-Copilot, this study devises an agent group management strategy, an effective human-agent interaction mechanism, a shared interdisciplinary knowledge database, and continuous learning strategies for the agents. We comprehensively compare Bio-Copilot against GPT-4o and several leading AI agents across diverse bioinformatics tasks, using a broad range of evaluation metrics. Bio-Copilot achieves overall state-of-the-art performance across all tasks, while showcasing exceptional task completeness. Furthermore, on application to constructing a large-scale human lung cell atlas, Bio-Copilot not only reproduces the intricate data integration process detailed in a seminal study but also introduces a recursive, multilevel annotation strategy to capture the continuous nature of cellular states and uncovers the characteristics of rare cell types, highlighting its potential to unravel hidden complexities in biological systems. Beyond the technical achievements, this study also underscores the profound implications of integrating AI capabilities with expert knowledge in accelerating impactful biological discoveries and exploring uncharted territories.

PMID:40639418 | PMC:PMC12245162 | DOI:10.1093/bib/bbaf312

Interplay between gut microbial communities and metabolites modulates pan-cancer immunotherapy responses

Cell Metab. 2025 Jan 28:S1550-4131(24)00495-9. doi: 10.1016/j.cmet.2024.12.013. Online ahead of print.

ABSTRACT

Immune checkpoint blockade (ICB) therapy has revolutionized cancer treatment but remains effective in only a subset of patients. Emerging evidence suggests that the gut microbiome and its metabolites critically influence ICB efficacy. In this study, we performed a multi-omics analysis of fecal microbiomes and metabolomes from 165 patients undergoing anti-programmed cell death protein 1 (PD-1)/programmed death ligand 1 (PD-L1) therapy, identifying microbial and metabolic entities associated with treatment response. Integration of data from four public metagenomic datasets (n = 568) uncovered cross-cohort microbial and metabolic signatures, validated in an independent cohort (n = 138). An integrated predictive model incorporating these features demonstrated robust performance. Notably, we characterized five response-associated enterotypes, each linked to specific bacterial taxa and metabolites. Among these, the metabolite phenylacetylglutamine (PAGln) was negatively correlated with response and shown to attenuate anti-PD-1 efficacy in vivo. This study sheds light on the interplay among the gut microbiome, the gut metabolome, and immunotherapy response, identifying potential biomarkers to improve treatment outcomes.

PMID:39909032 | DOI:10.1016/j.cmet.2024.12.013

❌