❌

Normal view

AI Deception: Risks, Dynamics, and Controls

arXiv:2511.22619v2 Announce Type: replace Abstract: As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an empirically demonstrated risk across language models, AI agents, and emerging frontier systems. This project provides a comprehensive and up-to-date overview of the AI deception field, covering its core concepts, methodologies, genesis, and potential mitigations. First, we identify a formal definition of AI deception, grounded in signaling theory from studies of animal deception. We then review existing empirical studies and associated risks, highlighting deception as a sociotechnical safety challenge. We organize the landscape of AI deception research as a deception cycle, consisting of two key components: deception emergence and deception treatment. Deception emergence reveals the mechanisms underlying AI deception: systems with sufficient capability and incentive potential inevitably engage in deceptive behaviors when triggered by external conditions. Deception treatment, in turn, focuses on detecting and addressing such behaviors. On deception emergence, we analyze incentive foundations across three hierarchical levels and identify three essential capability preconditions required for deception. We further examine contextual triggers, including supervision gaps, distributional shifts, and environmental pressures. On deception treatment, we conclude detection methods covering benchmarks and evaluation protocols in static and interactive settings. Building on the three core factors of deception emergence, we outline potential mitigation strategies and propose auditing approaches that integrate technical, community, and governance efforts to address sociotechnical challenges and future AI risks. To support ongoing work in this area, we release a living resource at www.deceptionsurvey.com.

PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing

arXiv:2512.02589v1 Announce Type: new Abstract: Large language models are increasingly embedded into academic writing workflows, yet existing assistants remain external to the editor, preventing deep interaction with document state, structure, and revision history. This separation makes it impossible to support agentic, context-aware operations directly within LaTeX editors such as Overleaf. We present PaperDebugger, an in-editor, multi-agent, and plugin-based academic writing assistant that brings LLM-driven reasoning directly into the writing environment. Enabling such in-editor interaction is technically non-trivial: it requires reliable bidirectional synchronization with the editor, fine-grained version control and patching, secure state management, multi-agent scheduling, and extensible communication with external tools. PaperDebugger addresses these challenges through a Chrome-approved extension, a Kubernetes-native orchestration layer, and a Model Context Protocol (MCP) toolchain that integrates literature search, reference lookup, document scoring, and revision pipelines. Our demo showcases a fully integrated workflow, including localized edits, structured reviews, parallel agent execution, and diff-based updates, encapsulated within a minimal-intrusion user interface (UI). Early aggregated analytics demonstrate active user engagement and validate the practicality of an editor-native, agentic writing assistant. More details about this demo and video could be found at https://github.com/PaperDebugger/PaperDebugger.

AutoSurvey2: Empowering Researchers with Next Level Automated Literature Surveys

arXiv:2510.26012v3 Announce Type: replace Abstract: The rapid growth of research literature, particularly in large language models (LLMs), has made producing comprehensive and current survey papers increasingly difficult. This paper introduces autosurvey2, a multi-stage pipeline that automates survey generation through retrieval-augmented synthesis and structured evaluation. The system integrates parallel section generation, iterative refinement, and real-time retrieval of recent publications to ensure both topical completeness and factual accuracy. Quality is assessed using a multi-LLM evaluation framework that measures coverage, structure, and relevance in alignment with expert review standards. Experimental results demonstrate that autosurvey2 consistently outperforms existing retrieval-based and automated baselines, achieving higher scores in structural coherence and topical relevance while maintaining strong citation fidelity. By combining retrieval, reasoning, and automated evaluation into a unified framework, autosurvey2 provides a scalable and reproducible solution for generating long-form academic surveys and contributes a solid foundation for future research on automated scholarly writing. All code and resources are available at https://github.com/annihi1ation/auto_research.

Whole-genome landscapes of 1,364 breast cancers

Nature, Published online: 03 December 2025; doi:10.1038/s41586-025-09812-3

Whole-genome and transcriptome analysis of 1,364 cases of breast cancer from South Korea broadens our understanding of breast cancer biology and reveals genomic features that connect tumour biology with treatment responses and clinical outcomes.

SelfAI: Building a Self-Training AI System with LLM Agents

arXiv:2512.00403v1 Announce Type: cross Abstract: Recent work on autonomous scientific discovery has leveraged LLM-based agents to integrate problem specification, experiment planning, and execution into end-to-end systems. However, these frameworks are often confined to narrow application domains, offer limited real-time interaction with researchers, and lack principled mechanisms for determining when to halt exploration, resulting in inefficiencies, reproducibility challenges, and under-utilized human expertise. To address these gaps, we propose \textit{SelfAI}, a general multi-agent platform that combines a User Agent for translating high-level research objectives into standardized experimental configurations, a Cognitive Agent powered by LLMs with optimal stopping criteria to iteratively refine hyperparameter searches, and an Experiment Manager responsible for orchestrating parallel, fault-tolerant training workflows across heterogeneous hardware while maintaining a structured knowledge base for continuous feedback. We further introduce two novel evaluation metrics, Score and $\text{AUP}_D$, to quantify discovery efficiency and search diversity. Across regression, NLP, computer vision, scientific computing, medical imaging, and drug discovery benchmarks, SelfAI consistently achieves strong performance and reduces redundant trials compared to classical Bayesian optimization and LLM-based baselines, while enabling seamless interaction with human researchers.

Human Decision-making is Susceptible to AI-driven Manipulation

arXiv:2502.07663v3 Announce Type: replace Abstract: AI systems are increasingly intertwined with daily life, assisting users with various tasks and guiding decision-making. This integration introduces risks of AI-driven manipulation, where such systems may exploit users' cognitive biases and emotional vulnerabilities to steer them toward harmful outcomes. Through a randomized between-subjects experiment with 233 participants, we examined human susceptibility to such manipulation in financial (e.g., purchases) and emotional (e.g., conflict resolution) decision-making contexts. Participants interacted with one of three AI agents: a neutral agent (NA) optimizing for user benefit without explicit influence, a manipulative agent (MA) designed to covertly influence beliefs and behaviors, or a strategy-enhanced manipulative agent (SEMA) equipped with established psychological tactics, allowing it to select and apply them adaptively during interactions to reach its hidden objectives. By analyzing participants' preference ratings, we found significant susceptibility to AI-driven manipulation. Particularly across both decision-making domains, interacting with the manipulative agents significantly increased the odds of rating hidden incentives higher than optimal options (Financial, MA: OR=5.24, SEMA: OR=7.96; Emotional, MA: OR=5.52, SEMA: OR=5.71) compared to the NA group. Notably, we found no clear evidence that employing psychological strategies (SEMA) was overall more effective than simple manipulative objectives (MA) on our primary outcomes. Hence, AI-driven manipulation could become widespread even without requiring sophisticated tactics and expertise. While our findings are preliminary and derived from hypothetical, low-stakes scenarios, we highlight a critical vulnerability in human-AI interactions, emphasizing the need for ethical safeguards and regulatory frameworks to protect human autonomy.

Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models

arXiv:2509.24365v2 Announce Type: replace-cross Abstract: Unified Multimodal Models (UMMs) built on shared autoregressive (AR) transformers are attractive for their architectural simplicity. However, we identify a critical limitation: when trained on multimodal inputs, modality-shared transformers suffer from severe gradient conflicts between vision and text, particularly in shallow and deep layers. We trace this issue to the fundamentally different low-level statistical properties of images and text, while noting that conflicts diminish in middle layers where representations become more abstract and semantically aligned. To overcome this challenge, we propose Uni-X, a two-end-separated, middle-shared architecture. Uni-X dedicates its initial and final layers to modality-specific processing, while maintaining shared parameters in the middle layers for high-level semantic fusion. This X-shaped design not only eliminates gradient conflicts at both ends but also further alleviates residual conflicts in the shared layers. Extensive experiments validate the effectiveness of Uni-X. Under identical training conditions, Uni-X achieves superior training efficiency compared to strong baselines. When scaled to 3B parameters with larger training data, Uni-X matches or surpasses 7B AR-based UMMs, achieving a GenEval score of 82 for image generation alongside strong performance in text and vision understanding tasks. These results establish Uni-X as a parameter-efficient and scalable foundation for future unified multimodal modeling. Our code is available at https://github.com/CURRENTF/Uni-X

The AI Productivity Index (APEX)

arXiv:2509.25721v3 Announce Type: replace-cross Abstract: We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Single-Cell Multi-Omics in Type 2 Diabetes Mellitus: Revealing Cellular Heterogeneity and Mechanistic Insights

Int J Mol Sci. 2025 Nov 13;26(22):11005. doi: 10.3390/ijms262211005.

ABSTRACT

Type 2 diabetes mellitus (T2DM) is a prevalent and complex metabolic disorder characterized by insulin resistance, progressive β-cell dysfunction, and severe systemic complications. Advances in single-cell multi-omics-transcriptomics, chromatin accessibility profiling, and integrative analyses-have offered unprecedented insights into the cellular heterogeneity and regulatory networks of pancreatic islets. We highlight recent discoveries in islet cell heterogeneity and β-cell pathophysiology, with a particular focus on dysfunction and dedifferentiation. We further underscore the computational frameworks that enable these discoveries, spanning data preprocessing, multi-omics integration, and machine learning-driven analyses, which collectively enable the dissection of disease-relevant cell subpopulations and the reconstruction of developmental and regulatory trajectories. We also examine how impaired signaling within islets and chronic adipose inflammation contribute to T2DM pathogenesis. Finally, we discuss key challenges in clinical translation-including limited population diversity in single-cell atlases and the interpretability of computational models-and propose future directions toward precision diagnostics and therapeutic innovation in T2DM.

PMID:41303487 | PMC:PMC12652634 | DOI:10.3390/ijms262211005

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

How Do Companies Manage the Environmental Sustainability of AI? An Interview Study About Green AI Efforts and Regulations

arXiv:2505.07317v2 Announce Type: replace-cross Abstract: With the ever-growing adoption of artificial intelligence (AI), AI-based software and its negative impact on the environment are no longer negligible, and studying and mitigating this impact has become a critical area of research. However, it is currently unclear which role environmental sustainability plays during AI adoption in industry and how AI regulations influence Green AI practices and decision-making in industry. We therefore aim to investigate the Green AI perception and management of industry practitioners. To this end, we conducted a total of 11 interviews with participants from 10 different organizations that adopted AI-based software. The interviews explored three main themes: AI adoption, current efforts in mitigating the negative environmental impact of AI, and the influence of the EU AI Act and the Corporate Sustainability Reporting Directive (CSRD). Our findings indicate that 9 of 11 participants prioritized business efficiency during AI adoption, with minimal consideration of environmental sustainability. Monitoring and mitigation of AI's environmental impact were very limited. Only one participant monitored negative environmental effects. Regarding applied mitigation practices, six participants reported no actions, with the others sporadically mentioning techniques like prompt engineering, relying on smaller models, or not overusing AI. Awareness and compliance with the EU AI Act are low, with only one participant reporting on its influence, while the CSRD drove sustainability reporting efforts primarily in larger companies. All in all, our findings reflect a lack of urgency and priority for sustainable AI among these companies. We suggest that current regulations are not very effective, which has implications for policymakers. Additionally, there is a need to raise industry awareness, but also to provide user-friendly techniques and tools for Green AI practices.

scGALA advances graph link prediction-based cell alignment for comprehensive data integration and harmonization

Nat Commun. 2025 Nov 26. doi: 10.1038/s41467-025-66644-5. Online ahead of print.

ABSTRACT

Single-cell technologies have transformed our understanding of cellular heterogeneity through multimodal data acquisition. However, robust cell alignment remains a major challenge for data integration and harmonization, including batch correction, label transfer, and multi-omics integration. Many existing methods constrain alignment based on rigid feature-wise distance metrics, limiting their ability to capture accurate cell correspondence across diverse cell populations and conditions. We introduce scGALA, a graph-based learning framework that redefines cell alignment by combining graph attention networks with a score-driven, task-independent optimization strategy. scGALA constructs enriched graphs of cell-cell relationships by integrating gene expression profiles with auxiliary information, such as spatial coordinates, and iteratively refines alignment via self-supervised graph link prediction, where a deep neural network is trained to identify and reinforce high-confidence correspondences across datasets. In extensive benchmarks, scGALA identifies over 25 percent more high-confidence alignments without compromising accuracy. By improving the core step of cell alignment, scGALA serves as a versatile enhancer for a wide range of single-cell data integration tasks.

PMID:41298467 | DOI:10.1038/s41467-025-66644-5

Clinician-Directed Large Language Model Software Generation for Therapeutic Interventions in Physical Rehabilitation

arXiv:2511.18274v1 Announce Type: cross Abstract: Digital health interventions are increasingly used in physical and occupational therapy to deliver home exercise programs via sensor equipped devices such as smartphones, enabling remote monitoring of adherence and performance. However, digital interventions are typically programmed as software before clinical encounters as libraries of parametrized exercise modules targeting broad patient populations. At the point of care, clinicians can only select modules and adjust a narrow set of parameters like repetitions, so patient specific needs that emerge during encounters, such as distinct movement limitations, and home environments, are rarely reflected in the software. We evaluated a digital intervention paradigm that uses large language models (LLMs) to translate clinicians' exercise prescriptions into intervention software. In a prospective single arm feasibility study with 20 licensed physical and occupational therapists and a standardized patient, clinicians created 40 individualized upper extremity programs (398 instructions) that were automatically translated into executable software. Our results show a 45% increase in the proportion of personalized prescriptions that can be implemented as software compared with a template based benchmark, with unanimous consensus among therapists on ease of use. The LLM generated software correctly delivered 99.78% (397/398) of instructions as prescribed and monitored performance with 88.4% (352/398) accuracy, with 90% (18/20) of therapists judged it safe to interact with patients, and 75% (15/20) expressed willingness to adopt it. To our knowledge, this is the first prospective evaluation of clinician directed intervention software generation with LLMs in healthcare, demonstrating feasibility and motivating larger trials to assess clinical effectiveness and safety in real patient populations.

AI-Assisted Cardiovascular Risk Assessment by General Practitioners in Resource-Constrained Indonesian Settings Using a Conceptual Prototype: Randomized Controlled Study

Background: Preventive strategies integrated with digital health and artificial intelligence (AI), have significant potential to mitigate the global burden of atherosclerotic cardiovascular disease (ASCVD). AI-enabled clinical decision support (CDS) systems increasingly provide patient-specific insights beyond traditional risk factors. Despite these advances, their capacity to enhance clinical decision-making in resource-constrained settings remains largely unexplored. Objective: We conducted a randomised controlled study to assess the effect of AI-based CDS on 10-year ASCVD risk assessment and management in primary prevention. Methods: In a three-way within-subject randomised design, doctors completed nine clinical vignettes representative of primary care presentations in a resource-constrained outpatient setting. For each vignette, participants assessed 10-year ASCVD risk and made management decisions using either a conceptual prototype of AI-based CDS, automated CDS, or no decision support. The conceptual prototype represented contemporary risk calculators based on traditional machine learning models (e.g., random forest, neural networks, logistic regression) that incorporate additional predictors alongside traditional risk factors. Primary outcomes were correct risk assessment and patient management (prescription of aspirin, statins, and anti-hypertensives; referral for advanced examinations). Decision-making time and perceptions about AI utility were also measured. Results: 102 doctors from all seven geographical regions of Indonesia participated. Most participants were 26–35 years (83%), 56% male, with a median of six years of clinical experience (IQR=4.75). AI-based CDS improved risk assessment by 27% (2(2, n=102) = 48.875, P<.001 compared to unassisted or one additional correct risk classification for every patients where doctors use ai needed treat nnt="3.7;" ci the prescription of statins also improved by n="102)=" p in pairwise comparisons assisted with ai-based cds correctly assessed significantly more cases adjusted and prescribed appropriate statin often medium effect size r=".35)" control. ai-assisted required less time marginal means sec vs f however improvements aspirin anti-hypertensives did not reach statistical significance. no improvement was observed referral decisions. participants generally viewed positively agreeing strongly that they would follow its recommendations indicating it if given access. believed could enhance efficiency assessment particularly high-volume primary care settings while noting need verify against clinical guidelines each patient. conclusions: coupled reduced decision-making highlight potential utility ascvd resource-constrained efficient healthcare resources is crucial. further research ascertain whether this online study translate real-world low-resource settings.>

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

25 November 2025 at 19:00

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | PMC:PMC12647489 | DOI:10.1007/s12672-025-04010-z

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | DOI:10.1007/s12672-025-04010-z

When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models

arXiv:2511.16203v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) have recently demonstrated remarkable progress in embodied environments, enabling robots to perceive, reason, and act through unified multimodal understanding. Despite their impressive capabilities, the adversarial robustness of these systems remains largely unexplored, especially under realistic multimodal and black-box conditions. Existing studies mainly focus on single-modality perturbations and overlook the cross-modal misalignment that fundamentally affects embodied reasoning and decision-making. In this paper, we introduce VLA-Fool, a comprehensive study of multimodal adversarial robustness in embodied VLA models under both white-box and black-box settings. VLA-Fool unifies three levels of multimodal adversarial attacks: (1) textual perturbations through gradient-based and prompt-based manipulations, (2) visual perturbations via patch and noise distortions, and (3) cross-modal misalignment attacks that intentionally disrupt the semantic correspondence between perception and instruction. We further incorporate a VLA-aware semantic space into linguistic prompts, developing the first automatically crafted and semantically guided prompting framework. Experiments on the LIBERO benchmark using a fine-tuned OpenVLA model reveal that even minor multimodal perturbations can cause significant behavioral deviations, demonstrating the fragility of embodied multimodal alignment.

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.
❌