❌

Normal view

An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0

arXiv:2510.25813v1 Announce Type: new Abstract: We present a novel framework for Industry 5.0 that simplifies the deployment of AI models on edge devices in various industrial settings. The design reduces latency and avoids external data transfer by enabling local inference and real-time processing. Our implementation is agent-based, which means that individual agents, whether human, algorithmic, or collaborative, are responsible for well-defined tasks, enabling flexibility and simplifying integration. Moreover, our framework supports modular integration and maintains low resource requirements. Preliminary evaluations concerning the food industry in real scenarios indicate improved deployment time and system adaptability performance. The source code is publicly available at https://github.com/AI-REDGIO-5-0/ci-component.

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

arXiv:2510.25908v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated transformative potential in scientific research, yet their deployment in high-stakes contexts raises significant trustworthiness concerns. Here, we introduce SciTrust 2.0, a comprehensive framework for evaluating LLM trustworthiness in scientific applications across four dimensions: truthfulness, adversarial robustness, scientific safety, and scientific ethics. Our framework incorporates novel, open-ended truthfulness benchmarks developed through a verified reflection-tuning pipeline and expert validation, alongside a novel ethics benchmark for scientific research contexts covering eight subcategories including dual-use research and bias. We evaluated seven prominent LLMs, including four science-specialized models and three general-purpose industry models, using multiple evaluation metrics including accuracy, semantic similarity measures, and LLM-based scoring. General-purpose industry models overall outperformed science-specialized models across each trustworthiness dimension, with GPT-o4-mini demonstrating superior performance in truthfulness assessments and adversarial robustness. Science-specialized models showed significant deficiencies in logical and ethical reasoning capabilities, along with concerning vulnerabilities in safety evaluations, particularly in high-risk domains such as biosecurity and chemical weapons. By open-sourcing our framework, we provide a foundation for developing more trustworthy AI systems and advancing research on model safety and ethics in scientific contexts.

Human-AI Complementarity: A Goal for Amplified Oversight

arXiv:2510.26518v1 Announce Type: new Abstract: Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can leverage AI to improve the quality of human oversight. We focus on an important safety problem that is already challenging for humans: fact-verification of AI outputs. We find that combining AI ratings and human ratings based on AI rater confidence is better than relying on either alone. Giving humans an AI fact-verification assistant further improves their accuracy, but the type of assistance matters. Displaying AI explanation, confidence, and labels leads to over-reliance, but just showing search results and evidence fosters more appropriate trust. These results have implications for Amplified Oversight -- the challenge of combining humans and AI to supervise AI systems even as they surpass human expert performance.

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling

arXiv:2510.26603v1 Announce Type: new Abstract: The electricity sector transition requires substantial increases in residential demand response capacity, yet Home Energy Management Systems (HEMS) adoption remains limited by user interaction barriers requiring translation of everyday preferences into technical parameters. While large language models have been applied to energy systems as code generators and parameter extractors, no existing implementation deploys LLMs as autonomous coordinators managing the complete workflow from natural language input to multi-appliance scheduling. This paper presents an agentic AI HEMS where LLMs autonomously coordinate multi-appliance scheduling from natural language requests to device control, achieving optimal scheduling without example demonstrations. A hierarchical architecture combining one orchestrator with three specialist agents uses the ReAct pattern for iterative reasoning, enabling dynamic coordination without hardcoded workflows while integrating Google Calendar for context-aware deadline extraction. Evaluation across three open-source models using real Austrian day-ahead electricity prices reveals substantial capability differences. Llama-3.3-70B successfully coordinates all appliances across all scenarios to match cost-optimal benchmarks computed via mixed-integer linear programming, while other models achieve perfect single-appliance performance but struggle to coordinate all appliances simultaneously. Progressive prompt engineering experiments demonstrate that analytical query handling without explicit guidance remains unreliable despite models' general reasoning capabilities. We open-source the complete system including orchestration logic, agent prompts, tools, and web interfaces to enable reproducibility, extension, and future research.

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.

Multi-Agent Reinforcement Learning for Market Making: Competition without Collusion

arXiv:2510.25929v1 Announce Type: cross Abstract: Algorithmic collusion has emerged as a central question in AI: Will the interaction between different AI agents deployed in markets lead to collusion? More generally, understanding how emergent behavior, be it a cartel or market dominance from more advanced bots, affects the market overall is an important research question. We propose a hierarchical multi-agent reinforcement learning framework to study algorithmic collusion in market making. The framework includes a self-interested market maker (Agent~A), which is trained in an uncertain environment shaped by an adversary, and three bottom-layer competitors: the self-interested Agent~B1 (whose objective is to maximize its own PnL), the competitive Agent~B2 (whose objective is to minimize the PnL of its opponent), and the hybrid Agent~B$^\star$, which can modulate between the behavior of the other two. To analyze how these agents shape the behavior of each other and affect market outcomes, we propose interaction-level metrics that quantify behavioral asymmetry and system-level dynamics, while providing signals potentially indicative of emergent interaction patterns. Experimental results show that Agent~B2 secures dominant performance in a zero-sum setting against B1, aggressively capturing order flow while tightening average spreads, thus improving market execution efficiency. In contrast, Agent~B$^\star$ exhibits a self-interested inclination when co-existing with other profit-seeking agents, securing dominant market share through adaptive quoting, yet exerting a milder adverse impact on the rewards of Agents~A and B1 compared to B2. These findings suggest that adaptive incentive control supports more sustainable strategic co-existence in heterogeneous agent environments and offers a structured lens for evaluating behavioral design in algorithmic trading systems.

The Quest for Reliable Metrics of Responsible AI

arXiv:2510.26007v1 Announce Type: cross Abstract: The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less work on assessing the robustness and reliability of the metrics themselves. We reflect on prior work that examines the robustness of fairness metrics for recommender systems as a type of AI application and summarise their key takeaways into a set of non-exhaustive guidelines for developing reliable metrics of responsible AI. Our guidelines apply to a broad spectrum of AI applications, including AIS.

Integrating Genomics into Multimodal EHR Foundation Models

arXiv:2510.23639v2 Announce Type: replace-cross Abstract: This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health profiles. Leveraging the extensive and diverse data from the All of Us (AoU) Research Program, this multimodal framework aims to learn complex relationships between clinical data and genetic predispositions. The methodology extends advancements in generative AI to the EHR foundation model space, enhancing predictive capabilities and interpretability. Evaluation on AoU data demonstrates the model's predictive value for the onset of various conditions, particularly Type 2 Diabetes (T2D), and illustrates the interplay between PRS and EHR data. The work also explores transfer learning for custom classification tasks, showcasing the architecture's versatility and efficiency. This approach is pivotal for unlocking new insights into disease prediction, proactive health management, risk stratification, and personalized treatment strategies, laying the groundwork for more personalized, equitable, and actionable real-world evidence generation in healthcare.

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.

Advancing Non-Small-Cell Lung Cancer Management Through Multi-Omics Integration: Insights from Genomics, Metabolomics, and Radiomics

Diagnostics (Basel). 2025 Oct 14;15(20):2586. doi: 10.3390/diagnostics15202586.

ABSTRACT

The integration of multi-omics technologies is transforming the landscape of cancer management, offering unprecedented insights into tumor biology, early diagnosis, and personalized therapy. This review provides a comprehensive overview of the current state of omics approaches, with a particular focus on the application of genomics, NMR-based metabolomics, and radiomics in non-small cell lung cancer (NSCLC). Genomics currently represents one of the most established omics technologies in oncology, as it enables the identification of genetic alterations that drive tumor initiation, progression, and therapeutic response. Interestingly, genomic analyses have revealed that many tumors harbor mutations in genes encoding metabolic enzymes, thus establishing a tight connection between genomics and tumor metabolism. In parallel, metabolomics profiling-by capturing the metabolic phenotype of tumors-has, in recent years, identified specific biomarkers associated with tumor burden, progression, and prognosis. Such findings have catalyzed growing interest in metabolomics as a complementary approach to better characterize cancer biology and discover novel diagnostic and therapeutic targets. Moreover, radiomics, through the extraction of quantitative features from standard imaging modalities, captures tumor heterogeneity and contributes predictive information on tumor biology, treatment response, and clinical outcomes. As a non-invasive and widely available technique, radiomics has the potential to support longitudinal monitoring and individualized treatment planning. Both metabolomics and radiomics, when integrated with genomic data, could support a more comprehensive understanding of NSCLC and pave the way for the development of non-invasive, predictive models and personalized therapeutic strategies. In addition, we explore the specific contributions of these technologies in enhancing clinical decision-making for lung cancer patients, with particular attention to their potential in early diagnosis, treatment selection, and real-time monitoring.

PMID:41153258 | DOI:10.3390/diagnostics15202586

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

arXiv:2510.24161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.

Integrating Genomics into Multimodal EHR Foundation Models

arXiv:2510.23639v1 Announce Type: cross Abstract: This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health profiles. Leveraging the extensive and diverse data from the All of Us (AoU) Research Program, this multimodal framework aims to learn complex relationships between clinical data and genetic predispositions. The methodology extends advancements in generative AI to the EHR foundation model space, enhancing predictive capabilities and interpretability. Evaluation on AoU data demonstrates the model's predictive value for the onset of various conditions, particularly Type 2 Diabetes (T2D), and illustrates the interplay between PRS and EHR data. The work also explores transfer learning for custom classification tasks, showcasing the architecture's versatility and efficiency. This approach is pivotal for unlocking new insights into disease prediction, proactive health management, risk stratification, and personalized treatment strategies, laying the groundwork for more personalized, equitable, and actionable real-world evidence generation in healthcare.

Closing Gaps: An Imputation Analysis of ICU Vital Signs

arXiv:2510.24217v1 Announce Type: cross Abstract: As more Intensive Care Unit (ICU) data becomes available, the interest in developing clinical prediction models to improve healthcare protocols increases. However, the lack of data quality still hinders clinical prediction using Machine Learning (ML). Many vital sign measurements, such as heart rate, contain sizeable missing segments, leaving gaps in the data that could negatively impact prediction performance. Previous works have introduced numerous time-series imputation techniques. Nevertheless, more comprehensive work is needed to compare a representative set of methods for imputing ICU vital signs and determine the best practice. In reality, ad-hoc imputation techniques that could decrease prediction accuracy, like zero imputation, are still used. In this work, we compare established imputation techniques to guide researchers in improving the performance of clinical prediction models by selecting the most accurate imputation technique. We introduce an extensible and reusable benchmark with currently 15 imputation and 4 amputation methods, created for benchmarking on major ICU datasets. We hope to provide a comparative basis and facilitate further ML development to bring more models into clinical practice.

The Role of Omentin in Gastrointestinal Cancer: Diagnostic, Prognostic, and Therapeutic Perspectives

Metabolites. 2025 Sep 30;15(10):649. doi: 10.3390/metabo15100649.

ABSTRACT

Background/Objectives: Omentin, also known as intelectin-1, is a secreted adipokine with anti-inflammatory, insulin-sensitizing, and immune-modulatory functions, primarily expressed in visceral adipose tissue. While omentin has been associated with favorable metabolic outcomes, its role in cancer pathogenesis appears context-dependent and remains poorly understood. This review investigates the biological functions, expression patterns, and clinical relevance of omentin across gastrointestinal malignancies. Methods: A comprehensive review of the literature was conducted using PubMed, Scopus, and Web of Science up to August 2025 to evaluate the role of omentin in gastrointestinal cancers. Both preclinical and clinical studies evaluating omentin, its analogues and omentin-enhancing agents in gastric, colorectal, hepatic, pancreatic, and esophageal cancers were included. Results: Omentin exhibits anti-proliferative, anti-inflammatory, and anti-angiogenic effects within the tumor microenvironment in several GI malignancies. However, evidence also indicates a dual role. High intratumoral omentin expression correlates with improved prognosis in colorectal, gastric, and hepatic cancers; in contrast, elevated circulating levels-particularly in colorectal and pancreatic cancers-have been paradoxically associated with increased cancer risk and poor outcomes. Mechanistically, omentin modulates PI3K/Akt, NF-κB, AMPK, and oxidative stress pathways, and interacts with TMEM207. However, most available studies are small-scale and heterogeneous, with methodological inconsistencies and limited multi-omics integration, leaving major knowledge gaps. Conclusions: This review highlights omentin's distinct systemic and local roles across GI cancers, underscoring its translational implications. Omentin emerges as a promising but context-dependent biomarker and therapeutic target, with future research needed to address heterogeneity, standardize assays, and validate its clinical utility in large-scale prospective studies.

PMID:41149627 | PMC:PMC12566161 | DOI:10.3390/metabo15100649

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

arXiv:2510.21293v2 Announce Type: replace Abstract: Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. However, current research often adopts techno-centric approaches, focusing primarily on technical attributes such as reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality

arXiv:2503.05860v2 Announce Type: replace-cross Abstract: Benchmarks are essential for unified evaluation and reproducibility. The rapid rise of Artificial Intelligence for Software Engineering (AI4SE) has produced numerous benchmarks for tasks such as code generation and bug repair. However, this proliferation has led to major challenges: (1) fragmented knowledge across tasks, (2) difficulty in selecting contextually relevant benchmarks, (3) lack of standardization in benchmark creation, and (4) flaws that limit utility. Addressing these requires a dual approach: systematically mapping existing benchmarks for informed selection and defining unified guidelines for robust, adaptable benchmark development. We conduct a review of 247 studies, identifying 273 AI4SE benchmarks since 2014. We categorize them, analyze limitations, and expose gaps in current practices. Building on these insights, we introduce BenchScout, an extensible semantic search tool for locating suitable benchmarks. BenchScout employs automated clustering with contextual embeddings of benchmark-related studies, followed by dimensionality reduction. In a user study with 22 participants, BenchScout achieved usability, effectiveness, and intuitiveness scores of 4.5, 4.0, and 4.1 out of 5. To improve benchmarking standards, we propose BenchFrame, a unified framework for enhancing benchmark quality. Applying BenchFrame to HumanEval yielded HumanEvalNext, featuring corrected errors, improved language conversion, higher test coverage, and greater difficulty. Evaluating 10 state-of-the-art code models on HumanEval, HumanEvalPlus, and HumanEvalNext revealed average pass-at-1 drops of 31.22% and 19.94%, respectively, underscoring the need for continuous benchmark refinement. We further examine BenchFrame's scalability through an agentic pipeline and confirm its generalizability on the MBPP dataset. All review data, user study materials, and enhanced benchmarks are publicly released.

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

ProfileXAI: User-Adaptive Explainable AI

arXiv:2510.22998v1 Announce Type: new Abstract: ProfileXAI is a model- and domain-agnostic framework that couples post-hoc explainers (SHAP, LIME, Anchor) with retrieval - augmented LLMs to produce explanations for different types of users. The system indexes a multimodal knowledge base, selects an explainer per instance via quantitative criteria, and generates grounded narratives with chat-enabled prompting. On Heart Disease and Thyroid Cancer datasets, we evaluate fidelity, robustness, parsimony, token use, and perceived quality. No explainer dominates: LIME achieves the best fidelity--robustness trade-off (Infidelity $\le 0.30$, $L
❌