❌

Reading view

An Agentic Framework for Rapid Deployment of Edge AI Solutions in Industry 5.0

arXiv:2510.25813v1 Announce Type: new Abstract: We present a novel framework for Industry 5.0 that simplifies the deployment of AI models on edge devices in various industrial settings. The design reduces latency and avoids external data transfer by enabling local inference and real-time processing. Our implementation is agent-based, which means that individual agents, whether human, algorithmic, or collaborative, are responsible for well-defined tasks, enabling flexibility and simplifying integration. Moreover, our framework supports modular integration and maintains low resource requirements. Preliminary evaluations concerning the food industry in real scenarios indicate improved deployment time and system adaptability performance. The source code is publicly available at https://github.com/AI-REDGIO-5-0/ci-component.
  •  

Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling

arXiv:2510.26603v1 Announce Type: new Abstract: The electricity sector transition requires substantial increases in residential demand response capacity, yet Home Energy Management Systems (HEMS) adoption remains limited by user interaction barriers requiring translation of everyday preferences into technical parameters. While large language models have been applied to energy systems as code generators and parameter extractors, no existing implementation deploys LLMs as autonomous coordinators managing the complete workflow from natural language input to multi-appliance scheduling. This paper presents an agentic AI HEMS where LLMs autonomously coordinate multi-appliance scheduling from natural language requests to device control, achieving optimal scheduling without example demonstrations. A hierarchical architecture combining one orchestrator with three specialist agents uses the ReAct pattern for iterative reasoning, enabling dynamic coordination without hardcoded workflows while integrating Google Calendar for context-aware deadline extraction. Evaluation across three open-source models using real Austrian day-ahead electricity prices reveals substantial capability differences. Llama-3.3-70B successfully coordinates all appliances across all scenarios to match cost-optimal benchmarks computed via mixed-integer linear programming, while other models achieve perfect single-appliance performance but struggle to coordinate all appliances simultaneously. Progressive prompt engineering experiments demonstrate that analytical query handling without explicit guidance remains unreliable despite models' general reasoning capabilities. We open-source the complete system including orchestration logic, agent prompts, tools, and web interfaces to enable reproducibility, extension, and future research.
  •  

Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world

arXiv:2510.25819v1 Announce Type: cross Abstract: The rapid rise of AI agents presents urgent challenges in authentication, authorization, and identity management. Current agent-centric protocols (like MCP) highlight the demand for clarified best practices in authentication and authorization. Looking ahead, ambitions for highly autonomous agents raise complex long-term questions regarding scalable access control, agent-centric identities, AI workload differentiation, and delegated authority. This OpenID Foundation whitepaper is for stakeholders at the intersection of AI agents and access management. It outlines the resources already available for securing today's agents and presents a strategic agenda to address the foundational authentication, authorization, and identity problems pivotal for tomorrow's widespread autonomous systems.
  •  

Multi-Agent Reinforcement Learning for Market Making: Competition without Collusion

arXiv:2510.25929v1 Announce Type: cross Abstract: Algorithmic collusion has emerged as a central question in AI: Will the interaction between different AI agents deployed in markets lead to collusion? More generally, understanding how emergent behavior, be it a cartel or market dominance from more advanced bots, affects the market overall is an important research question. We propose a hierarchical multi-agent reinforcement learning framework to study algorithmic collusion in market making. The framework includes a self-interested market maker (Agent~A), which is trained in an uncertain environment shaped by an adversary, and three bottom-layer competitors: the self-interested Agent~B1 (whose objective is to maximize its own PnL), the competitive Agent~B2 (whose objective is to minimize the PnL of its opponent), and the hybrid Agent~B$^\star$, which can modulate between the behavior of the other two. To analyze how these agents shape the behavior of each other and affect market outcomes, we propose interaction-level metrics that quantify behavioral asymmetry and system-level dynamics, while providing signals potentially indicative of emergent interaction patterns. Experimental results show that Agent~B2 secures dominant performance in a zero-sum setting against B1, aggressively capturing order flow while tightening average spreads, thus improving market execution efficiency. In contrast, Agent~B$^\star$ exhibits a self-interested inclination when co-existing with other profit-seeking agents, securing dominant market share through adaptive quoting, yet exerting a milder adverse impact on the rewards of Agents~A and B1 compared to B2. These findings suggest that adaptive incentive control supports more sustainable strategic co-existence in heterogeneous agent environments and offers a structured lens for evaluating behavioral design in algorithmic trading systems.
  •  

Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning

arXiv:2510.25992v1 Announce Type: cross Abstract: Large Language Models (LLMs) often struggle with problems that require multi-step reasoning. For small-scale open-source models, Reinforcement Learning with Verifiable Rewards (RLVR) fails when correct solutions are rarely sampled even after many attempts, while Supervised Fine-Tuning (SFT) tends to overfit long demonstrations through rigid token-by-token imitation. To address this gap, we propose Supervised Reinforcement Learning (SRL), a framework that reformulates problem solving as generating a sequence of logical "actions". SRL trains the model to generate an internal reasoning monologue before committing to each action. It provides smoother rewards based on the similarity between the model's actions and expert actions extracted from the SFT dataset in a step-wise manner. This supervision offers richer learning signals even when all rollouts are incorrect, while encouraging flexible reasoning guided by expert demonstrations. As a result, SRL enables small models to learn challenging problems previously unlearnable by SFT or RLVR. Moreover, initializing training with SRL before refining with RLVR yields the strongest overall performance. Beyond reasoning benchmarks, SRL generalizes effectively to agentic software engineering tasks, establishing it as a robust and versatile training framework for reasoning-oriented LLMs.
  •  

The Quest for Reliable Metrics of Responsible AI

arXiv:2510.26007v1 Announce Type: cross Abstract: The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quantified through evaluation metrics, yet there has been less work on assessing the robustness and reliability of the metrics themselves. We reflect on prior work that examines the robustness of fairness metrics for recommender systems as a type of AI application and summarise their key takeaways into a set of non-exhaustive guidelines for developing reliable metrics of responsible AI. Our guidelines apply to a broad spectrum of AI applications, including AIS.
  •  

A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI

arXiv:2510.26275v1 Announce Type: cross Abstract: Generative AI (GenAI) is rapidly transforming software engineering (SE) practices, influencing how SE processes are executed, as well as how software systems are developed, operated, and evolved. This paper applies design science research to build a roadmap for GenAI-augmented SE. The process consists of three cycles that incrementally integrate multiple sources of evidence, including collaborative discussions from the FSE 2025 "Software Engineering 2030" workshop, rapid literature reviews, and external feedback sessions involving peers. McLuhan's tetrads were used as a conceptual instrument to systematically capture the transforming effects of GenAI on SE processes and software products.The resulting roadmap identifies four fundamental forms of GenAI augmentation in SE and systematically characterizes their related research challenges and opportunities. These insights are then consolidated into a set of future research directions. By grounding the roadmap in a rigorous multi-cycle process and cross-validating it among independent author teams and peers, the study provides a transparent and reproducible foundation for analyzing how GenAI affects SE processes, methods and tools, and for framing future research within this rapidly evolving area. Based on these findings, the article finally makes ten predictions for SE in the year 2030.
  •  

Multi-Agent Evolve: LLM Self-Improve through Co-evolution

arXiv:2510.23595v3 Announce Type: replace Abstract: Reinforcement Learning (RL) has demonstrated significant potential in enhancing the reasoning capabilities of large language models (LLMs). However, the success of RL for LLMs heavily relies on human-curated datasets and verifiable rewards, which limit their scalability and generality. Recent Self-Play RL methods, inspired by the success of the paradigm in games and Go, aim to enhance LLM reasoning capabilities without human-annotated data. However, their methods primarily depend on a grounded environment for feedback (e.g., a Python interpreter or a game engine); extending them to general domains remains challenging. To address these challenges, we propose Multi-Agent Evolve (MAE), a framework that enables LLMs to self-evolve in solving diverse tasks, including mathematics, reasoning, and general knowledge Q&A. The core design of MAE is based on a triplet of interacting agents (Proposer, Solver, Judge) that are instantiated from a single LLM, and applies reinforcement learning to optimize their behaviors. The Proposer generates questions, the Solver attempts solutions, and the Judge evaluates both while co-evolving. Experiments on Qwen2.5-3B-Instruct demonstrate that MAE achieves an average improvement of 4.54% on multiple benchmarks. These results highlight MAE as a scalable, data-efficient method for enhancing the general reasoning abilities of LLMs with minimal reliance on human-curated supervision.
  •  

Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models

arXiv:2406.05948v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs), especially those accessed via APIs, have demonstrated impressive capabilities across various domains. However, users without technical expertise often turn to (untrustworthy) third-party services, such as prompt engineering, to enhance their LLM experience, creating vulnerabilities to adversarial threats like backdoor attacks. Backdoor-compromised LLMs generate malicious outputs to users when inputs contain specific "triggers" set by attackers. Traditional defense strategies, originally designed for small-scale models, are impractical for API-accessible LLMs due to limited model access, high computational costs, and data requirements. To address these limitations, we propose Chain-of-Scrutiny (CoS) which leverages LLMs' unique reasoning abilities to mitigate backdoor attacks. It guides the LLM to generate reasoning steps for a given input and scrutinizes for consistency with the final output -- any inconsistencies indicating a potential attack. It is well-suited for the popular API-only LLM deployments, enabling detection at minimal cost and with little data. User-friendly and driven by natural language, it allows non-experts to perform the defense independently while maintaining transparency. We validate the effectiveness of CoS through extensive experiments on various tasks and LLMs, with results showing greater benefits for more powerful LLMs.
  •  

Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM

arXiv:2410.14879v4 Announce Type: replace-cross Abstract: Passive tracking methods, such as phone and wearable sensing, have become dominant in monitoring human behaviors in modern ubiquitous computing studies. While there have been significant advances in machine-learning approaches to translate periods of raw sensor data to model momentary behaviors, (e.g., physical activity recognition), there still remains a significant gap in the translation of these sensing streams into meaningful, high-level, context-aware insights that are required for various applications (e.g., summarizing an individual's daily routine). To bridge this gap, experts often need to employ a context-driven sensemaking process in real-world studies to derive insights. This process often requires manual effort and can be challenging even for experienced researchers due to the complexity of human behaviors. We conducted three rounds of user studies with 21 experts to explore solutions to address challenges with sensemaking. We follow a human-centered design process to identify needs and design, iterate, build, and evaluate Vital Insight (VI), a novel, LLM-assisted, prototype system to enable human-in-the-loop inference (sensemaking) and visualizations of multi-modal passive sensing data from smartphones and wearables. Using the prototype as a technology probe, we observe experts' interactions with it and develop an expert sensemaking model that explains how experts move between direct data representations and AI-supported inferences to explore, question, and validate insights. Through this iterative process, we also synthesize and discuss a list of design implications for the design of future AI-augmented visualization systems to better assist experts' sensemaking processes in multi-modal health sensing data.
  •  

Epistemic Diversity and Knowledge Collapse in Large Language Models

arXiv:2510.04226v4 Announce Type: replace-cross Abstract: Large language models (LLMs) tend to generate lexically, semantically, and stylistically homogenous texts. This poses a risk of knowledge collapse, where homogenous LLMs mediate a shrinking in the range of accessible information over time. Existing works on homogenization are limited by a focus on closed-ended multiple-choice setups or fuzzy semantic features, and do not look at trends across time and cultural contexts. To overcome this, we present a new methodology to measure epistemic diversity, i.e., variation in real-world claims in LLM outputs, which we use to perform a broad empirical study of LLM knowledge collapse. We test 27 LLMs, 155 topics covering 12 countries, and 200 prompt variations sourced from real user chats. For the topics in our study, we show that while newer models tend to generate more diverse claims, nearly all models are less epistemically diverse than a basic web search. We find that model size has a negative impact on epistemic diversity, while retrieval-augmented generation (RAG) has a positive impact, though the improvement from RAG varies by the cultural context. Finally, compared to a traditional knowledge source (Wikipedia), we find that country-specific claims reflect the English language more than the local one, highlighting a gap in epistemic representation
  •  

Integrating Genomics into Multimodal EHR Foundation Models

arXiv:2510.23639v2 Announce Type: replace-cross Abstract: This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health profiles. Leveraging the extensive and diverse data from the All of Us (AoU) Research Program, this multimodal framework aims to learn complex relationships between clinical data and genetic predispositions. The methodology extends advancements in generative AI to the EHR foundation model space, enhancing predictive capabilities and interpretability. Evaluation on AoU data demonstrates the model's predictive value for the onset of various conditions, particularly Type 2 Diabetes (T2D), and illustrates the interplay between PRS and EHR data. The work also explores transfer learning for custom classification tasks, showcasing the architecture's versatility and efficiency. This approach is pivotal for unlocking new insights into disease prediction, proactive health management, risk stratification, and personalized treatment strategies, laying the groundwork for more personalized, equitable, and actionable real-world evidence generation in healthcare.
  •  

Nanomaterial-assisted immunodiagnostic profiling and therapeutic targeting of hepatocellular carcinoma: from molecular biomarkers to clinical applications

Front Immunol. 2025 Oct 14;16:1668630. doi: 10.3389/fimmu.2025.1668630. eCollection 2025.

ABSTRACT

AIMS AND OBJECTIVES: This study aimed to identify immunologically relevant transcriptomic and proteomic biomarkers in hepatocellular carcinoma (HCC) and to characterize their B-cell epitopes for potential integration into nanomaterial-based biosensors and immunomodulatory platforms for early diagnosis and targeted therapy.

METHODS: We conducted a comprehensive multi-omics analysis by integrating transcriptomic (TCGA-LIHC) and proteomic data to identify differentially expressed genes (DEGs) in HCC. Protein-protein interaction networks and pathway enrichment were used to prioritize hub genes. Five candidate biomarkers, RFC2, HSP90AB1, YWHAZ, CYP2E1, and ADH4, were selected for qRT-PCR and serum ELISA validation in clinical cohorts comprising 85 HCC patients and 50 healthy controls. B-cell epitope prediction was performed using BepiPred 2.0 and validated through synthetic peptide-based ELISA in the same cohort to assess immunoreactivity. Diagnostic performance was evaluated using ROC curve analysis.

RESULTS: RFC2, HSP90AB1, and YWHAZ were significantly upregulated (|log2FC|>0.2) and showed high serological expression, whereas CYP2E1 and ADH4 were consistently downregulated. Predicted B-cell epitopes from RFC2, HSP90AB1, and YWHAZ exhibited strong immunoreactivity (AUC>0.84), indicating their diagnostic potential. Enrichment analysis revealed that upregulated DEGs were involved in cell cycle and mitotic progression, while downregulated genes were linked to immune suppression and metabolic dysfunction. These validated immunogenic epitopes offer promising anchors for nanomaterial-functionalized biosensors, such as gold nanoparticle-conjugated ELISA, graphene-based electrochemical platforms, and peptide-coated quantum dots, for ultrasensitive and multiplexed HCC detection.

CONCLUSION: By integrating transcriptomic and proteomic screening with epitope-level validation, we identified a novel panel of immunogenic biomarkers suitable for nanomaterial-enabled diagnostics in HCC. These findings support the translational potential of peptide-nano scaffold conjugates in developing minimally invasive, immune-responsive biosensing and therapeutic tools tailored for early-stage liver cancer management.

PMID:41164201 | PMC:PMC12558944 | DOI:10.3389/fimmu.2025.1668630

  •  

Building a high performance data and AI organization (2nd edition)

Four years is a lifetime when it comes to artificial intelligence. Since the first edition of this study was published in 2021, AI’s capabilities have been advancing at speed, and the advances have not slowed since generative AI’s breakthrough. For example, multimodality— the ability to process information not only as text but also as audio, video, and other unstructured formats—is becoming a common feature of AI models. AI’s capacity to reason and act autonomously has also grown, and organizations are now starting to work with AI agents that can do just that.

Amid all the change, there remains a constant: the quality of an AI model’s outputs is only ever as good as the data
that feeds it. Data management technologies and practices have also been advancing, but the second edition of this study suggests that most organizations are not leveraging those fast enough to keep up with AI’s development. As a result of that and other hindrances, relatively few organizations are delivering the desired business results from their AI strategy. No more than 2% of senior executives we surveyed rate their organizations highly in terms of delivering results from AI.

To determine the extent to which organizational data performance has improved as generative AI and other AI advances have taken hold, MIT Technology Review Insights surveyed 800 senior data and technology executives. We also conducted in-depth interviews with 15 technology and business leaders.

Key findings from the report include the following:

• Few data teams are keeping pace with AI. Organizations are doing no better today at delivering on data strategy than in pre-generative AI days. Among those surveyed in 2025, 12% are self-assessed data “high achievers” compared with 13% in 2021. Shortages of skilled talent remain a constraint, but teams also struggle with accessing fresh data, tracing lineage, and dealing with security complexity—important requirements for AI success.

• Partly as a result, AI is not fully firing yet. There are even fewer “high achievers” when it comes to AI. Just 2% of respondents rate their organizations’ AI performance highly today in terms of delivering measurable business results. In fact, most are still struggling to scale generative AI. While two thirds have deployed it, only 7% have done so widely.

Download the report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

  •  

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.
  •  

Advancing Non-Small-Cell Lung Cancer Management Through Multi-Omics Integration: Insights from Genomics, Metabolomics, and Radiomics

Diagnostics (Basel). 2025 Oct 14;15(20):2586. doi: 10.3390/diagnostics15202586.

ABSTRACT

The integration of multi-omics technologies is transforming the landscape of cancer management, offering unprecedented insights into tumor biology, early diagnosis, and personalized therapy. This review provides a comprehensive overview of the current state of omics approaches, with a particular focus on the application of genomics, NMR-based metabolomics, and radiomics in non-small cell lung cancer (NSCLC). Genomics currently represents one of the most established omics technologies in oncology, as it enables the identification of genetic alterations that drive tumor initiation, progression, and therapeutic response. Interestingly, genomic analyses have revealed that many tumors harbor mutations in genes encoding metabolic enzymes, thus establishing a tight connection between genomics and tumor metabolism. In parallel, metabolomics profiling-by capturing the metabolic phenotype of tumors-has, in recent years, identified specific biomarkers associated with tumor burden, progression, and prognosis. Such findings have catalyzed growing interest in metabolomics as a complementary approach to better characterize cancer biology and discover novel diagnostic and therapeutic targets. Moreover, radiomics, through the extraction of quantitative features from standard imaging modalities, captures tumor heterogeneity and contributes predictive information on tumor biology, treatment response, and clinical outcomes. As a non-invasive and widely available technique, radiomics has the potential to support longitudinal monitoring and individualized treatment planning. Both metabolomics and radiomics, when integrated with genomic data, could support a more comprehensive understanding of NSCLC and pave the way for the development of non-invasive, predictive models and personalized therapeutic strategies. In addition, we explore the specific contributions of these technologies in enhancing clinical decision-making for lung cancer patients, with particular attention to their potential in early diagnosis, treatment selection, and real-time monitoring.

PMID:41153258 | DOI:10.3390/diagnostics15202586

  •  

Prospective proteomics for discovering biomarkers in lung adenocarcinoma: a literature review

Transl Cancer Res. 2025 Sep 30;14(9):6102-6117. doi: 10.21037/tcr-2025-1092. Epub 2025 Sep 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung adenocarcinoma (LUAD), as the main subtype of non-small cell lung cancer (NSCLC), faces clinical challenges including molecular heterogeneity, late diagnosis, and aggressive growth, leading to a low 5-year survival rate. Biomarkers are critical for early detection, accurate differentiation of benign/malignant lesions, and guiding personalized treatment strategies. Proteomic technologies using liquid biopsy show potential by analyzing protein changes and post-translational modifications (PTMs) to identify novel biomarkers and unravel cancer mechanisms. This review examines proteomic advances in LUAD, compares platform strengths, lists validated protein markers, and discusses challenges like specificity and regulations. It aims to develop a precision medicine framework by integrating multi-omics data for improved diagnosis and treatment.

METHODS: This study conducted a literature review by searching the PubMed and Web of Science databases for original articles written in English from 2002 to 2025, using the keywords "lung adenocarcinoma" OR "LUAD" AND "biomarkers" AND "proteomics" OR "SomaScan" OR "spatial proteomics" to identify the latest research findings in the field of proteomics technology and LUAD biomarkers. The included studies mainly focused on the current landscape of biomarkers in the diagnosis, treatment, and prognosis of LUAD.

KEY CONTENT AND FINDINGS: This review discusses high-throughput methods for comprehensive protein profiling in accessible biospecimens (tissues, blood, urine) to identify biomarkers for LUAD. We systematically evaluate emerging proteomic strategies, including mass spectrometry (MS), proximity extension assays (PEAs), spatial proteomics techniques, and SomaScan platforms-coupled with innovative computational frameworks have revolutionized biomarkers discovery and their translational potential in developing precision diagnostics and targeted therapies. Additionally, the review addresses challenges in integrating proteomics with genomics, transcriptomics, and metabolomics, offering new methodologies and expanding research in life sciences. As technological advancements continue, it is anticipated that more potential biomarkers will be conducted to validate the broader application in LUAD treatment, addressing early-stage disease complexities and aiding in selecting more effective treatment strategies.

CONCLUSIONS: By synthesizing cutting-edge evidence on proteome-driven LUAD biomarkers, this review elucidates actionable strategies to refine early detection protocols and mechanism-informed personalized treatment frameworks, directly advancing precision oncology initiatives for this prevalent malignancy through biomarker-guided clinical decision-making and multi-omics integration.

PMID:41158224 | PMC:PMC12554480 | DOI:10.21037/tcr-2025-1092

  •  

STAT+: Natera, known for spotting cancer recurrence, wades into early detection

Want to stay on top of the science and politics driving biotech today? Sign up to get our biotech newsletter in your inbox.

Good morning. It seems everyone I know has been getting sick lately — hope you are all taking care of yourselves! Onto the news today.

BridgeBio notches another Phase 3 win

BridgeBio said this morning that its investigational drug succeeded in a late-stage trial of patients with autosomal dominant hypocalcemia type 1, a rare genetic condition that causes low calcium levels in the blood.

Continue to STAT+ to read the full story…

© Adobe

  •  
❌