❌

Reading view

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.
  •  

Advancing Non-Small-Cell Lung Cancer Management Through Multi-Omics Integration: Insights from Genomics, Metabolomics, and Radiomics

Diagnostics (Basel). 2025 Oct 14;15(20):2586. doi: 10.3390/diagnostics15202586.

ABSTRACT

The integration of multi-omics technologies is transforming the landscape of cancer management, offering unprecedented insights into tumor biology, early diagnosis, and personalized therapy. This review provides a comprehensive overview of the current state of omics approaches, with a particular focus on the application of genomics, NMR-based metabolomics, and radiomics in non-small cell lung cancer (NSCLC). Genomics currently represents one of the most established omics technologies in oncology, as it enables the identification of genetic alterations that drive tumor initiation, progression, and therapeutic response. Interestingly, genomic analyses have revealed that many tumors harbor mutations in genes encoding metabolic enzymes, thus establishing a tight connection between genomics and tumor metabolism. In parallel, metabolomics profiling-by capturing the metabolic phenotype of tumors-has, in recent years, identified specific biomarkers associated with tumor burden, progression, and prognosis. Such findings have catalyzed growing interest in metabolomics as a complementary approach to better characterize cancer biology and discover novel diagnostic and therapeutic targets. Moreover, radiomics, through the extraction of quantitative features from standard imaging modalities, captures tumor heterogeneity and contributes predictive information on tumor biology, treatment response, and clinical outcomes. As a non-invasive and widely available technique, radiomics has the potential to support longitudinal monitoring and individualized treatment planning. Both metabolomics and radiomics, when integrated with genomic data, could support a more comprehensive understanding of NSCLC and pave the way for the development of non-invasive, predictive models and personalized therapeutic strategies. In addition, we explore the specific contributions of these technologies in enhancing clinical decision-making for lung cancer patients, with particular attention to their potential in early diagnosis, treatment selection, and real-time monitoring.

PMID:41153258 | DOI:10.3390/diagnostics15202586

  •  

Improving Recruitment Into Research Studies via Electronically Collected Patient-Entered Data: Mixed Methods Study

Background: Patient recruitment remains a critical challenge in clinical research. Although the integration of electronically collected patient-entered data within clinical practices enables innovative recruitment approaches, existing methods present challenges such as increased patient burden and potential violation of autonomy. A more nuanced approach involves identifying patient attributes associated with higher propensity for research participation, enabling research teams to efficiently prioritize outreach efforts. Objective: This study aims to (1) develop patient-reported questions reflecting perceptions about research participation and (2) determine whether patient responses are predictive of interest in joining a precision medicine registry. Methods: This mixed methods study used an exploratory sequential design in 2 phases. Phase 1 involved cognitive interviews with 32 patients recruited through the Cleveland Clinic Healthcare Partners program to develop “research perception” questions. Participants evaluated 9 candidate questions that were based on a literature review of research participation factors. Three questions were selected for implementation. Phase 2 was a cross-sectional cohort study incorporating these 3 questions into routine electronic questionnaires completed by primary care patients through the patient portal. The study population included 1077 patients who completed both “research perception” and “research recruitment” questions between August 2018 and April 2019. Diagnostic accuracy was assessed using receiver operating characteristic curve analysis, and multivariable logistic regression models evaluated associations while adjusting for demographic and health factors. Results: Phase 1 revealed strong research support among participants, with 97% (31/32) agreeing that research should be part of the institution’s mission and 100% (32/32) affirming that research enhances patient care. Phase 2 included 1077 patients (mean age 48.3, SD 16.3 years; 625/1065 female, 58.68%; 661/1005 White, 65.77%), of whom 278 (25.8%) expressed interest in being contacted about the precision medicine registry. Patients expressing interest were older and had worse self-reported health, more depressive symptoms, and greater social needs. “Strongly agree” and “very important” responses to any “research perception” question were significantly associated with study interest, with adjusted odds ratios ranging from 6.36 (95% CI 2.77-14.6) to 17.6 (95% CI 5.08-61.1; P<.001). The “research perception” questions demonstrated high sensitivity (>80%) but limited specificity (24%-31%). Conclusions: Patient-reported questions assessing research participation likelihood can help identify patients more likely to enroll in clinical studies. This approach enables effective recruitment prioritization while preserving patient autonomy and reducing patient burden. High sensitivity makes these questions valuable as screening tools, although limited specificity suggests use for prioritizing rather than excluding participants. Further validation across different trial types and populations is warranted.
  •  

Test-Time Tuned Language Models Enable End-to-end De Novo Molecular Structure Generation from MS/MS Spectra

arXiv:2510.23746v1 Announce Type: new Abstract: Tandem Mass Spectrometry enables the identification of unknown compounds in crucial fields such as metabolomics, natural product discovery and environmental analysis. However, current methods rely on database matching from previously observed molecules, or on multi-step pipelines that require intermediate fragment or fingerprint prediction. This makes finding the correct molecule highly challenging, particularly for compounds absent from reference databases. We introduce a framework that, by leveraging test-time tuning, enhances the learning of a pre-trained transformer model to address this gap, enabling end-to-end de novo molecular structure generation directly from the tandem mass spectra and molecular formulae, bypassing manual annotations and intermediate steps. We surpass the de-facto state-of-the-art approach DiffMS on two popular benchmarks NPLIB1 and MassSpecGym by 100% and 20%, respectively. Test-time tuning on experimental spectra allows the model to dynamically adapt to novel spectra, and the relative performance gain over conventional fine-tuning is of 62% on MassSpecGym. When predictions deviate from the ground truth, the generated molecular candidates remain structurally accurate, providing valuable guidance for human interpretation and more reliable identification.
  •  

Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives

arXiv:2510.24551v1 Announce Type: new Abstract: Generative Artificial Intelligence (GenAI) is taking the world by storm. It promises transformative opportunities for advancing and disrupting existing practices, including healthcare. From large language models (LLMs) for clinical note synthesis and conversational assistance to multimodal systems that integrate medical imaging, electronic health records, and genomic data for decision support, GenAI is transforming the practice of medicine and the delivery of healthcare, such as diagnosis and personalized treatments, with great potential in reducing the cognitive burden on clinicians, thereby improving overall healthcare delivery. However, GenAI deployment in healthcare requires an in-depth understanding of healthcare tasks and what can and cannot be achieved. In this paper, we propose a data-centric paradigm in the design and deployment of GenAI systems for healthcare. Specifically, we reposition the data life cycle by making the medical data ecosystem as the foundational substrate for generative healthcare systems. This ecosystem is designed to sustainably support the integration, representation, and retrieval of diverse medical data and knowledge. With effective and efficient data processing pipelines, such as semantic vector search and contextual querying, it enables GenAI-powered operations for upstream model components and downstream clinical applications. Ultimately, it not only supplies foundation models with high-quality, multimodal data for large-scale pretraining and domain-specific fine-tuning, but also serves as a knowledge retrieval backend to support task-specific inference via the agentic layer. The ecosystem enables the deployment of GenAI for high-quality and effective healthcare delivery.
  •  

Integrating Genomics into Multimodal EHR Foundation Models

arXiv:2510.23639v1 Announce Type: cross Abstract: This paper introduces an innovative Electronic Health Record (EHR) foundation model that integrates Polygenic Risk Scores (PRS) as a foundational data modality, moving beyond traditional EHR-only approaches to build more holistic health profiles. Leveraging the extensive and diverse data from the All of Us (AoU) Research Program, this multimodal framework aims to learn complex relationships between clinical data and genetic predispositions. The methodology extends advancements in generative AI to the EHR foundation model space, enhancing predictive capabilities and interpretability. Evaluation on AoU data demonstrates the model's predictive value for the onset of various conditions, particularly Type 2 Diabetes (T2D), and illustrates the interplay between PRS and EHR data. The work also explores transfer learning for custom classification tasks, showcasing the architecture's versatility and efficiency. This approach is pivotal for unlocking new insights into disease prediction, proactive health management, risk stratification, and personalized treatment strategies, laying the groundwork for more personalized, equitable, and actionable real-world evidence generation in healthcare.
  •  

Closing Gaps: An Imputation Analysis of ICU Vital Signs

arXiv:2510.24217v1 Announce Type: cross Abstract: As more Intensive Care Unit (ICU) data becomes available, the interest in developing clinical prediction models to improve healthcare protocols increases. However, the lack of data quality still hinders clinical prediction using Machine Learning (ML). Many vital sign measurements, such as heart rate, contain sizeable missing segments, leaving gaps in the data that could negatively impact prediction performance. Previous works have introduced numerous time-series imputation techniques. Nevertheless, more comprehensive work is needed to compare a representative set of methods for imputing ICU vital signs and determine the best practice. In reality, ad-hoc imputation techniques that could decrease prediction accuracy, like zero imputation, are still used. In this work, we compare established imputation techniques to guide researchers in improving the performance of clinical prediction models by selecting the most accurate imputation technique. We introduce an extensible and reusable benchmark with currently 15 imputation and 4 amputation methods, created for benchmarking on major ICU datasets. We hope to provide a comparative basis and facilitate further ML development to bring more models into clinical practice.
  •  

Dynamic Monitoring of Recurrent Ovarian Cancer Using Serial ctDNA: A Real-World Case Series

Curr Oncol. 2025 Oct 21;32(10):585. doi: 10.3390/curroncol32100585.

ABSTRACT

Recurrent ovarian cancer (OC) is challenging to detect early using current methods like CA-125 and imaging. Circulating tumor DNA (ctDNA) may improve disease monitoring. Here, we assess the real-world clinical utility of serial ctDNA analyses in patients with recurrent OC. We analyzed serial plasma samples (N = 23) from six patients with recurrent OC using a tumor-informed next-generation sequencing assay targeting 68 cancer-related genes developed at the University of Washington. ctDNA variant allele frequencies (VAFs) were correlated with CA-125 levels, radiographic findings, and clinical outcomes. ctDNA levels generally reflected clinical status, accurately mirroring disease progression and therapeutic response. In one patient, rising ctDNA preceded clinical recurrence by four months, despite normal CA-125 and imaging, highlighting its potential advantage. Conversely, some patients exhibited clinical progression with undetectable ctDNA, indicating limitations in assay sensitivity, biological factors, or metastatic sites (e.g., brain metastases). ctDNA and CA-125 showed complementary value in most cases, suggesting potential combined use in clinical monitoring. Our findings demonstrate that ctDNA is a promising biomarker to complement existing monitoring approaches for recurrent OC. In some cases, capable of predicting relapse and treatment response ahead of current clinical indicators. However, identified discordances underscore technical and biological challenges that warrant further investigation. Larger prospective studies are necessary to refine ctDNA's clinical utility and integration into personalized OC care.

PMID:41149505 | PMC:PMC12563156 | DOI:10.3390/curroncol32100585

  •  

Tongyi DeepResearch Technical Report

arXiv:2510.24701v1 Announce Type: cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically designed for long-horizon, deep information-seeking research tasks. To incentivize autonomous deep research agency, Tongyi DeepResearch is developed through an end-to-end training framework that combines agentic mid-training and agentic post-training, enabling scalable reasoning and information seeking across complex tasks. We design a highly scalable data synthesis pipeline that is fully automatic, without relying on costly human annotation, and empowers all training stages. By constructing customized environments for each stage, our system enables stable and consistent interactions throughout. Tongyi DeepResearch, featuring 30.5 billion total parameters, with only 3.3 billion activated per token, achieves state-of-the-art performance across a range of agentic deep research benchmarks, including Humanity's Last Exam, BrowseComp, BrowseComp-ZH, WebWalkerQA, xbench-DeepSearch, FRAMES and xbench-DeepSearch-2510. We open-source the model, framework, and complete solutions to empower the community.
  •  

Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents

arXiv:2510.24702v1 Announce Type: cross Abstract: Public research results on large-scale supervised finetuning of AI agents remain relatively rare, since the collection of agent training data presents unique challenges. In this work, we argue that the bottleneck is not a lack of underlying data sources, but that a large variety of data is fragmented across heterogeneous formats, tools, and interfaces. To this end, we introduce the agent data protocol (ADP), a light-weight representation language that serves as an "interlingua" between agent datasets in diverse formats and unified agent training pipelines downstream. The design of ADP is expressive enough to capture a large variety of tasks, including API/tool use, browsing, coding, software engineering, and general agentic workflows, while remaining simple to parse and train on without engineering at a per-dataset level. In experiments, we unified a broad collection of 13 existing agent training datasets into ADP format, and converted the standardized ADP data into training-ready formats for multiple agent frameworks. We performed SFT on these data, and demonstrated an average performance gain of ~20% over corresponding base models, and delivers state-of-the-art or near-SOTA performance on standard coding, browsing, tool use, and research benchmarks, without domain-specific tuning. All code and data are released publicly, in the hope that ADP could help lower the barrier to standardized, scalable, and reproducible agent training.
  •  

The Role of Omentin in Gastrointestinal Cancer: Diagnostic, Prognostic, and Therapeutic Perspectives

Metabolites. 2025 Sep 30;15(10):649. doi: 10.3390/metabo15100649.

ABSTRACT

Background/Objectives: Omentin, also known as intelectin-1, is a secreted adipokine with anti-inflammatory, insulin-sensitizing, and immune-modulatory functions, primarily expressed in visceral adipose tissue. While omentin has been associated with favorable metabolic outcomes, its role in cancer pathogenesis appears context-dependent and remains poorly understood. This review investigates the biological functions, expression patterns, and clinical relevance of omentin across gastrointestinal malignancies. Methods: A comprehensive review of the literature was conducted using PubMed, Scopus, and Web of Science up to August 2025 to evaluate the role of omentin in gastrointestinal cancers. Both preclinical and clinical studies evaluating omentin, its analogues and omentin-enhancing agents in gastric, colorectal, hepatic, pancreatic, and esophageal cancers were included. Results: Omentin exhibits anti-proliferative, anti-inflammatory, and anti-angiogenic effects within the tumor microenvironment in several GI malignancies. However, evidence also indicates a dual role. High intratumoral omentin expression correlates with improved prognosis in colorectal, gastric, and hepatic cancers; in contrast, elevated circulating levels-particularly in colorectal and pancreatic cancers-have been paradoxically associated with increased cancer risk and poor outcomes. Mechanistically, omentin modulates PI3K/Akt, NF-κB, AMPK, and oxidative stress pathways, and interacts with TMEM207. However, most available studies are small-scale and heterogeneous, with methodological inconsistencies and limited multi-omics integration, leaving major knowledge gaps. Conclusions: This review highlights omentin's distinct systemic and local roles across GI cancers, underscoring its translational implications. Omentin emerges as a promising but context-dependent biomarker and therapeutic target, with future research needed to address heterogeneity, standardize assays, and validate its clinical utility in large-scale prospective studies.

PMID:41149627 | PMC:PMC12566161 | DOI:10.3390/metabo15100649

  •  

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

  •  

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

arXiv:2510.21293v2 Announce Type: replace Abstract: Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. However, current research often adopts techno-centric approaches, focusing primarily on technical attributes such as reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.
  •  

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

  •  

Benchmarking AI Models in Software Engineering: A Review, Search Tool, and Unified Approach for Elevating Benchmark Quality

arXiv:2503.05860v2 Announce Type: replace-cross Abstract: Benchmarks are essential for unified evaluation and reproducibility. The rapid rise of Artificial Intelligence for Software Engineering (AI4SE) has produced numerous benchmarks for tasks such as code generation and bug repair. However, this proliferation has led to major challenges: (1) fragmented knowledge across tasks, (2) difficulty in selecting contextually relevant benchmarks, (3) lack of standardization in benchmark creation, and (4) flaws that limit utility. Addressing these requires a dual approach: systematically mapping existing benchmarks for informed selection and defining unified guidelines for robust, adaptable benchmark development. We conduct a review of 247 studies, identifying 273 AI4SE benchmarks since 2014. We categorize them, analyze limitations, and expose gaps in current practices. Building on these insights, we introduce BenchScout, an extensible semantic search tool for locating suitable benchmarks. BenchScout employs automated clustering with contextual embeddings of benchmark-related studies, followed by dimensionality reduction. In a user study with 22 participants, BenchScout achieved usability, effectiveness, and intuitiveness scores of 4.5, 4.0, and 4.1 out of 5. To improve benchmarking standards, we propose BenchFrame, a unified framework for enhancing benchmark quality. Applying BenchFrame to HumanEval yielded HumanEvalNext, featuring corrected errors, improved language conversion, higher test coverage, and greater difficulty. Evaluating 10 state-of-the-art code models on HumanEval, HumanEvalPlus, and HumanEvalNext revealed average pass-at-1 drops of 31.22% and 19.94%, respectively, underscoring the need for continuous benchmark refinement. We further examine BenchFrame's scalability through an agentic pipeline and confirm its generalizability on the MBPP dataset. All review data, user study materials, and enhanced benchmarks are publicly released.
  •  

Navigating Cancer Complexity: Integrative Multi-Omics Methodologies for Clinical Insights

Clin Med Insights Oncol. 2025 Oct 21;19:11795549251384582. doi: 10.1177/11795549251384582. eCollection 2025.

ABSTRACT

Recent advancements in cancer multi-omics have transformed our understanding of cancer biology by integrating genomics, transcriptomics, proteomics, and metabolomics. These integrative approaches have led to the identification of novel biomarkers and therapeutic targets, offering deeper insights into the molecular intricacies of various cancers, including breast, lung, gastric, pancreatic, and glioblastoma. Despite these advances, challenges remain, such as the integration of disparate data types and the interpretation of complex biological interactions. However, developments in proteogenomics and mass spectrometry have enhanced the correlation between molecular profiles and clinical features, refining the prediction of therapeutic responses. Future research in cancer drug discovery is poised to benefit from multi-omics approaches, improving the precision and efficacy of personalized therapies. By developing integrative network-based models, researchers aim to address challenges related to heterogeneity, reproducibility, and data interpretation. A standardized framework for multi-omics data integration could revolutionize cancer research, optimizing the identification of novel drug targets and enhancing our understanding of cancer biology. This complete approach holds the promise of advancing personalized therapies by fully characterizing the molecular landscape of cancer, ultimately improving patient outcomes through more effective and targeted treatment strategies. This narrative review underscores the potential of multi-omics approaches to transform cancer research and improve patient outcomes through more precise and effective treatments.

PMID:41147019 | PMC:PMC12553891 | DOI:10.1177/11795549251384582

  •  

Robustness is Important: Limitations of LLMs for Data Fitting

arXiv:2508.19563v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are being applied in a wide array of settings, well beyond the typical language-oriented use cases. In particular, LLMs are increasingly used as a plug-and-play method for fitting data and generating predictions. Prior work has shown that LLMs, via in-context learning or supervised fine-tuning, can perform competitively with many tabular supervised learning techniques in terms of predictive performance. However, we identify a critical vulnerability of using LLMs for data fitting -- making changes to data representation that are completely irrelevant to the underlying learning task can drastically alter LLMs' predictions on the same data. For example, simply changing variable names can sway the size of prediction error by as much as 82% in certain settings. Such prediction sensitivity with respect to task-irrelevant variations manifests under both in-context learning and supervised fine-tuning, for both close-weight and open-weight general-purpose LLMs. Moreover, by examining the attention scores of an open-weight LLM, we discover a non-uniform attention pattern: training examples and variable names/values which happen to occupy certain positions in the prompt receive more attention when output tokens are generated, even though different positions are expected to receive roughly the same attention. This partially explains the sensitivity in the presence of task-irrelevant variations. We also consider a state-of-the-art tabular foundation model (TabPFN) trained specifically for data fitting. Despite being explicitly designed to achieve prediction robustness, TabPFN is still not immune to task-irrelevant variations. Overall, despite LLMs' impressive predictive capabilities, currently they lack even the basic level of robustness to be used as a principled data-fitting tool.
  •  

Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms

arXiv:2510.22052v1 Announce Type: new Abstract: The field of artificial intelligence (AI) has taken a tight hold on broad aspects of society, industry, business, and governance in ways that dictate the prosperity and might of the world's economies. The AI market size is projected to grow from 189 billion USD in 2023 to 4.8 trillion USD by 2033. Currently, AI is dominated by large language models that exhibit linguistic and visual intelligence. However, training these models requires a massive amount of data scraped from the web as well as large amounts of energy (50--60 GWh to train GPT-4). Despite these costs, these models often hallucinate, a characteristic that prevents them from being deployed in critical application domains. In contrast, the human brain consumes only 20~W of power. What is needed is the next level of AI evolution in which lightweight domain-specific multimodal models with higher levels of intelligence can reason, plan, and make decisions in dynamic environments with real-time data and prior knowledge, while learning continuously and evolving in ways that enhance future decision-making capability. This will define the next wave of AI, progressing from today's large models, trained with vast amounts of data, to nimble energy-efficient domain-specific agents that can reason and think in a world full of uncertainty. To support such agents, hardware will need to be reimagined to allow energy efficiencies greater than 1000x over the state of the art. Such a vision of future AI systems is developed in this work.
  •  

How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations

arXiv:2510.22780v1 Announce Type: new Abstract: AI agents are continually optimized for tasks related to human work, such as software engineering and professional writing, signaling a pressing trend with significant impacts on the human workforce. However, these agent developments have often not been grounded in a clear understanding of how humans execute work, to reveal what expertise agents possess and the roles they can play in diverse workflows. In this work, we study how agents do human work by presenting the first direct comparison of human and agent workers across multiple essential work-related skills: data analysis, engineering, computation, writing, and design. To better understand and compare heterogeneous computer-use activities of workers, we introduce a scalable toolkit to induce interpretable, structured workflows from either human or agent computer-use activities. Using such induced workflows, we compare how humans and agents perform the same tasks and find that: (1) While agents exhibit promise in their alignment to human workflows, they take an overwhelmingly programmatic approach across all work domains, even for open-ended, visually dependent tasks like design, creating a contrast with the UI-centric methods typically used by humans. (2) Agents produce work of inferior quality, yet often mask their deficiencies via data fabrication and misuse of advanced tools. (3) Nonetheless, agents deliver results 88.3% faster and cost 90.4-96.2% less than humans, highlighting the potential for enabling efficient collaboration by delegating easily programmable tasks to agents.
  •  

Reduced AI Acceptance After the Generative AI Boom: Evidence From a Two-Wave Survey Study

arXiv:2510.23578v1 Announce Type: new Abstract: The rapid adoption of generative artificial intelligence (GenAI) technologies has led many organizations to integrate AI into their products and services, often without considering user preferences. Yet, public attitudes toward AI use, especially in impactful decision-making scenarios, are underexplored. Using a large-scale two-wave survey study (n_wave1=1514, n_wave2=1488) representative of the Swiss population, we examine shifts in public attitudes toward AI before and after the launch of ChatGPT. We find that the GenAI boom is significantly associated with reduced public acceptance of AI (see Figure 1) and increased demand for human oversight in various decision-making contexts. The proportion of respondents finding AI "not acceptable at all" increased from 23% to 30%, while support for human-only decision-making rose from 18% to 26%. These shifts have amplified existing social inequalities in terms of widened educational, linguistic, and gender gaps post-boom. Our findings challenge industry assumptions about public readiness for AI deployment and highlight the critical importance of aligning technological development with evolving public preferences.
  •  
❌