❌

Normal view

Empa: An AI-Powered Virtual Mentor for Developing Global Collaboration Skills in HPC Education

arXiv:2511.17669v1 Announce Type: cross Abstract: High-performance computing (HPC) and parallel computing increasingly rely on global collaboration among diverse teams, yet traditional computing curricula inadequately prepare students for cross-cultural teamwork essential in modern computational research environments. This paper presents Empa, an AI-powered virtual mentor that integrates intercultural collaboration training into undergraduate computing education. Built using large language models and deployed through a progressive web application, Empa guides students through structured activities covering cultural dimensions, communication styles, and conflict resolution that are critical for effective multicultural teamwork. Our system addresses the growing need for culturally competent HPC professionals by helping computing students develop skills to collaborate effectively in international research teams, contribute to global computational projects, and navigate the cultural complexities inherent in distributed computing environments. Pilot preparation for deployment in computing courses demonstrates the feasibility of AI-mediated intercultural training and provides insights into scalable approaches for developing intercultural collaboration skills essential for HPC workforce development.

TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG

arXiv:2412.05447v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is one of the leading and most widely used techniques for enhancing LLM retrieval capabilities, but it still faces significant limitations in commercial use cases. RAG primarily relies on the query-chunk text-to-text similarity in the embedding space for retrieval and can fail to capture deeper semantic relationships across chunks, is highly sensitive to chunking strategies, and is prone to hallucinations. To address these challenges, we propose TOBUGraph, a graph-based retrieval framework that first constructs the knowledge graph from unstructured data dynamically and automatically. Using LLMs, TOBUGraph extracts structured knowledge and diverse relationships among data, going beyond RAG's text-to-text similarity. Retrieval is achieved through graph traversal, leveraging the extracted relationships and structures to enhance retrieval accuracy, eliminating the need for chunking configurations while reducing hallucination. We demonstrate TOBUGraph's effectiveness in TOBU, a real-world application in production for personal memory organization and retrieval. Our evaluation using real user data demonstrates that TOBUGraph outperforms multiple RAG implementations in both precision and recall, significantly improving user experience through improved retrieval accuracy.

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

arXiv:2510.21652v1 Announce Type: new Abstract: AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging from general-purpose "deep research" systems to specialized science-specific agents, such as AI Scientist and AIGS. Rigorous evaluation of these agents is critical for progress. Yet existing benchmarks fall short on several fronts: they (1) fail to provide holistic, product-informed measures of real-world use cases such as science research; (2) lack reproducible agent tools necessary for a controlled comparison of core agentic capabilities; (3) do not account for confounding variables such as model cost and tool access; (4) do not provide standardized interfaces for quick agent prototyping and evaluation; and (5) lack comprehensive baseline agents necessary to identify true advances. In response, we define principles and tooling for more rigorously benchmarking agents. Using these, we present AstaBench, a suite that provides the first holistic measure of agentic ability to perform scientific research, comprising 2400+ problems spanning the entire scientific discovery process and multiple scientific domains, and including many problems inspired by actual user requests to deployed Asta agents. Our suite comes with the first scientific research environment with production-grade search tools that enable controlled, reproducible evaluation, better accounting for confounders. Alongside, we provide a comprehensive suite of nine science-optimized classes of Asta agents and numerous baselines. Our extensive evaluation of 57 agents across 22 agent classes reveals several interesting findings, most importantly that despite meaningful progress on certain individual aspects, AI remains far from solving the challenge of science research assistance.

Best Practices for Data Modernization Across the United States Public Health System: Scoping Review

Background: The adoption of new technologies and data modernization approaches in public health aims to enhance the use of health data to inform decision-making and improve population health. However, public health departments struggle with legacy systems, siloed data, and privacy concerns, hampering new technology adoption and data sharing with stakeholders. This paper maps how to address these shortcomings by identifying data modernization challenges, initiatives, and progress. Objective: To characterize the evidence for data modernization associated gaps and best practices in public health. Methods: This scoping review was conducted using the five-stage framework developed by Arksey and O’Malley and was reported according to the PRISMA-ScR guidelines. A structured search was performed in databases PubMed, Scopus, CINAHL, PsycINFO, and was complemented by a further search in the Google Scholar search engine, covering publications from January 1, 2019, to April 30, 2024. Eligible studies were peer-reviewed, published in English, and focused on data modernization initiatives within U.S. public health and reported on best practices, challenges, and outcomes. Search terms combined concepts such as “Data Modernization,” “Interoperability,” and “Public Health” using Boolean operators. Two reviewers independently screened titles, abstracts, and full texts using Rayyan QCRI, with conflicts resolved through consultation with a third reviewer. Data was extracted into Microsoft Excel and thematically analyzed. Results: This review analyzed 22 studies focused on public health data modernization. Across the literature, common components included transitioning to cloud-based systems, consolidating fragmented data into unified platforms, applying governance frameworks, and implementing analytics tools to support decision-making. Primary data sources were electronic health records, insurance claims, and disease surveillance registries. Key challenges identified across studies involved data quality issues, lack of interoperability, and limited resources, particularly in underfunded settings. Notable benefits included more timely and accessible data, improved integration across systems, and enhanced analytical capabilities, which collectively support more responsive and effective public health interventions when guided by clear standards and policy alignment. Conclusions: Progress hinges on balancing local adaptability with national coordination, improving data governance practices, and enhancing collaboration across institutions. These steps are vital to ensure public health systems can deliver timely, accurate, and actionable information to support effective public health efforts.

Digital Health Technology Infrastructure Challenges to Support Health Equity in the United States: Scoping Review

Background: Even though Digital Health Technology (DHT) is widely utilized in the United States (U.S.) at both hospital provider and individual levels, it is beset with several challenges that have contributed to inequities in the health service delivery. Previous studies have shown that health inequities observed may be amplified many by DHT requirements. Objective: The objectives of this scoping review are aimed at synthesizing information on DHT inequities by exploring evidence that describes DHT infrastructure needs focused on promoting health equity in the U.S. and identifying key challenges at both the individual/patient level and at the health service provider's level. Methods: We adapted Arksey and O'Malley's scoping review guidelines in our review. We searched PubMed, Web of Science, CINAHL, and PsycINFO were searched. We also conducted supplementary searches on Google Scholar. The inclusion criteria were peer-reviewed publications that broadly conceptualize or analyze DHT infrastructure from a health equity perspective and the challenges of DHT requirements between 2020 and 2024. Following a full-text screening using eligibility criteria such as studies were included if they examined DHT infrastructure in the U.S. from a health equity perspective, discussed health disparities resulting from DHT interventions, or investigated the variables influencing health inequities connected to DHT. Two researchers evaluated each citation’s individually at the title and abstract levels. Thematic approach and qualitative analysis determined this scoping review’s outcome. Results: Of the 628 research articles from the search, 27 were included in the analysis based on the inclusion criteria. In this review, we discussed factors such as elderly population, education, race, ethnicity, and socioeconomic status leading to health inequities in DHT. Patients and Service providers challenges that exist in health inequities related to DHT. The most common challenges for service providers were infrastructure and technical issues such as inadequate integration with existing workflows, user-unfriendly health information exchange (HIE) interfaces, and lack of skilled staff, while for individuals or patients, this included limited broadband internet access, cultural or linguistic appropriateness, and access to digital tools. Conclusions: The study identified that in the U.S., DHT is an essential part of the delivery of health services, yet it is saddled with key challenges leading to health inequities. Finding pragmatic solutions to these challenges can improve health equity in DHT.

Reconstruction of a fatty acid synthesis cycle from acyl carrier protein and cofactor structural snapshots

Structural snapshots of cofactors in the yeast fatty acid synthase at 1.9 Å resolution and ACP structural snapshots representing intermediates in the fatty acid biosynthesis cycle open avenues for diversifying the product profile and discovering anti-fungal fatty acid synthase inhibitors.
❌