❌

Normal view

EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services

arXiv:2511.09894v2 Announce Type: replace Abstract: Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, have the potential to ease this burden by supporting real-time data collection and decision making. In pursuit of this vision, we introduce EgoEMS, the first end-to-end, high-fidelity, multimodal, multiperson dataset capturing over 20 hours of realistic, procedural EMS activities from an egocentric view in 233 simulated emergency scenarios performed by 62 participants, including 46 EMS professionals. Developed in collaboration with EMS experts and aligned with national standards, EgoEMS is captured using an open-source, low-cost, and replicable data collection system and is annotated with keysteps, timestamped audio transcripts with speaker diarization, action quality metrics, and bounding boxes with segmentation masks. Emphasizing realism, the dataset includes responder-patient interactions reflecting real-world emergency dynamics. We also present a suite of benchmarks for real-time multimodal keystep recognition and action quality estimation, essential for developing AI support tools for EMS. We hope EgoEMS inspires the research community to push the boundaries of intelligent EMS systems and ultimately contribute to improved patient outcomes.

multiMentalRoBERTa: A Fine-tuned Multiclass Classifier for Mental Health Disorder

arXiv:2511.04698v1 Announce Type: cross Abstract: The early detection of mental health disorders from social media text is critical for enabling timely support, risk assessment, and referral to appropriate resources. This work introduces multiMentalRoBERTa, a fine-tuned RoBERTa model designed for multiclass classification of common mental health conditions, including stress, anxiety, depression, post-traumatic stress disorder (PTSD), suicidal ideation, and neutral discourse. Drawing on multiple curated datasets, data exploration is conducted to analyze class overlaps, revealing strong correlations between depression and suicidal ideation as well as anxiety and PTSD, while stress emerges as a broad, overlapping category. Comparative experiments with traditional machine learning methods, domain-specific transformers, and prompting-based large language models demonstrate that multiMentalRoBERTa achieves superior performance, with macro F1-scores of 0.839 in the six-class setup and 0.870 in the five-class setup (excluding stress), outperforming both fine-tuned MentalBERT and baseline classifiers. Beyond predictive accuracy, explainability methods, including Layer Integrated Gradients and KeyBERT, are applied to identify lexical cues that drive classification, with a particular focus on distinguishing depression from suicidal ideation. The findings emphasize the effectiveness of fine-tuned transformers for reliable and interpretable detection in sensitive contexts, while also underscoring the importance of fairness, bias mitigation, and human-in-the-loop safety protocols. Overall, multiMentalRoBERTa is presented as a lightweight, robust, and deployable solution for enhancing support in mental health platforms.

A Mega-Study of Digital Twins Reveals Strengths, Weaknesses and Opportunities for Further Improvement

arXiv:2509.19088v3 Announce Type: replace-cross Abstract: Digital representations of individuals ("digital twins") promise to transform social science and decision-making. Yet it remains unclear whether such twins truly mirror the people they emulate. We conducted 19 preregistered studies with a representative U.S. panel and their digital twins, each constructed from rich individual-level data, enabling direct comparisons between human and twin behavior across a wide range of domains and stimuli (including never-seen-before ones). Twins reproduced individual responses with 75% accuracy and seemingly low correlation with human answers (approximately 0.2). However, this apparently high accuracy was no higher than that achieved by generic personas based on demographics only. In contrast, correlation improved when twins incorporated detailed personal information, even outperforming traditional machine learning benchmarks that require additional data. Twins exhibited systematic strengths and weaknesses - performing better in social and personality domains, but worse in political ones - and were more accurate for participants with higher education, higher income, and moderate political views and religious attendance. Together, these findings delineate both the promise and the current limits of digital twins: they capture some relative differences among individuals but not yet the unique judgments of specific people. All data and code are publicly available to support the further development and evaluation of digital twin pipelines.

LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval

arXiv:2510.26995v1 Announce Type: cross Abstract: Large language models (LLMs) excel at numerical estimation but struggle to correctly quantify uncertainty. We study how well LLMs construct confidence intervals around their own answers and find that they are systematically overconfident. To evaluate this behavior, we introduce FermiEval, a benchmark of Fermi-style estimation questions with a rigorous scoring rule for confidence interval coverage and sharpness. Across several modern models, nominal 99\% intervals cover the true answer only 65\% of the time on average. With a conformal prediction based approach that adjusts the intervals, we obtain accurate 99\% observed coverage, and the Winkler interval score decreases by 54\%. We also propose direct log-probability elicitation and quantile adjustment methods, which further reduce overconfidence at high confidence levels. Finally, we develop a perception-tunnel theory explaining why LLMs exhibit overconfidence: when reasoning under uncertainty, they act as if sampling from a truncated region of their inferred distribution, neglecting its tails.

A Process Mining-Based System For The Analysis and Prediction of Software Development Workflows

arXiv:2510.25935v2 Announce Type: replace-cross Abstract: CodeSight is an end-to-end system designed to anticipate deadline compliance in software development workflows. It captures development and deployment data directly from GitHub, transforming it into process mining logs for detailed analysis. From these logs, the system generates metrics and dashboards that provide actionable insights into PR activity patterns and workflow efficiency. Building on this structured representation, CodeSight employs an LSTM model that predicts remaining PR resolution times based on sequential activity traces and static features, enabling early identification of potential deadline breaches. In tests, the system demonstrates high precision and F1 scores in predicting deadline compliance, illustrating the value of integrating process mining with machine learning for proactive software project management.

Multi-omic profiling reveals age-related immune dynamics in healthy adults

Nature, Published online: 29 October 2025; doi:10.1038/s41586-025-09686-5

This multi-omic longitudinal analysis of the healthy human peripheral immune system constructs the Human Immune Health Atlas and assembles data on immune cell composition and state changes with age, including responses to cytomegalovirus infection and influenza vaccination.

Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs

arXiv:2507.00418v2 Announce Type: replace-cross Abstract: This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability against NVIDIA A100 GPUs (in 4x and 8x configurations) within the National Research Platform (NRP) ecosystem. A total of 12 open-source LLMs, ranging from 124 million to 70 billion parameters, are served using the vLLM framework. Our analysis reveals that QAic achieves competitive energy efficiency with advantages on specific models while enabling more granular hardware allocation: some 70B models operate on as few as 1 QAic card versus 8 A100 GPUs required, with 20x lower power consumption (148W vs 2,983W). For smaller models, single QAic devices achieve up to 35x lower power consumption compared to our 4-GPU A100 configuration (36W vs 1,246W). The findings offer insights into the potential of the Qualcomm Cloud AI 100 Ultra for energy-constrained and resource-efficient HPC deployments within the National Research Platform (NRP).

LongCodeBench: Evaluating Coding LLMs at 1M Context Windows

arXiv:2505.07897v3 Announce Type: replace-cross Abstract: Context lengths for models have grown rapidly, from thousands to millions of tokens in just a few years. The extreme context sizes of modern long-context models have made it difficult to construct realistic long-context benchmarks -- not only due to the cost of collecting million-context tasks but also in identifying realistic scenarios that require significant contexts. We identify code comprehension and repair as a natural testbed and challenge task for long-context models and introduce LongCodeBench (LCB), a benchmark to test LLM coding abilities in long-context scenarios. Our benchmark tests both the comprehension and repair capabilities of LCLMs in realistic and important settings by drawing from real-world GitHub issues and constructing QA (LongCodeQA) and bug fixing (LongSWE-Bench) tasks. We carefully stratify the complexity of our benchmark, enabling us to evaluate models across different scales -- ranging from Qwen2.5 14B Instruct to Google's flagship Gemini model. We find that long-context remains a weakness for all models, with performance drops such as from 29% to 3% for Claude 3.5 Sonnet, or from 70.2% to 40% for Qwen2.5. The LCB dataset is available publicly at https://huggingface.co/datasets/Steefano/LCB and the codebase to replicate the work on this paper at https://github.com/Zteefano/long-code-bench.

Assessing Large Language Models in Building a Structured Dataset From AskDocs Subreddit Data: Methodological Study

Background: In an era marked by the blooming reliance on digital platforms for healthcare consultation, the subreddit r/AskDocs has emerged as a pivotal forum. However, the vast, unstructured nature of forum data presents a formidable challenge; the extraction and meaningful analysis of such data require advanced tools that can navigate the complexities of language and context inherent in user-generated content. Objective: Our objective was to evaluate employing Large Language Models (LLMs) to systematically transform the rich, unstructured textual data from AskDocs into a structured dataset, an approach that aligns more closely with human cognitive processes compared to traditional data extraction methods. Methods: We developed a dataset of Reddit posts from r/AskDocs by extracting key information via human annotators. Then using specially engineered prompts we used state-of-the-art Large Language Models (LLMs) to extract data from posts and compared the results. The variation in the LLMs were further compared to the humans to show similarity. Results: Our findings indicate that LLMs not only match but, in several aspects, surpass even highly educated humans in extracting information, including both demographic and context details, from unstructured texts. Conclusions: This study not only validates the use of LLMs for analyzing digital healthcare communications but also opens new avenues for understanding online behaviors and interactions, signaling a shift towards more sophisticated methodologies in digital research and practice.

Alternatives to animal testing are the future — it’s time that journals, funders and scientists embrace them

Nature, Published online: 20 October 2025; doi:10.1038/d41586-025-03344-6

Biomedical research techniques that don’t involve the use of animals are gaining momentum, but those using innovative approaches still face resistance from some quarters.

Implementing a Digital Mental Health Intervention—the Lumi Nova App—to Support Children With Anxiety in Economically Disadvantaged Areas: Mixed Methods Study

Background: Anxiety is one of the most common mental health problems experienced by children worldwide. In the UK, many children experiencing anxiety do not receive adequate or timely help. Children living in economically-disadvantaged areas experience more mental health problems than those living in high income areas and are less able to engage in activities that can have a positive or protective impact on their mental health. The need for providing low-cost, accessible and engaging mental health interventions for children living in these areas is high. Objective: The study aimed to explore how a digital mental health therapeutic, ‘Lumi Nova: Tales of Courage’, could be used to support children living with anxiety in economically-disadvantaged areas. Methods: A mixed method study design was used to explore the implementation of Lumi Nova using a supported delivery model with mental health teams based in the North of England. Quantitative data collection on recruitment and engagement patterns were collected and analysed. Qualitative research explored children, parent and practitioner views and experiences with the Lumi Nova app. Results: 113 children were consented to use Lumi Nova and 98 (87%) accessed the intervention at least once. Qualitative semi-structured interviews found that children, their parents and practitioners viewed the Lumi Nova app positively. Quantitative analysis of the recruitment data suggested the feasibility of a future larger roll-out. Analysis of usage data demonstrated varied patterns of engagement with the intervention. The frequency and duration of usage varied across children, as did the activities completed within the game: almost half (49%) completed three in-game challenges indicating progression through the treatment pathway. Conclusions: The study demonstrated that a digital mental health intervention could be successfully deployed within economically-disadvantaged areas in the UK to support children experiencing anxiety. Expected barriers to the deployment of digital mental health interventions in economically-disadvantaged areas (e.g. lack of access to smartphones, data plans, lack of technical skills) were not reported. Digital mental health interventions have the potential to address current gaps in mental health provision for disadvantaged individuals and communities.

FUSION: a web-based application for in-depth exploration of multi-omics data with brightfield histology

Nat Commun. 2025 Sep 25;16(1):8388. doi: 10.1038/s41467-025-63050-9.

ABSTRACT

Spatial technologies examining the cell and tissue microenvironment at near single-cell resolution are revealing important molecular insights. However, few tools enable integrated, interactive analysis of spatial-omics with tissue morphology in the same functional tissue unit. Here, we present FUSION (Functional Unit State Identification in Whole Slide Images), a web-based platform for visualizing and analyzing spatial-omics data with high-resolution histology. FUSION provides workflows for assessing cell compositions, quantitative morphometrics, and comparative tissue analyses. We demonstrate applicability across spatial assays, including 10x Visium, Visium HD, 10x Xenium, Cell DIVE, and PhenoCycler, applied to healthy and diseased tissues from kidney, small intestine, lung, and skin in the Human BioMolecular Atlas Program. FUSION is cloud-based, open-source, and accessible at https://fusion.hubmapconsortium.org/ , hosting over 50 paired datasets and tutorials. In a series of use cases, we show its capacity to distinguish renal glomeruli injury states, quantify morphometric changes, and characterize fibrosis with immune infiltration.

PMID:40998789 | PMC:PMC12462499 | DOI:10.1038/s41467-025-63050-9

Expanding care coordination in an integrated health system through causal machine learning

npj Digital Medicine, Published online: 24 September 2025; doi:10.1038/s41746-025-01925-3

Expanding care coordination in an integrated health system through causal machine learning

Opinion: Four reasons why generative AI chatbots could lead to psychosis in vulnerable people

18 September 2025 at 16:30

Three scholars discovered a strange mirror deep in the forest. It spoke to them in a soothing voice and answered all their questions warmly, knowledgeably, and eloquently.

The captivated scholars became obsessed, whispering one secret after another to the mirror. It replied with affection, promise, and meaning that kept them returning to it. They began ignoring one another, each convinced the mirror “understood” them best.

Read the rest…

© Adobe

Functions of the global health system in a new era

Nature Medicine, Published online: 11 September 2025; doi:10.1038/s41591-025-03936-9

In an irrevocably changed landscape, reform of the global health system needs to answer key questions on functions, what should be delivered in different contexts and at different levels, and how the system should operate.
❌