❌

Normal view

Telehealth Delivery of the Homeostasis–Enrichment–Plasticity Approach for Premature Infants With Developmental Risks: Exploratory Feasibility Study

Background: Preterm delivery is an increasing worldwide health concern linked to increased neurodevelopmental risks. Early intervention is crucial for harnessing neuroplasticity to enhance developmental and functional performance outcomes; however, access to early intervention is frequently hindered by logistical, financial, and labor constraints. The Homeostasis–Enrichment–Plasticity (HEP) Approach is a family-centered early intervention model based on enriched environments, designed to improve infants’ sensory-motor, cognitive, and socio-emotional development. Objective: This study aimed to assess the feasibility, safety, acceptability, and outcomes sensitivity to change of implementing the HEP Approach through telehealth for premature infants at developmental risk. Methods: A pre-post exploratory feasibility study was performed, including 16 preterm infants (aged 4-12 months corrected age), of whom 14 completed the study. The 12-week intervention included weekly remote sessions focused on environmental enrichment, active exploration, and parental guidance. The feasibility and acceptability were evaluated using a 24-item questionnaire. Developmental outcomes were assessed with the Young Children’s Participation and Environment Measure, Ages and Stages Questionnaire (ASQ), Alberta Infant Motor Scale, Infant Motor Profile, and Depression Anxiety Stress Scales. Results: High adherence (14/14, 100%) and retention (14/16, 87.5%) rates demonstrated robust feasibility. Parents indicated 86%-100% agreement across all feasible criteria, affirming safety, satisfaction, and acceptability. No adverse incidents were reported. Changes were identified in participation (Young Children’s Participation and Environment Measure), motor development (Alberta Infant Motor Scale, Infant Motor Profile, and ASQ), communication and social-emotional domains (ASQ), and caregiver well-being (Depression Anxiety Stress Scales) (P<.05). Conclusions: The telehealth implementation of the HEP Approach demonstrated feasibility, safety, and strong acceptance among families, along with quantifiable developmental and psychosocial changes. These initial findings endorse the model’s viability as an accessible, family-oriented telehealth framework for infants born preterm. Future randomized controlled and longitudinal studies are necessary to validate intervention efficacy and scalability.

Preferences for Personalized Text Message Appointment Reminders Among Outpatients in a Universal Health System: Cross-Sectional Study

Background: SMS text messaging reminders are widely used to reduce missed outpatient appointments; however, evidence remains limited regarding which types of reminder content patients prefer, particularly within East Asian universal health systems. In Taiwan, minimal financial barriers to care and unrestricted access to secondary and tertiary hospitals contribute to high outpatient visit volumes and persistent no-show rates. These contextual features underscore the need for behaviorally informed and demographically tailored reminder strategies rather than uniform messaging approaches. Objective: This study aimed to examine patient preferences for 6 theory-guided SMS appointment reminder types and to identify the predictors of reminder preference related to demographic characteristics and health care utilization, with the goal of informing personalized reminder design for a forthcoming randomized controlled trial. Methods: We conducted a cross-sectional online survey among adults in Taiwan with prior outpatient experience. Six SMS reminder prototypes were developed based on behavioral communication principles and validated by a multidisciplinary expert panel using item-level content validity indices. Participants selected their preferred SMS reminder type and reported sociodemographic characteristics and recent health care utilization. Bivariate associations were examined using chi-square tests and one-way ANOVA, with Benjamini-Hochberg false discovery rate correction applied to control for multiple testing. To identify independent predictors of SMS reminder preference while adjusting for potential confounding, we fitted a multinomial logistic regression model with all covariates entered simultaneously. Results: A total of 1095 respondents completed the survey. General reminders and messages referencing prior missed appointments were most frequently preferred, whereas empathy-based or relationally framed messages were selected less often. In false discovery rate–adjusted univariate analyses, both age and sex were associated with SMS reminder preference. However, in the fully adjusted multinomial logistic regression model, age emerged as the only statistically significant independent predictor. Participants younger than 50 years were significantly more likely to prefer alternative reminder message types compared with the general reminder (adjusted odds ratio 1.64, 95% CI 1.18‐2.28; =.003). Sex did not retain statistical significance after multivariable adjustment. Other sociodemographic characteristics and health care utilization variables, including education level, employment status, residential region, outpatient visit frequency, and recent missed appointments history, were not independently associated with reminder preference. Conclusions: Preferences for outpatient SMS reminder content vary systematically, with age representing the most robust independent predictor. Across the sample, concise and behavior-focused reminders were preferred over empathy-oriented or relational formats. These findings support age-informed tailoring of SMS reminder content and provide content-validated SMS prototypes for use in subsequent interventional research. The results offer formative evidence to guide the design of randomized trials aimed at reducing outpatient no-shows and improving the efficiency of ambulatory care delivery in Taiwan’s universal health care system.

A Gamified Mobile Health Intervention to Promote Physical Activity, Executive Function, and Mental Health in College Students: Randomized Controlled Trial

Background: College students commonly experience suboptimal health conditions, including insufficient physical activity (PA), excessive body weight, and declining physical fitness. Traditional interventions face low adherence, while gamified mobile health (mHealth) programs may improve engagement and outcomes. Objective: This study aimed to evaluate the feasibility and effectiveness of a novel gamified, incentive-based mHealth intervention on primary outcomes (PA and adherence) and secondary outcomes (physical fitness, body composition, executive function [EF], and mental health). Methods: A 2-arm parallel-group randomized controlled trial (RCT) was conducted in 2025 at Yantai University with 160 college students (18‐25 years; BMI 18.5‐30.0) who were randomized 1:1 (computer-generated, sex-stratified blocks of 4; concealed allocation) to the intervention group (IG) or control group (CG; n=80 each); major exclusions were contraindications to exercise, severe physical/mental illness, recent PA interventions, or psychotropic medication use. Both used the same fitness watch–app system and identical PA targets (≥150 min moderate-to-vigorous physical activity [MVPA] per week or ≥900 metabolic equivalent-minutes [MET-min] per week); IG additionally received team-based gamification (competition, points/leaderboards, feedback, and rewards), while CG received monitoring only. PA and adherence were monitored throughout the 8-week intervention; other outcomes were assessed at baseline and 8 weeks (fitness, body composition, EF, and mental health). Open-label with blinded outcome assessors/analysts; intention-to-treat (ITT) with multiple imputation. Results: At 8 weeks, data were available for 154 participants (IG 78; CG 76); all 160 were analyzed per ITT. Compared to the CG, the IG demonstrated significantly higher mean levels in all primary PA outcomes over 8 weeks (daily steps: mean 10,356, SD 1245 versus 8242, SD 1087; Δ=2114; =1.81, 95% CI 1.44‐2.18;

Initial Insights Into an Institutional Secure Large Language Model for Magnetic Resonance Imaging Examination Requests: Retrospective Study

Background: Incomplete clinical details on magnetic resonance imaging (MRI) examination requests (MERs) can lead to suboptimal protocol selection. An institutional secure large language model (sLLM) with access to manually retrieved salient data from the electronic medical record (EMR) may improve request completeness and protocol accuracy across multiple MRI subspecialties. Objective: The objective of this study was to compare clinician MERs with sLLM-augmented MERs for information quality and to evaluate the protocoling accuracy of the sLLM versus board-certified radiologists across body, musculoskeletal, and neuroradiology MRI. Methods: This retrospective study included 608 random outpatient MRI examinations performed between September 2023 and July 2024 (body 206, musculoskeletal 203, neuroradiology 199). The cohort comprised 528 patients (mean 51.2 years, SD 19.2; range 4‐93; n=279, 52.8% women, n=249, 47.2% men). MERs without EMR access were excluded. A privately hosted Anthropic Claude 3.5 model (temperature 0) augmented each MER with manually retrieved salient EMR data and, via rule-based parsing, mapped the extracted elements onto predefined institutional criteria to recommend region or coverage and contrast use. Two experienced radiologists established a consensus reference standard. Two board-certified general radiologists (Rad 3 and Rad 4) and the sLLM were compared with this standard. Clinical information quality was graded using the Reason-for-Exam Imaging Reporting and Data System (RI-RADS). Interrater reliability was quantified with Gwet AC1. Paired accuracies were compared with the McNemar test to determine whether there was a statistically significant difference. Results: Interreader agreement for RI-RADS was almost perfect for sLLM-augmented MERs (AC1 0.97, 95% CI 0.94‐0.99) and moderate for clinician MERs (AC1 0.43, 95% CI 0.34‐0.52). Limited or deficient clinical information (RI-RADS C/D) fell to 0% to 0.7% (0/608 to 4/608) with sLLM augmentation vs 4.1% to 20.4% (25/608 to 124/608) for clinician MERs. Overall protocol accuracy was 93.1% (566/608; 95% CI 89.6‐96.6) for the sLLM, 91.4% (556/608; 95% CI 87.6‐95.3) for Rad 3, and 92.1% (560/608; 95% CI 88.4‐95.8) for Rad 4 (sLLM vs Rad 3 =.23 vs Rad 4 =.40). Region or coverage accuracy was similar (sLLM: 579/608, 95.2%; Rad 3: 585/608, 96.2%; Rad 4: 573/608, 94.2%; =.46 and =.36). Contrast decisions were more accurate using the sLLM at 94.4% (574/608; 95% CI 91.3‐97.5) vs Rad 3 at 92.1% (560/608; 95% CI 88.4‐95.8; =.027) and were not significantly different to Rad 4 at 92.9% (565/608; 95% CI 89.4‐96.4; =.16). Subspecialty analyses showed similar patterns, with the sLLM outperforming Rad 4 for musculoskeletal MRI contrast decisions (96.6% vs 91.1%; =.006) and matching readers elsewhere. Manual review indicated that sLLM improvements arose from EMR details not listed on the MER (infection/inflammation, tumor history, prior surgery). No clinically significant hallucinations were identified in a manual review of discordant cases. Conclusions: Across body, musculoskeletal, and neuroradiology MRI, sLLM-augmented examination requests improved clinical context and enhanced contrast selection while demonstrating accuracy comparable to general radiologists for region or coverage. Integrating sLLMs into routine vetting workflows may reduce manual workload in protocol selection for more efficient, standardized protocoling.

Social Media Intervention Based on the Information-Motivation-Behavioral Skills Model Promotes HIV Testing and Reduces High-Risk Behaviors Among Men Who Have Sex With Men in Resource-Limited Settings in China: Randomized Controlled Trial

Background: Social media intervention may enhance HIV prevention among men who have sex with men, but the effect of this intervention in resource-limited settings remains unclear. Objective: This randomized controlled trial evaluated whether a social media intervention grounded in the information-motivation-behavioral skills (IMB) model could be beneficial for HIV prevention among men who have sex with men in resource-limited settings. Methods: Participants were recruited in Nanning, China, between April 2023 and April 2024. Eligible participants were randomly assigned to either the social media intervention group or the routine HIV prevention services control group. Participants in the intervention group received a 3-month social media intervention, which included completing video-based tasks. Baseline surveys were conducted, followed by follow-up surveys every 3 months, for a total of 2 follow-ups. Outcomes included HIV testing uptake, high-risk behavior, AIDS-related knowledge, safe sex self-efficacy, and attitude. Results: A total of 180 eligible men who have sex with men were enrolled (90 per group). Follow-up rates were 97.8% (88/90) and 95.5% (86/90) for the intervention and control groups, respectively. At the follow-ups, the intervention group demonstrated significantly higher uptake of HIV testing, a lower proportion of participants reporting high-risk sexual behaviors, and higher condom use self-efficacy compared to the control group (all

WNT7A correlates with immunosuppression and predicts adverse prognosis in lung adenocarcinoma: Potential implication of the NF-kappaB/CCL2 Axis

Cytokine. 2026 Apr 6;202:157144. doi: 10.1016/j.cyto.2026.157144. Online ahead of print.

ABSTRACT

BACKGROUND: The remodeling of the tumor immune microenvironment (TME) is a pivotal determinant of therapeutic efficacy and clinical outcome in Lung Adenocarcinoma (LUAD). While WNT signaling is a known oncogenic driver, the specific immunomodulatory role of WNT7A and its potential crosstalk with inflammatory pathways in LUAD remain to be fully elucidated. We sought to define the prognostic value of WNT7A and explore the molecular mechanisms by which it may foster an immunosuppressive TME.

METHODS: We performed a multi-omics analysis utilizing the TCGA-LUAD cohort (N = 508) and validated findings in an independent external cohort (GSE30219, N = 293). The prognostic significance of WNT7A was evaluated using Kaplan-Meier and multivariate Cox regression analyses. TME composition was dissected via ssGSEA, focusing on myeloid-derived suppressor cell (MDSC) infiltration. Mechanistic pathways were identified using Gene Set Enrichment Analysis (GSEA) and gene co-expression networks.

RESULTS: High WNT7A expression was identified as a significant predictor of poor Overall Survival (OS) in the TCGA cohort (P < 0.05) and validated in the external cohort (P < 0.05). Multivariate analysis confirmed WNT7A as an independent prognostic risk factor (HR = 1.085, P = 0.036). Immunologically, WNT7A expression was positively correlated with MDSC infiltration (R = 0.43, P < 0.001), suggesting a shift towards an immune-tolerant phenotype. Mechanistically, GSEA revealed a robust activation of inflammatory signaling in the high-WNT7A group. Specifically, the TNFA Signaling via NF-κB pathway was significantly enriched(NES = 2.52, P < 0.001). Consistent with this pathway activation, WNT7A showed a statistically significant positive correlation with CCL2 (P < 0.001), a critical chemokine for MDSC recruitment, implicating the NF-κB/CCL2 axis in this process.

CONCLUSION: WNT7A serves as a prognostic biomarker linked to immune evasion in LUAD, potentially by modulating the NF-κB/CCL2/MDSC axis. This study identifies WNT7A as a potential therapeutic target to remodel the immune microenvironment, providing a rationale for future investigations into WNT-targeted strategies to improve immunotherapy efficacy.

PMID:41946008 | DOI:10.1016/j.cyto.2026.157144

Iron Physiology and Its Impact on Atopic Diseases: An EAACI Taskforce Report

Allergy. 2026 Apr 6. doi: 10.1111/all.70325. Online ahead of print.

ABSTRACT

Iron is essential for oxygen transport, energy metabolism, and immune regulation. Yet iron deficiency is the most common micronutrient disorder across all age groups, affecting nearly one quarter of the global population. Iron deficiency triggers nutritional immunity, a host defense mechanism that withholds and redistributes iron, contributing to increased morbidity and mortality. This review outlines normal iron physiology, distribution and absorption pathways and on the consequences of deficiency across body compartments, with particular attention to type 2-driven diseases. Beyond anemia, insufficient iron availability disrupts immune homeostasis by promoting type 2 inflammation, elevating IgE, and activating mast cells and eosinophils. Regulatory macrophages, the central hub of iron cycling, adopt an inflammatory, iron-sequestering state that reinforces malabsorption and redistribution. Epidemiology studies show higher iron-deficiency risk in allergic individuals; low maternal iron or early-life iron predisposes to eczema, wheeze, and asthma, while food-allergen elimination (notably cow's milk) further worsens anemia risk. Clinical evidence indicates that restoring iron status through diet, supplementation, or fortification lowers IgE levels, improves lung function, and alleviates symptoms of rhinitis, urticaria, and asthma. Iron may therefore represent a modifiable determinant of allergic disease development and severity. Integrating iron assessment and nutritional care into allergy management may reduce disease burden and slow the progression of allergic march.

PMID:41943501 | DOI:10.1111/all.70325

WNT7A correlates with immunosuppression and predicts adverse prognosis in lung adenocarcinoma: Potential implication of the NF-kappaB/CCL2 Axis

Cytokine. 2026 Apr 6;202:157144. doi: 10.1016/j.cyto.2026.157144. Online ahead of print.

ABSTRACT

BACKGROUND: The remodeling of the tumor immune microenvironment (TME) is a pivotal determinant of therapeutic efficacy and clinical outcome in Lung Adenocarcinoma (LUAD). While WNT signaling is a known oncogenic driver, the specific immunomodulatory role of WNT7A and its potential crosstalk with inflammatory pathways in LUAD remain to be fully elucidated. We sought to define the prognostic value of WNT7A and explore the molecular mechanisms by which it may foster an immunosuppressive TME.

METHODS: We performed a multi-omics analysis utilizing the TCGA-LUAD cohort (N = 508) and validated findings in an independent external cohort (GSE30219, N = 293). The prognostic significance of WNT7A was evaluated using Kaplan-Meier and multivariate Cox regression analyses. TME composition was dissected via ssGSEA, focusing on myeloid-derived suppressor cell (MDSC) infiltration. Mechanistic pathways were identified using Gene Set Enrichment Analysis (GSEA) and gene co-expression networks.

RESULTS: High WNT7A expression was identified as a significant predictor of poor Overall Survival (OS) in the TCGA cohort (P < 0.05) and validated in the external cohort (P < 0.05). Multivariate analysis confirmed WNT7A as an independent prognostic risk factor (HR = 1.085, P = 0.036). Immunologically, WNT7A expression was positively correlated with MDSC infiltration (R = 0.43, P < 0.001), suggesting a shift towards an immune-tolerant phenotype. Mechanistically, GSEA revealed a robust activation of inflammatory signaling in the high-WNT7A group. Specifically, the TNFA Signaling via NF-κB pathway was significantly enriched(NES = 2.52, P < 0.001). Consistent with this pathway activation, WNT7A showed a statistically significant positive correlation with CCL2 (P < 0.001), a critical chemokine for MDSC recruitment, implicating the NF-κB/CCL2 axis in this process.

CONCLUSION: WNT7A serves as a prognostic biomarker linked to immune evasion in LUAD, potentially by modulating the NF-κB/CCL2/MDSC axis. This study identifies WNT7A as a potential therapeutic target to remodel the immune microenvironment, providing a rationale for future investigations into WNT-targeted strategies to improve immunotherapy efficacy.

PMID:41946008 | DOI:10.1016/j.cyto.2026.157144

  • ✇STAT
  • STAT+: FDA backs proposals to entice pharma companies to test, make drugs domestically John Wilkerson and Lizzy Lawrence
    WASHINGTON — The Food and Drug Administration used the president’s budget to propose policies aimed at encouraging domestic development and manufacturing of drugs.   FDA Commissioner Marty Makary has said the agency needs “giant, big ideas” to counter China’s dominance in early-stage clinical development of drugs. Among the FDA’s ideas are proposals to make it easier to run early-stage trials in the U.S. and to hand an advantage to U.S.-based generics manufacturers. The Trump administrat
     

STAT+: FDA backs proposals to entice pharma companies to test, make drugs domestically

7 April 2026 at 16:30

WASHINGTON — The Food and Drug Administration used the president’s budget to propose policies aimed at encouraging domestic development and manufacturing of drugs.  

FDA Commissioner Marty Makary has said the agency needs “giant, big ideas” to counter China’s dominance in early-stage clinical development of drugs. Among the FDA’s ideas are proposals to make it easier to run early-stage trials in the U.S. and to hand an advantage to U.S.-based generics manufacturers.

The Trump administration has been using a variety of policy levers to try and bring drug manufacturing to the U.S. For example, many of the brand drugmakers that struck deals to lower U.S. prices also promised to increase domestic manufacturing, under the threat of tariffs.

Continue to STAT+ to read the full story…

© DREW ANGERER/AFP via Getty Images

Evaluating Artificial Intelligence Through a Christian Understanding of Human Flourishing

arXiv:2604.03356v1 Announce Type: new Abstract: Artificial intelligence (AI) alignment is fundamentally a formation problem, not only a safety problem. As Large Language Models (LLMs) increasingly mediate moral deliberation and spiritual inquiry, they do more than provide information; they function as instruments of digital catechesis, actively shaping and ordering human understanding, decision-making, and moral reflection. To make this formative influence visible and measurable, we introduce the Flourishing AI Benchmark: Christian Single-Turn (FAI-C-ST), a framework designed to evaluate Frontier Model responses against a Christian understanding of human flourishing across seven dimensions. By comparing 20 Frontier Models against both pluralistic and Christian-specific criteria, we show that current AI systems are not worldview-neutral. Instead, they default to a Procedural Secularism that lacks the grounding necessary to sustain theological coherence, resulting in a systematic performance decline of approximately 17 points across all dimensions of flourishing. Most critically, there is a 31-point decline in the Faith and Spirituality dimension. These findings suggest that the performance gap in values alignment is not a technical limitation, but arises from training objectives that prioritize broad acceptability and safety over deep, internally coherent moral or theological reasoning.

VERT: Reliable LLM Judges for Radiology Report Evaluation

arXiv:2604.03376v1 Announce Type: new Abstract: Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when applied to reports from other modalities and anatomies. Which model and prompt configurations are best suited to serve as LLM judges for radiology evaluation? We conduct a thorough correlation analysis between expert and LLM-based ratings. We compare three existing LLM-as-a-judge metrics (RadFact, GREEN, and FineRadScore) alongside VERT, our proposed LLM-based metric, using open- and closed-source models (reasoning and non-reasoning) of different sizes across two expert-annotated datasets, RadEval and RaTE-Eval, spanning multiple modalities and anatomies. We further evaluate few-shot approaches, ensembling, and parameter-efficient fine-tuning using RaTE-Eval. To better understand metric behavior, we perform a systematic error detection and categorization study to assess alignment of these metrics against expert judgments and identify areas of lower and higher agreement. Our results show that VERT improves correlation with radiologist judgments by up to 11.7% relative to GREEN. Furthermore, fine-tuning Qwen3 30B yield gains of up to 25% using only 1,300 training samples. The fine-tuned model also reduces inference time up to 37.2 times. These findings highlight the effectiveness of LLM-based judges and demonstrate that reliable evaluation can be achieved with lightweight adaptation.

Hume's Representational Conditions for Causal Judgment: What Bayesian Formalization Abstracted Away

arXiv:2604.03387v1 Announce Type: new Abstract: Hume's account of causal judgment presupposes three representational conditions: experiential grounding (ideas must trace to impressions), structured retrieval (association must operate through organized networks exceeding pairwise connection), and vivacity transfer (inference must produce felt conviction, not merely updated probability). This paper extracts these conditions from Hume's texts and argues that they are integral to his causal psychology. It then traces their fate through the formalization trajectory from Hume to Bayesian epistemology and predictive processing, showing that later frameworks preserve the updating structure of Hume's insight while abstracting away these further representational conditions. Large language models serve as an illustrative contemporary case: they exhibit a form of statistical updating without satisfying the three conditions, thereby making visible requirements that were previously background assumptions in Hume's framework.

TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering

arXiv:2604.03393v1 Announce Type: new Abstract: Multimodal reasoning has emerged as a powerful framework for enhancing reasoning capabilities of reasoning models. While multi-turn table reasoning methods have improved reasoning accuracy through tool use and reward modeling, they rely on fixed text serialization for table state readouts. This introduces representation errors in table encoding that significantly accumulate over multiple turns. Such accumulation is alleviated by tabular grounding methods in the expense of inference compute and cost, rendering real world deployment impractical. To address this, we introduce TABQAWORLD, a table reasoning framework that jointly optimizes tabular action through representation and estimation. For representation, TABQAWORLD employs an action-conditioned multimodal selection policy, which dynamically switches between visual and textual representations to maximize table state readout reliability. For estimation, TABQAWORLD optimizes stepwise reasoning trajectory through table metadata including dimension, data types and key values, safely planning trajectory and compressing low-complexity actions to reduce conversation turns and latency. Designed as a training-free framework, empirical evaluations show that TABQAWORLD achieves state-of-the-art performance with 4.87% accuracy improvements over baselines, with 5.42% accuracy gain and 33.35% inference latency reduction over static settings, establishing a new standard for reliable and efficient table reasoning.

ActionNex: A Virtual Outage Manager for Cloud

arXiv:2604.03512v1 Announce Type: new Abstract: Outage management in large-scale cloud operations remains heavily manual, requiring rapid triage, cross-team coordination, and experience-driven decisions under partial observability. We present \textbf{ActionNex}, a production-grade agentic system that supports end-to-end outage assistance, including real-time updates, knowledge distillation, and role- and stage-conditioned next-best action recommendations. ActionNex ingests multimodal operational signals (e.g., outage content, telemetry, and human communications) and compresses them into critical events that represent meaningful state transitions. It couples this perception layer with a hierarchical memory subsystem: long-term Key-Condition-Action (KCA) knowledge distilled from playbooks and historical executions, episodic memory of prior outages, and working memory of the live context. A reasoning agent aligns current critical events to preconditions, retrieves relevant memories, and generates actionable recommendations; executed human actions serve as an implicit feedback signal to enable continual self-evolution in a human-agent hybrid system. We evaluate ActionNex on eight real Azure outages (8M tokens, 4,000 critical events) using two complementary ground-truth action sets, achieving 71.4\% precision and 52.8-54.8\% recall. The system has been piloted in production and has received positive early feedback.
❌