❌

Normal view

UniAI-GraphRAG: Synergizing Ontology-Guided Extraction, Multi-Dimensional Clustering, and Dual-Channel Fusion for Robust Multi-Hop Reasoning

arXiv:2603.25152v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems face significant challenges in complex reasoning, multi-hop queries, and domain-specific QA. While existing GraphRAG frameworks have made progress in structural knowledge organization, they still have limitations in cross-industry adaptability, community report integrity, and retrieval performance. This paper proposes UniAI-GraphRAG, an enhanced framework built upon open-source GraphRAG. The framework introduces three core innovations: (1) Ontology-Guided Knowledge Extraction that uses predefined Schema to guide LLMs in accurately identifying domain-specific entities and relations; (2) Multi-Dimensional Community Clustering Strategy that improves community completeness through alignment completion, attribute-based clustering, and multi-hop relationship clustering; (3) Dual-Channel Graph Retrieval Fusion that balances QA accuracy and performance through hybrid graph and community retrieval. Evaluation results on MultiHopRAG benchmark show that UniAI-GraphRAG outperforms mainstream open source solutions (e.g.LightRAG) in comprehensive F1 scores, particularly in inference and temporal queries. The code is available at https://github.com/UnicomAI/wanwu/tree/main/rag/rag_open_source/rag_core/graph.

Ran Score: a LLM-based Evaluation Score for Radiology Report Generation

arXiv:2603.22935v1 Announce Type: new Abstract: Chest X-ray report generation and automated evaluation are limited by poor recognition of low-prevalence abnormalities and inadequate handling of clinically important language, including negation and ambiguity. We develop a clinician-guided framework combining human expertise and large language models for multi-label finding extraction from free-text chest X-ray reports and use it to define Ran Score, a finding-level metric for report evaluation. Using three non-overlapping MIMIC-CXR-EN cohorts from a public chest X-ray dataset and an independent ChestX-CN validation cohort, we optimize prompts, establish radiologist-derived reference labels and evaluate report generation models. The optimized framework improves the macro-averaged score from 0.753 to 0.956 on the MIMIC-CXR-EN development cohort, exceeds the CheXbert benchmark by 15.7 percentage points on directly comparable labels, and shows robust generalization on the ChestX-CN validation cohort. Here we show that clinician-guided prompt optimization improves agreement with a radiologist-derived reference standard and that Ran Score enables finding-level evaluation of report fidelity, particularly for low-prevalence abnormalities.

STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving

arXiv:2603.22577v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated potential in code generation, yet they struggle with the multi-step, stateful reasoning required for offensive cybersecurity operations. Existing research often relies on static benchmarks that fail to capture the dynamic nature of real-world vulnerabilities. In this work, we introduce STRIATUM-CTF (A Search-based Test-time Reasoning Inference Agent for Tactical Utility Maximization in Cybersecurity), a modular agentic framework built upon the Model Context Protocol (MCP). By standardizing tool interfaces for system introspection, decompilation, and runtime debugging, STRIATUM-CTF enables the agent to maintain a coherent context window across extended exploit trajectories. We validate this approach not merely on synthetic datasets, but in a live competitive environment. Our system participated in a university-hosted Capture-the-Flag (CTF) competition in late 2025, where it operated autonomously to identify and exploit vulnerabilities in real-time. STRIATUM-CTF secured First Place, outperforming 21 human teams and demonstrating strong adaptability in a dynamic problem-solving setting. We analyze the agent's decision-making logs to show how MCP-based tool abstraction significantly reduces hallucination compared to naive prompting strategies. These results suggest that standardized context protocols are a critical path toward robust autonomous cyber-reasoning systems.

Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework

npj Digital Medicine, Published online: 24 March 2026; doi:10.1038/s41746-026-02560-2

Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework

The Effects of Digital Health Interventions on Motor Symptoms, Nonmotor Symptoms, and Quality of Life in Patients With Parkinson Disease: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Parkinson disease (PD) is a progressive neurodegenerative disorder with increasing global prevalence, necessitating innovative management. Digital health interventions (DHIs) offer potential advantages for PD care; yet, a comprehensive systematic review and synthesis across all DHI types and core outcomes is still lacking. Objective: This review aimed to assess the effectiveness of DHIs for improving motor symptoms, nonmotor symptoms, and quality of life in patients with PD and to summarize the reach, uptake, and feasibility. Methods: We searched PubMed, Ovid Embase, Web of Science, CINAHL, Cochrane Central Register of Controlled Trials, and APA PsycINFO up to November 2025. Pooled standardized mean differences (SMDs) were calculated using random-effects models. We calculated 95% prediction intervals (PIs) to estimate the true effects. The revised Cochrane Risk of Bias 2 tool was used to assess risk of bias. Heterogeneity was assessed using I2, τ2, and 95% PI. Subgroup analyses, meta-regression, and sensitivity analyses were conducted to address heterogeneity and potential bias. The quality of evidence was assessed using GRADE (Grading of Recommendations Assessment, Development, and Evaluation). Results: The review included 112 randomized controlled trials involving 5594 participants. Significant postintervention improvements were identified in motor symptoms (SMD=–0.39, 95% CI –0.60 to –0.18, 95% PI –1.75 to 0.99; I2=80.3%) and overall nonmotor symptoms (SMD=–0.26, 95% CI –0.49 to –0.03, 95% PI –0.56 to 0.03; I2=13.8%), including cognitive function (SMD=0.47, 95% CI 0.22 to 0.72, 95% PI –0.41 to 1.35; I2=63.5%) and psychiatric symptoms (SMD=–0.42, 95% CI –0.74 to –0.09, 95% PI –1.82 to –0.99; I2=85.4%); however, there was no significant enhancement in quality of life (SMD=–0.19, 95% CI –0.47 to 0.09, 95% PI –1.50 to 1.12; I2=81.2%). The certainty of evidence was very low for quality of life, motor, and psychiatric symptoms and low for cognitive function and overall nonmotor symptoms. Improvements in motor symptoms and cognitive function remained stable at follow-up. Meta-regression analysis indicated that age, percentage of female participants, and supervision mode were possible sources of heterogeneity. Overall, 94 studies reported reach (median 37.5%), 38 reported fidelity (95.7%), and 105 reported dropout rates (9.1%). Conclusions: In contrast to previous reviews focused on single technologies or outcomes, this review provided the first comprehensive synthesis across all DHI types on multiple outcomes and indicated their potential as nonpharmacological interventions for PD management. However, current evidence is of low to very low certainty, and wide 95% PIs, together with high risk of bias and substantial heterogeneity, indicate considerable uncertainty regarding the true effect in future implementations. Therefore, findings should be interpreted with caution. These findings provide integrated evidence to guide the design and prioritization of future research. The results have important real-world implications, supporting cautious implementation while underscoring the need for more robust trials, particularly in resource-limited settings. Trial Registration: PROSPERO CRD42023492123; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023492123

Efficient Personalized Reranking with Semi-Autoregressive Generation and Online Knowledge Distillation

arXiv:2603.07107v1 Announce Type: cross Abstract: Generative models offer a promising paradigm for the final stage reranking in multi-stage recommender systems, with the ability to capture inter-item dependencies within reranked lists. However, their practical deployment still faces two key challenges: (1) an inherent conflict between achieving high generation quality and ensuring low-latency inference, making it difficult to balance the two, and (2) insufficient interaction between user and item features in existing methods. To address these challenges, we propose a novel Personalized Semi-Autoregressive with online knowledge Distillation (PSAD) framework for reranking. In this framework, the teacher model adopts a semi-autoregressive generator to balance generation quality and efficiency, while its ranking knowledge is distilled online into a lightweight scoring network during joint training, enabling real-time and efficient inference. Furthermore, we propose a User Profile Network (UPN) that injects user intent and models interest dynamics, enabling deeper interactions between users and items. Extensive experiments conducted on three large-scale public datasets demonstrate that PSAD significantly outperforms state-of-the-art baselines in both ranking performance and inference efficiency.

Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO

arXiv:2602.17686v2 Announce Type: replace-cross Abstract: Distilling Chain-of-Thought (CoT) reasoning from large language models into compact student models presents a fundamental challenge: teacher rationales are often too verbose for smaller models to faithfully reproduce. Existing approaches either compress reasoning into single-step, losing the interpretability that makes CoT valuable. We present a three-stage curriculum learning framework that addresses this capacity mismatch through progressive skill acquisition. First, we establish structural understanding via masked shuffled reconstruction. Second, we apply Group Relative Policy Optimization (GRPO) on masked completion tasks, enabling the model to discover its own balance between accuracy and brevity. Third, we identify persistent failure cases and guide the student to internalize teacher knowledge through targeted rewriting, again optimized with GRPO. Experiments on GSM8K demonstrate that our approach enables Qwen2.5-3B-Base to achieve an 11.29 percent accuracy improvement while reducing output length by 27.4 percent, surpassing both instruction-tuned variants and prior distillation methods.

CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment

arXiv:2602.19574v1 Announce Type: cross Abstract: Large-language-model (LLM)-based text-to-speech (TTS) systems can generate natural speech, but most are not designed for low-latency dual-streaming synthesis. High-quality dual-streaming TTS depends on accurate text--speech alignment and well-designed training sequences that balance synthesis quality and latency. Prior work often relies on GMM-HMM based forced-alignment toolkits (e.g., MFA), which are pipeline-heavy and less flexible than neural aligners; fixed-ratio interleaving of text and speech tokens struggles to capture text--speech alignment regularities. We propose CTC-TTS, which replaces MFA with a CTC based aligner and introduces a bi-word based interleaving strategy. Two variants are designed: CTC-TTS-L (token concatenation along the sequence length) for higher quality and CTC-TTS-F (embedding stacking along the feature dimension) for lower latency. Experiments show that CTC-TTS outperforms fixed-ratio interleaving and MFA-based baselines on streaming synthesis and zero-shot tasks. Speech samples are available at https://ctctts.github.io/.
❌