❌

Normal view

Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems

arXiv:2603.28998v1 Announce Type: cross Abstract: As Large Language Models (LLMs) and multi-agent AI systems are demonstrating increasing potential in cybersecurity operations, organizations, policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabilities of such AI systems to achieve more autonomous SOCs (security operation centers) and reduce manual effort. In particular, the AI and cybersecurity communities have recently developed several benchmarks for evaluating the red team capabilities of multi-agent AI systems. However, because the operations in SOCs are dominated by blue team operations, the capabilities of AI systems & agents to achieve more autonomous SOCs cannot be evaluated without a benchmark focused on blue team operations. To our best knowledge, no systematic benchmark for evaluating coordinated multi-task blue team AI has been proposed in the literature. Existing blue team benchmarks focus on a particular task. The goal of this work is to develop a set of design principles for the construction of a benchmark, which is denoted as SOC-bench, to evaluate the blue team capabilities of AI. Following these design principles, we have developed a conceptual design of SOC-bench, which consists of a family of five blue team tasks in the context of large-scale ransomware attack incident response.

Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

arXiv:2603.25158v3 Announce Type: replace Abstract: Equipping Large Language Model (LLM) agents with domain-specific skills is critical for tackling complex tasks. Yet, manual authoring creates a severe scalability bottleneck. Conversely, automated skill generation often yields fragile or fragmented results because it either relies on shallow parametric knowledge or sequentially overfits to non-generalizable trajectory-local lessons. To overcome this, we introduce Trace2Skill, a framework that mirrors how human experts author skills: by holistically analyzing broad execution experience before distilling it into a single, comprehensive guide. Instead of reacting sequentially to individual trajectories, Trace2Skill dispatches a parallel fleet of sub-agents to analyze a diverse pool of executions. It extracts trajectory-specific lessons and hierarchically consolidates them into a unified, conflict-free skill directory via inductive reasoning. Trace2Skill supports both deepening existing human-written skills and creating new ones from scratch. Experiments in challenging domains, such as spreadsheet, VisionQA and math reasoning, show that Trace2Skill significantly improves upon strong baselines, including Anthropic's official xlsx skills. Crucially, this trajectory-grounded evolution does not merely memorize task instances or model-specific quirks: evolved skills transfer across LLM scales and generalize to OOD settings. For example, skills evolved by Qwen3.5-35B on its own trajectories improved a Qwen3.5-122B agent by up to 57.65 absolute percentage points on WikiTableQuestions. Ultimately, our results demonstrate that complex agent experience can be packaged into highly transferable, declarative skills -- requiring no parameter updates, no external retrieval modules, and utilizing open-source models as small as 35B parameters.

Two-step clinical care pathway to predict MASLD-related advanced fibrosis and long-term outcomes in type 2 diabetes

Gut. 2026 Feb 9;75(3):576-587. doi: 10.1136/gutjnl-2025-337506.

ABSTRACT

BACKGROUND: Current guidelines recommend a two-step approach for risk stratification of metabolic dysfunction-associated steatotic liver disease (MASLD), starting with Fibrosis-4 index (FIB-4) followed by liver stiffness measurement (LSM) using vibration-controlled transient elastography (VCTE).

OBJECTIVE: To evaluate this approach for predicting advanced fibrosis and liver-related events (LREs) in patients with type 2 diabetes (T2D).

DESIGN: A prospective liver biopsy cohort of T2D patients with histologically confirmed MASLD from seven centres in China was used to assess diagnostic performance for advanced fibrosis. The international VCTE-Prognosis cohort, including T2D patients with MASLD who underwent VCTE at 16 centres in the USA, Europe and Asia, with longitudinal follow-up, was used to assess LREs, defined as hepatic decompensation or hepatocellular carcinoma.

RESULTS: 4781 participants were included. In the liver biopsy cohort (n=352; 22.2% with advanced fibrosis), applying LSM thresholds of <8 kPa and >12 kPa after FIB-4 classified patients into 63.4% low-risk, 9.4% intermediate-risk and 27.3% high-risk, with a correct classification rate of 71%. In the VCTE-Prognosis cohort (n=4429; median follow-up 51.3 (IQR 27.4-70.7) months), 140 (3.2%) patients developed LREs (110 (2.5%) with hepatic decompensation and 59 (1.3%) with hepatocellular carcinoma). The two-step approach classified 72.6%, 6.8% and 20.6% of patients into low-risk, intermediate-risk and high-risk groups, with corresponding 5-year cumulative LRE incidences of 0.7%, 0.9% and 11.8%. Refining classification of intermediate FIB-4 patients using LSM <10 kPa (low-risk) and >15 kPa (high-risk) reduced the intermediate-risk group to 5.6% while preserving predictive accuracy.

CONCLUSION: The non-invasive two-step approach of FIB-4 followed by LSM effectively stratifies MASLD-related advanced fibrosis and LREs risk in T2D. Applying LSM cut-offs of 10 and 15 kPa further optimises risk stratification for future LREs.

PMID:41911049 | DOI:10.1136/gutjnl-2025-337506

❌