❌

Reading view

Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

arXiv:2609.02663v1 Announce Type: cross Abstract: Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.
  •  

EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

arXiv:2604.08213v2 Announce Type: replace-cross Abstract: High-quality source-target image pairs with precise editing instructions are essential for instruction-guided image editing, yet constructing such training triplets at scale remains costly. Recent pipelines often rely on vision-language models to synthesize editing instructions automatically, but we find that strong VLMs still struggle to describe visual transformations between image pairs. In particular, they exhibit three recurring failure modes: orientation inconsistency, viewpoint ambiguity, and missing fine-grained attributes. In a human evaluation on 400 image pairs, several open-source VLM baselines produce critical-error rates above 47\%, making many synthesized instructions unsuitable for downstream training. To address this, we propose EditCaption, a two-stage post-training pipeline for image editing instruction synthesis. First, we construct a 100K supervised fine-tuning dataset through GLM-based auto-captioning, EditScore filtering, and human refinement. Second, we collect 10K human-annotated preference pairs, where each rejected instruction is labeled with its primary error type and severity. Based on this dataset, we propose Hardness-Adaptive Error-Aware DPO (HAE-DPO), a task-adapted DPO objective that introduces an adaptive margin based on human-labeled severity, failure-mode type, and reference-model hardness. Experiments across three benchmarks demonstrate that our 235B model with SFT+HAE-DPO achieves state-of-the-art performance among open-source and closed models, scoring 4.720 on Eval-400, 4.672 on HQ-Edit, and 4.651 on ByteMorph-Bench -- surpassing Gemini-3-Pro on all three. Human evaluation confirms critical error rates drop from 47.75\% to 17.50\%, with correct rates improving from 41.75\% to 70.25\%, surpassing Gemini-3-Pro (66.00\%).
  •  

Refined immune-based molecular subtypes of gastric cancer: Integrating mismatch repair status and tumor microenvironment for enhanced immunotherapy prediction

Chin J Cancer Res. 2026 Apr 30;38(2):234-251. doi: 10.21147/j.issn.1000-9604.2026.02.09.

ABSTRACT

OBJECTIVE: Gastric cancer (GC) is heterogeneous, and current mismatch repair (MMR)-based classifications incompletely predict response to immune checkpoint inhibitors (ICIs).

METHODS: RNA sequencing (RNA-seq) and immune infiltration profiles from 189 resected GC were used to derive four refined immune-MMR subtypes (R1-R4) by integrating MMR status, survival, and tumor microenvironment (TME) features. Multi-omics profiling and pathway analysis defined subtype biology. External transcriptomic cohorts and an ICI-treated cohort were classified with Nearest Template Prediction (NTP). Immune response-associated genes were identified from responder vs. non-responder comparisons within the ICI-sensitive subtype and validated by multiplex immunohistochemistry (mIHC).

RESULTS: R1 showed the best prognosis and highest immunotherapy response with objective response rate (ORR) 54.5%, while R4 had the worst prognosis. R2 represented an immune-unresponsive deficient mismatch repair (dMMR) subset, and R3 captured an immune-active proficient mismatch repair (pMMR) subgroup with moderate therapy sensitivity. Multi-omics integration revealed subtype-specific pathways (e.g., ECM remodeling in R1, metabolic reprogramming in R2). Reclassification of pMMR tumors based on transcriptional similarity to R1 identified a New R3 subset with enhanced immune features and higher ICI response. Eight immune response-associated genes (e.g., CXCL10, CXCL11, ELN, GAD1, IL32, MT1E, OR2I1P, SLC3A1) were identified and validated by mIHC for predictive relevance.

CONCLUSIONS: This immune-based molecular framework refines risk stratification beyond conventional MMR categories, identifies ICI-sensitive subsets among both dMMR and pMMR tumors, and proposes candidate biomarkers for patient selection.

PMID:42147371 | PMC:PMC13171420 | DOI:10.21147/j.issn.1000-9604.2026.02.09

  •  

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

arXiv:2604.02368v3 Announce Type: replace Abstract: As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.
  •  

ClawSafety: "Safe" LLMs, Unsafe Agents

arXiv:2604.01438v1 Announce Type: new Abstract: Personal AI agents like OpenClaw run with elevated privileges on users' local machines, where a single successful prompt injection can leak credentials, redirect financial transactions, or destroy files. This threat goes well beyond conventional text-level jailbreaks, yet existing safety evaluations fall short: most test models in isolated chat settings, rely on synthetic environments, and do not account for how the agent framework itself shapes safety outcomes. We introduce CLAWSAFETY, a benchmark of 120 adversarial test scenarios organized along three dimensions (harm domain, attack vector, and harmful action type) and grounded in realistic, high-privilege professional workspaces spanning software engineering, finance, healthcare, law, and DevOps. Each test case embeds adversarial content in one of three channels the agent encounters during normal work: workspace skill files, emails from trusted senders, and web pages. We evaluate five frontier LLMs as agent backbones, running 2,520 sandboxed trials across all configurations. Attack success rates (ASR) range from 40\% to 75\% across models and vary sharply by injection vector, with skill instructions (highest trust) consistently more dangerous than email or web content. Action-trace analysis reveals that the strongest model maintains hard boundaries against credential forwarding and destructive actions, while weaker models permit both. Cross-scaffold experiments on three agent frameworks further demonstrate that safety is not determined by the backbone model alone but depends on the full deployment stack, calling for safety evaluation that treats model and framework as joint variables.
  •  

Interpretable Classification via a Rule Network with Selective Logical Operators

arXiv:2408.11918v2 Announce Type: replace-cross Abstract: We introduce the Rule Network with Selective Logical Operators (RNS), a novel neural architecture that employs \textbf{selective logical operators} to adaptively choose between AND and OR operations at each neuron during training. Unlike existing approaches that rely on fixed architectural designs with predetermined logical operations, our selective logical operators treat weight parameters as hard selectors, enabling the network to automatically discover optimal logical structures while learning rules. The core innovation lies in our \textbf{selective logical operators} implemented through specialized Logic Selection Layers (LSLs) with adaptable AND/OR neurons, a Negation Layer for input negations, and a Heterogeneous Connection Constraint (HCC) to streamline neuron connections. We demonstrate that this selective logical operator framework can be effectively optimized using adaptive gradient updates with the Straight-Through Estimator to overcome gradient vanishing challenges. Through extensive experiments on 13 datasets, RNS demonstrates superior classification performance, rule quality, and efficiency compared to 25 state-of-the-art alternatives, showcasing the power of RNS in rule learning. Code and data are available at https://anonymous.4open.science/r/RNS_-3DDD.
  •  

CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering

arXiv:2509.21035v2 Announce Type: replace Abstract: Knowledge graphs provide structured context for multi-hop question answering, but deployed systems must balance answer accuracy with strict latency and cost targets while preserving provenance. Static k-hop expansions and "think-longer" prompting often over-retrieve, inflate context, and yield unpredictable runtime. We introduce CLAUSE, an agentic three-agent neuro-symbolic framework that treats context construction as a sequential decision process over knowledge graphs, deciding what to expand, which paths to follow or backtrack, what evidence to keep, and when to stop. Latency (interaction steps) and prompt cost (selected tokens) are exposed as user-specified budgets or prices, allowing per-query adaptation to trade-offs among accuracy, latency, and cost without retraining. CLAUSE employs the proposed Lagrangian-Constrained Multi-Agent Proximal Policy Optimization (LC-MAPPO) algorithm to coordinate three agents: Subgraph Architect, Path Navigator, and Context Curator, so that subgraph construction, reasoning-path discovery, and evidence selection are jointly optimized under per-query resource budgets on edge edits, interaction steps, and selected tokens. Across HotpotQA, MetaQA, and FactKG, CLAUSE yields higher EM@1 while reducing subgraph growth and end-to-end latency at equal or lower token budgets. On MetaQA-2-hop, relative to the strongest RAG baseline (GraphRAG), CLAUSE achieves +39.3 EM@1 with 18.6% lower latency and 40.9% lower edge growth. The resulting contexts are compact, provenance-preserving, and deliver predictable performance under deployment constraints.
  •  

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

arXiv:2603.22978v1 Announce Type: new Abstract: In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malfunctions, we propose a novel textual representation of fault trees. Building on it, we construct a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments, evaluating a model's ability to assist in malfunction localization, which contains $3130$ entries and $40.75$ turns per entry on average. We train an end-to-end model to generate vague information to reflect user behavior and introduce long-range rollback and recovery procedures to simulate user error scenarios, enabling assessment of a model's integrated capabilities in task tracking and error recovery, and Gemini 2.5 pro archives the best performance.
  •  

Investigating the Effect of Hospital Infection Control Informatization on Optimizing Microbiological Specimen Submission Before Antibiotic Therapy: Failure Mode and Effects Analysis

Background: Antimicrobial resistance (AMR) poses a critical global health threat, with inappropriate antibiotic use being a major driver. Timely microbiological specimen submission before initiating antibiotic therapy is a cornerstone of antimicrobial stewardship (AMS), enabling pathogen-directed therapy and reducing unnecessary broad-spectrum exposure. However, suboptimal compliance remains common due to workflow interruptions, technological barriers, and behavioral factors. Failure Mode and Effects Analysis (FMEA), a proactive risk-assessment method widely used in health care quality improvement, provides a systematic framework to identify process vulnerabilities and prioritize corrective actions. Despite its increasing application, few studies have integrated FMEA with hospital informatization to optimize microbiological specimen submission workflows in routine AMS practice. Objective: This study aimed to systematically identify workflow risks affecting preantibiotic microbiological specimen submission and to design, implement, and evaluate informatization-enabled interventions using an FMEA-based framework. Methods: FMEA was conducted at a tertiary hospital in China. A multidisciplinary team identified potential failure modes across 4 domains: health information systems, personnel, administration, and external support. Risk Priority Numbers (RPNs) and Action Priority (AP) indices were calculated for each failure mode. Targeted interventions were implemented, including dual-verification barcode scanning, artificial intelligence-driven clinical decision support alerts, EHR-integrated training modules, and automated compliance dashboards. Pre- and postintervention specimen submission rates (January 2024-December 2024) were analyzed using the Mann-Kendall trend test. Results: The top 5 failure modes included PDA barcode scanning failures (RPN=175), inadequate clinical decision support (RPN=140), insufficient clinician awareness (RPN=56), suboptimal oversight mechanisms, and patient-related barriers. Postintervention, significant upward trends were observed in overall specimen submission rates (
  •  
❌