❌

Normal view

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study

Background: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. Objective: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. Methods: The benchmark comprises 3139 objective items from educational resources, 3175 research article–derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article–derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. Results: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07‐91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68‐70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95‐69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial ≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator’s per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%‐34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). Conclusions: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.

A Non-Canonical Role of SMAD4 in Regulating 3D Genome Architecture to Inhibit Lung Squamous Cell Carcinoma Development

Adv Sci (Weinh). 2026 May 26:e75839. doi: 10.1002/advs.75839. Online ahead of print.

ABSTRACT

Lung squamous cell carcinoma (LUSC) lacks clearly defined key drivers and effective targeted therapies, reflecting an incomplete understanding of its molecular pathogenesis. Here, we identify SMAD4 as a critical regulator of three-dimensional (3D) genome organization in LUSC and uncover a mechanistic link between tumor suppressor loss and oncogenic transcriptional activation. By integrating clinical datasets, genetically engineered mouse models, human and murine LUSC cell lines, and multi-omics analyses, we demonstrate that SMAD4 deficiency promotes LUSC progression by unleashing EP300-mediated enhancer-promoter looping at the SOX2 locus. Mechanistically, SMAD4 does not directly bind SOX2 regulatory elements but instead constrains chromatin looping by sequestering EP300 away from loop anchor regions. Loss of SMAD4 leads to enhanced H3K27ac deposition, aberrant SOX2 activation, and increased LUSC tumor cell proliferation. Together, these findings reveal a non-canonical role for a transcription factor (e.g., SMAD4) in regulating dysregulated 3D genome architecture to inhibit tumor development.

PMID:42189071 | DOI:10.1002/advs.75839

A Non-Canonical Role of SMAD4 in Regulating 3D Genome Architecture to Inhibit Lung Squamous Cell Carcinoma Development

Adv Sci (Weinh). 2026 May 26:e75839. doi: 10.1002/advs.75839. Online ahead of print.

ABSTRACT

Lung squamous cell carcinoma (LUSC) lacks clearly defined key drivers and effective targeted therapies, reflecting an incomplete understanding of its molecular pathogenesis. Here, we identify SMAD4 as a critical regulator of three-dimensional (3D) genome organization in LUSC and uncover a mechanistic link between tumor suppressor loss and oncogenic transcriptional activation. By integrating clinical datasets, genetically engineered mouse models, human and murine LUSC cell lines, and multi-omics analyses, we demonstrate that SMAD4 deficiency promotes LUSC progression by unleashing EP300-mediated enhancer-promoter looping at the SOX2 locus. Mechanistically, SMAD4 does not directly bind SOX2 regulatory elements but instead constrains chromatin looping by sequestering EP300 away from loop anchor regions. Loss of SMAD4 leads to enhanced H3K27ac deposition, aberrant SOX2 activation, and increased LUSC tumor cell proliferation. Together, these findings reveal a non-canonical role for a transcription factor (e.g., SMAD4) in regulating dysregulated 3D genome architecture to inhibit tumor development.

PMID:42189071 | DOI:10.1002/advs.75839

Test-Time Deep Thinking to Explore Implicit Rules

arXiv:2605.24828v1 Announce Type: new Abstract: With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in environments governed by implicit rules--hidden constraints that cannot be observed directly and must be inferred through interaction. This causes agents to fall into repetitive trial-and-error loops, ultimately leading to task failure. To address this challenge, we propose Test-Time Exploration (TTExplore), a framework where a thinker component analyzes interaction history to infer these implicit rules and guide an actor. Effective exploration in this setting critically depends on the reasoning ability of the thinker. However, evaluating deep reasoning trajectories is inherently unstable and difficult, which poses a major obstacle to effective training. To overcome this issue, we introduce a novel and stable reinforcement learning pipeline. The core idea is to use accurate task-level scores as indirect rewards to bypass the difficulty of evaluating intermediate reasoning, and to retain only a single thinking node per trajectory to alleviate reward sparsity. Using this pipeline, we train a specialized 7B model, Exp-Thinker. Experiments on five text-based embodied tasks show that TTExplore equipped with Exp-Thinker improves baseline agent performance by an average of $14$-$19$ points, demonstrating the effectiveness of explicitly reasoning about implicit rules.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

arXiv:2605.23954v1 Announce Type: cross Abstract: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18\%$\uparrow$ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02\%$\uparrow$ in Acc, 3.89\%$\uparrow$ in Noisy, and 4.53\%$\uparrow$ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models

arXiv:2512.18735v2 Announce Type: replace-cross Abstract: Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding. However, their capacity to comprehend the dynamic changes of objects within a shared spatial context between two distinct video observations, remains largely unexplored. This ability to reason about transformations within a consistent environment is particularly crucial for advancements in the field of spatial intelligence. In this paper, we introduce $M^3-Verse$, a Multi-Modal, Multi-State, Multi-Dimensional benchmark, to formally evaluate this capability. It is built upon paired videos that provide multi-perspective observations of an indoor scene before and after a state change. The benchmark contains a total of 270 scenes and 2,932 questions, which are categorized into over 50 subtasks that probe 4 core capabilities. We evaluate 16 state-of-the-art LMMs and observe their limitations in tracking state transitions. To address these challenges, we further propose a simple yet effective baseline that achieves significant performance improvements in multi-state perception. $M^3-Verse$ thus provides a challenging new testbed to catalyze the development of next-generation models with a more holistic understanding of our dynamic visual world. You can get the construction pipeline from https://github.com/Wal-K-aWay/M3-Verse_pipeline and full benchmark data from https://www.modelscope.cn/datasets/WalKaWay/M3-Verse.

Large Language Models for Combinatorial Optimization of Design Structure Matrix

arXiv:2506.09749v3 Announce Type: replace-cross Abstract: In complex engineering systems, the dependencies among components or development activities are often modeled and analyzed using Design Structure Matrix (DSM). Reorganizing elements within a DSM to minimize feedback loops and enhance modularity or process efficiency constitutes a challenging combinatorial optimization (CO) problem in engineering design and operations. As problem sizes increase and dependency networks become more intricate, traditional optimization methods that rely solely on mathematical heuristics often fail to capture the contextual nuances and struggle to deliver effective solutions. In this study, we explore the potential of Large Language Models (LLMs) to address such CO problems by leveraging their capabilities for advanced reasoning and contextual understanding. We propose a novel LLM-based framework that integrates network topology with contextual domain knowledge for iterative optimization of DSM sequencing-a common CO problem. Experiments on various DSM cases demonstrate that our method consistently achieves faster convergence and superior solution quality compared to both stochastic and deterministic baselines. Notably, incorporating contextual domain knowledge significantly enhances optimization performance regardless of the chosen LLM backbone. These findings highlight the potential of LLMs to solve complex engineering CO problems by combining semantic and mathematical reasoning. This approach paves the way towards a new paradigm in LLM-based engineering design optimization.

Investigating the Effect of Hospital Infection Control Informatization on Optimizing Microbiological Specimen Submission Before Antibiotic Therapy: Failure Mode and Effects Analysis

Background: Antimicrobial resistance (AMR) poses a critical global health threat, with inappropriate antibiotic use being a major driver. Timely microbiological specimen submission before initiating antibiotic therapy is a cornerstone of antimicrobial stewardship (AMS), enabling pathogen-directed therapy and reducing unnecessary broad-spectrum exposure. However, suboptimal compliance remains common due to workflow interruptions, technological barriers, and behavioral factors. Failure Mode and Effects Analysis (FMEA), a proactive risk-assessment method widely used in health care quality improvement, provides a systematic framework to identify process vulnerabilities and prioritize corrective actions. Despite its increasing application, few studies have integrated FMEA with hospital informatization to optimize microbiological specimen submission workflows in routine AMS practice. Objective: This study aimed to systematically identify workflow risks affecting preantibiotic microbiological specimen submission and to design, implement, and evaluate informatization-enabled interventions using an FMEA-based framework. Methods: FMEA was conducted at a tertiary hospital in China. A multidisciplinary team identified potential failure modes across 4 domains: health information systems, personnel, administration, and external support. Risk Priority Numbers (RPNs) and Action Priority (AP) indices were calculated for each failure mode. Targeted interventions were implemented, including dual-verification barcode scanning, artificial intelligence-driven clinical decision support alerts, EHR-integrated training modules, and automated compliance dashboards. Pre- and postintervention specimen submission rates (January 2024-December 2024) were analyzed using the Mann-Kendall trend test. Results: The top 5 failure modes included PDA barcode scanning failures (RPN=175), inadequate clinical decision support (RPN=140), insufficient clinician awareness (RPN=56), suboptimal oversight mechanisms, and patient-related barriers. Postintervention, significant upward trends were observed in overall specimen submission rates (
❌