❌

Reading view

Stopped-light-enhanced gravitational force sensing

Nature Nanotechnology, Published online: 29 September 2026; doi:10.1038/s41565-026-02298-8

A torsion-pendulum gravitational-force sensor with an optical microcavity readout leverages coupled photon–phonon effects to enhance sensitivity to the tiny gravitational pull from a millimetre-sized source mass.
  •  

Transformer-Based Multitask Framework Integrating Habitat and Deep Learning for Predicting Early Disease Control and Survival in Immunotherapy-Treated Hepatocellular Carcinoma

Adv Sci (Weinh). 2026 Sep 27:e78005. doi: 10.1002/advs.78005. Online ahead of print.

ABSTRACT

Hepatocellular carcinoma (HCC) patients show heterogeneous responses to immune checkpoint inhibitors (ICIs). This study developed ECOS-Net, a transformer-based multitask network integrating CT-derived habitat and 2.5-dimensional (2.5D) deep learning features for simultaneously predicting early disease control (DC) and overall survival (OS). Of 1,234 patients with HCC enrolled from eight institutions and public databases, 832 ICI-treated patients were used for model development. ECOS-Net fused features using multi-head attention and generated early DC probabilities and OS risk scores. ECOS-DC achieved AUCs of 0.836, 0.822, and 0.817 in training, internal validation, and external test sets, outperforming clinical models (all p values < 0.05). ECOS-OS yielded C-indices of 0.730, 0.722, and 0.720, respectively. Integrated models also showed favorable external performance (early DC AUC: 0.825; OS C-index: 0.741). Patients with higher ECOS-DC probabilities had a higher likelihood of early DC, whereas those with higher ECOS-OS risk had shorter OS, with directionally consistent associations across most subgroups. Exploratory biological analyses suggested that the higher ECOS-DC probability and lower ECOS-OS risk groups were associated with immune-active tumor microenvironment features. Therefore, ECOS-Net shows potential as a non-invasive imaging-based risk stratification framework for simultaneously predicting early DC and OS in ICI-treated HCC patients.

PMID:42801546 | PMC:PMC13616327 | DOI:10.1002/advs.78005

  •  

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.
  •  

UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction

arXiv:2609.12898v1 Announce Type: cross Abstract: Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundation models tend to be either generalized but object-aware, or part-aware but limited to closed-set taxonomies, which weakens zero-shot transfer. We study text-conditioned 3D part segmentation, where a free-form phrase selects a functional part on point cloud. We introduce UniPart, a feed-forward cross-modal 3D Transformer that conditions CLIP text embedding. To scale supervision, we build LangPart-1M with 160K+ Objaverse assets and 8M text to part pairs using multi-view consistent part generation. We further manually label a high-quality subset, LangPart-4K, for fine-tuning and evaluation. UniPart achieves strong zero-shot results on open-vocabulary part benchmarks and transfers to language-conditioned part grasping in real world.
  •  

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

arXiv:2507.01663v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.
  •  

BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation Models

arXiv:2609.10518v1 Announce Type: cross Abstract: fMRI foundation models increasingly aggregate heterogeneous data across brain states, cohorts, and acquisition settings, yet pretraining domains are commonly treated as a flat mixture and downstream tasks are adapted independently. We study whether measured learning relations can organize both stages without modifying the backbone. During pretraining, a lightweight Brain-DiT proxy estimates difficulty and directed facilitation across ten fMRI domains, yielding a priority-guided cumulative domain curriculum combined with high-to-low-noise timestep scheduling and joint consolidation. During adaptation, controlled first- and higher-order transfer across fifteen tasks constructs a directed taskonomy, from which budgeted integer programming (BIP) selects directly supervised source tasks and target-specific routes. The joint priority-domain and high-to-low-timestep curriculum reduces v-NMSE, PSD-NMSE, and FC-MSE by 6.5%, 16.3%, and 10.5%, respectively, relative to uniform sampling over both dimensions, and shows strong downstream performance across six in- and out-of-domain tasks. The taskonomy reveals asymmetric, target-dependent transfer, while exploratory sealed-test evaluation shows larger descriptive gains for BIP policies when higher-order route spaces are available than for matched random controls. Together, these findings support organizing fMRI pretraining and adaptation by measured learning relations rather than treating domains and tasks as independent flat sets.
  •  

Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks

arXiv:2505.03380v2 Announce Type: replace-cross Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains time-consuming and expertise-intensive. Existing artificial intelligence systems often require manual spatial prompts or task-specific retraining, while generic class labels provide limited semantic grounding for heterogeneous disease targets. Here we present SyRe, a promptable segmentation foundation model based on Synergistic vision-language Reinforcement. SyRe strengthens bidirectional interaction between visual and linguistic representations to improve semantically grounded spatial understanding. To support large-scale training, we introduce the Color Region Description strategy and construct SyReData, comprising 20 million image-mask-description triplets across 9 modalities and 229 segmentation tasks. Training with diversified prompt forms further enables open-ended prompting, invalid-prompt rejection and flexible switching between single- and multi-target analysis. SyRe achieves accurate text-prompted segmentation across diverse clinical scenarios, with particularly strong performance on disease-related targets. Across 28 unseen external datasets, including 20 cancer types and multinational in-house cohorts, SyRe generalizes robustly under real-world distribution shifts. SyRe-generated masks also preserve clinically relevant quantitative information in pathology and yield radiomics features that stratify survival and improve prognostic modeling across five retrospective CT and MRI tumor cohorts. Finally, clinician-in-the-loop refinement enables efficient case-level correction when greater precision is required. These results establish SyRe as a generalizable foundation for scalable quantitative oncology and clinician-guided segmentation refinement.
  •  

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

arXiv:2605.26114v1 Announce Type: new Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.
  •  

Multi-omics integration and Mendelian randomization reveal the mechanisms and experimental validation of curcumin targeting the RXRA-PI3K/AKT axis to enhance cisplatin sensitivity in gastric cancer

Front Oncol. 2026 Apr 15;16:1791971. doi: 10.3389/fonc.2026.1791971. eCollection 2026.

ABSTRACT

OBJECTIVE: This study aimed to integrate multi-omics analyses with genetic causal inference to identify key genes associated with cisplatin resistance in gastric cancer and to evaluate the potential mechanism by which curcumin enhances cisplatin sensitivity through relevant pathways.

METHODS: Cisplatin resistance-related transcriptomic datasets(GSE14210 and GSE31811) and a gastric cancer single-cell transcriptomic dataset (GSE183904) were obtained from the Gene Expression Omnibus(GEO)database. Differential expression analysis was performed to identify resistance-associated differentially expressed genes(DEGs),followed by GO and KEGG enrichment analyses. Putative curcumin targets were collected and intersected with DEGs to obtain candidate genes. Mendelian randomization (MR) analysis was conducted using the TwoSampleMR framework to evaluate the genetic association between RXRA expression and gastric cancer risk, with robustness and sensitivity analyses based on multiple MR methods. RXRA expression was further evaluated, along with pathway activity assessment using GSEA and GSVA, and molecular docking was performed to explore the potential binding of curcumin to RXRA. In vitro experiments were performed using the cisplatin-resistant gastric cancer cell lineNCI-N87/DDP. Drug effects and chemosensitization under combination treatment were assessed by CCK-8 assays, synergy was evaluated using the combination index(CI),and changes in key proteins in thePI3K/AKT pathway were measured by Western blotting.

RESULTS: A total of 595 DEGs associated with cisplatin resistance were identified. Functional enrichment analyses indicated that these DEGs were mainly involved in extracellular matrix remodeling and adhesion, secretion and vesicular transport, and signaling pathways including PI3K-Akt.The intersection of curcumin targets with DEGs highlighted RXRA as a key candidate gene. MR results indicated that genetically predicted increased RXRA expression was significantly associated with elevated gastric cancer risk (OR = 4.216,95%CI:1.201-14.797,P=0.025). GSEA and GSVA suggested that high RXRA expression was associated with altered activity of pathways related to lysosome, proteasome, oxidative phosphorylation, and the pentose phosphate pathway. Single-cell analysis indicated that RXRA was mainly expressed in tissue stem cells and fibroblasts. Molecular docking predicted a feasible interaction between curcumin and RXRA. In vitro experiments demonstrated that curcumin inhibited the viability of resistant cells and showed a synergistic trend when combined with cisplatin. Western blotting revealed decreased p-PI3K and p-AKT levels following curcumin treatment, supporting an inhibitory effect on the PI3K/AKT pathway.

CONCLUSION: These findings highlight RXRA as a candidate gene associated with cisplatin resistance-related programs in gastric cancer. Curcumin may enhance cisplatin sensitivity by influencing RXRA-associated transcriptional networks and suppressing PI3K/AKT signaling. This study provides new candidate targets and experimental evidence for mechanistic investigation and combination treatment strategies to overcome cisplatin resistance in gastric cancer.

PMID:42063729 | PMC:PMC13124633 | DOI:10.3389/fonc.2026.1791971

  •  

Mapping convergent regulators of melanoma drug resistance by PerturbFate

Nature, Published online: 15 April 2026; doi:10.1038/s41586-026-10367-0

PerturbFate is a high-throughput, cost-effective, single-cell platform that systematically profiles CRISPR interference perturbations to reveal common regulatory nodes and convergent phenotypic states across diverse genetic alterations linked to vemurafenib resistance in melanoma cells.
  •  

PRET is a few-shot system for pan-cancer recognition without example training

Nature Cancer, Published online: 03 April 2026; doi:10.1038/s43018-026-01141-2

Li et al. present PRET, a few-shot system for pan-cancer detection not requiring model fine-tuning, validated it in multicenter datasets and found that it outperformed existing approaches across tasks and pathologists in lymph node metastasis detection.
  •  

Unraveling the role of cuproptosis in pulmonary fibrosis pathogenesis and prognosis: an integrative single-cell transcriptomics and microarray analysis

Mol Cell Biochem. 2026 Mar 13. doi: 10.1007/s11010-026-05510-4. Online ahead of print.

ABSTRACT

Pulmonary fibrosis (PF), a progressive interstitial lung disease with elusive pathogenesis, remains a therapeutic challenge. Emerging evidence suggests cuproptosis-a copper-dependent cell death pathway-may play a regulatory role in disease progression. This study aims to elucidate cuproptosis's biological function and establish a prognostic model for PF. Through integrative analysis of single-cell RNA-seq data from bleomycin (BLM)-induced mouse models and bulk RNA-seq data from idiopathic pulmonary fibrosis (IPF) patients, we identified cuproptosis-related genes (CRGs) using LASSO regression and Cox regression. A novel 4-CRG signature (LIAS, LIPT1, ATP7A, PDHB) was constructed to stratify patients into distinct risk groups in the GSE70866 cohort, where high-risk individuals exhibited poorer survival and enhanced extracellular matrix/lipid metabolism activity via GO/KEGG analysis. Experimental validation in BLM-induced mouse models, TGF-β1-stimulated fibroblast-to-myofibroblast transition assays, and human IPF specimens demonstrated significant downregulation of CRGs through qRT-PCR and immunohistochemical analyses. Functional assays revealed impaired cell viability and elevated cuproptosis markers in fibrotic microenvironments. Our findings establish an inverse correlation between cuproptosis and PF progression, and propose a robust risk-score model for clinical prognosis prediction. This multi-omics approach provides new insights into copper-mediated regulatory mechanisms in fibrogenesis.

PMID:41824199 | DOI:10.1007/s11010-026-05510-4

  •  

Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System

arXiv:2603.07449v1 Announce Type: cross Abstract: Enterprises commonly deploy heterogeneous database systems, each of which owns a distinct SQL dialect with different syntax rules, built-in functions, and execution constraints. However, most existing NL2SQL methods assume a single dialect (e.g., SQLite) and struggle to produce queries that are both semantically correct and executable on target engines. Prompt-based approaches tightly couple intent reasoning with dialect syntax, rule-based translators often degrade native operators into generic constructs, and multi-dialect fine-tuning suffers from cross-dialect interference. In this paper, we present Dial, a knowledge-grounded framework for dialect-specific NL2SQL. Dial introduces: (1) a Dialect-Aware Logical Query Planning module that converts natural language into a dialect-aware logical query plan via operator-level intent decomposition and divergence-aware specification; (2) HINT-KB, a hierarchical intent-aware knowledge base that organizes dialect knowledge into (i) a canonical syntax reference, (ii) a declarative function repository, and (iii) a procedural constraint repository; and (3) an execution-driven debugging and semantic verification loop that separates syntactic recovery from logic auditing to prevent semantic drift. We construct DS-NL2SQL, a benchmark covering six major database systems with 2,218 dialect-specific test cases. Experimental results show that Dial consistently improves translation accuracy by 10.25% and dialect feature coverage by 15.77% over state-of-the-art baselines. The code is at https://github.com/weAIDB/Dial.
  •  

Solution to the 10th ABAW Expression Recognition Challenge: A Robust Multimodal Framework with Safe Cross-Attention and Modality Dropout

arXiv:2603.08034v1 Announce Type: cross Abstract: Emotion recognition in real-world environments is hindered by partial occlusions, missing modalities, and severe class imbalance. To address these issues, particularly for the Affective Behavior Analysis in-the-wild (ABAW) Expression challenge, we propose a multimodal framework that dynamically fuses visual and audio representations. Our approach uses a dual-branch Transformer architecture featuring a safe cross-attention mechanism and a modality dropout strategy. This design allows the network to rely on audio-based predictions when visual cues are absent. To mitigate the long-tail distribution of the Aff-Wild2 dataset, we apply focal loss optimization, combined with a sliding-window soft voting strategy to capture dynamic emotional transitions and reduce frame-level classification jitter. Experiments demonstrate that our framework effectively handles missing modalities and complex spatiotemporal dependencies, achieving an accuracy of 60.79% and an F1-score of 0.5029 on the Aff-Wild2 validation set.
  •  

SCL-GNN: Towards Generalizable Graph Neural Networks via Spurious Correlation Learning

arXiv:2603.08270v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have demonstrated remarkable success across diverse tasks. However, their generalization capability is often hindered by spurious correlations between node features and labels in the graph. Our analysis reveals that GNNs tend to exploit imperceptible statistical correlations in training data, even when such correlations are unreliable for prediction. To address this challenge, we propose the Spurious Correlation Learning Graph Neural Network (SCL-GNN), a novel framework designed to enhance generalization on both Independent and Identically Distributed (IID) and Out-of-Distribution (OOD) graphs. SCL-GNN incorporates a principled spurious correlation learning mechanism, leveraging the Hilbert-Schmidt Independence Criterion (HSIC) to quantify correlations between node representations and class scores. This enables the model to identify and mitigate irrelevant but influential spurious correlations effectively. Additionally, we introduce an efficient bi-level optimization strategy to jointly optimize modules and GNN parameters, preventing overfitting. Extensive experiments on real-world and synthetic datasets demonstrate that SCL-GNN consistently outperforms state-of-the-art baselines under various distribution shifts, highlighting its robustness and generalization capabilities.
  •  

HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts

arXiv:2603.01396v2 Announce Type: replace Abstract: Single-cell perturbation studies face dual heterogeneity bottlenecks: (i) semantic heterogeneity--identical biological concepts encoded under incompatible metadata schemas across datasets; and (ii) statistical heterogeneity--distribution shifts from biological variation demanding dataset-specific inductive biases. We propose HarmonyCell, an end-to-end agent framework resolving each challenge through a dedicated mechanism: an LLM-driven Semantic Unifier autonomously maps disparate metadata into a canonical interface without manual intervention; and an adaptive Monte Carlo Tree Search engine operates over a hierarchical action space to synthesize architectures with optimal statistical inductive biases for distribution shifts. Evaluated across diverse perturbation tasks under both semantic and distribution shifts, HarmonyCell achieves a 95% valid execution rate on heterogeneous input datasets (versus 0% for general agents) while matching or even exceeding expert-designed baselines in rigorous out-of-distribution evaluations. This dual-track orchestration enables scalable automatic virtual cell modeling without dataset-specific engineering.
  •  

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

arXiv:2507.01352v3 Announce Type: replace-cross Abstract: Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced human preferences. We hypothesize that this brittleness stems primarily from limitations in preference datasets, which are often narrowly scoped, synthetically labeled, or lack rigorous quality control. To address these challenges, we present SynPref-40M, a large-scale preference dataset comprising 40 million preference pairs. To enable data curation at scale, we design a human-AI synergistic two-stage pipeline that leverages the complementary strengths of human annotation quality and AI scalability. In this pipeline, humans provide verified annotations, while LLMs perform automatic curation based on human guidance. Training on this preference mixture, we introduce Skywork-Reward-V2, a suite of eight reward models ranging from 0.6B to 8B parameters, trained on a carefully curated subset of 26 million preference pairs from SynPref-40M. We demonstrate that Skywork-Reward-V2 is versatile across a wide range of capabilities, including alignment with human preferences, objective correctness, safety, resistance to stylistic biases, and best-of-N scaling. These reward models achieve state-of-the-art performance across seven major reward model benchmarks, outperform generative reward models, and demonstrate strong downstream performance. Ablation studies confirm that effectiveness stems not only from data scale but also from high-quality curation. The Skywork-Reward-V2 series represents substantial progress in open reward models, demonstrating how human-AI curation synergy can unlock significantly higher data quality.
  •  

Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation

arXiv:2602.19822v1 Announce Type: cross Abstract: Early detection of myometrial invasion is critical for the staging and life-saving management of endometrial carcinoma (EC), a prevalent global malignancy. Transvaginal ultrasound serves as the primary, accessible screening modality in resource-constrained primary care settings; however, its diagnostic reliability is severely hindered by low tissue contrast, high operator dependence, and a pronounced scarcity of positive pathological samples. Existing artificial intelligence solutions struggle to overcome this severe class imbalance and the subtle imaging features of invasion, particularly under the strict computational limits of primary care clinics. Here we present an automated, highly efficient two-stage deep learning framework that resolves both data and computational bottlenecks in EC screening. To mitigate pathological data scarcity, we develop a structure-guided cross-modal generation network that synthesizes diverse, high-fidelity ultrasound images from unpaired magnetic resonance imaging (MRI) data, strictly preserving clinically essential anatomical junctions. Furthermore, we introduce a lightweight screening network utilizing gradient distillation, which transfers discriminative knowledge from a high-capacity teacher model to dynamically guide sparse attention towards task-critical regions. Evaluated on a large, multicenter cohort of 7,951 participants, our model achieves a sensitivity of 99.5\%, a specificity of 97.2\%, and an area under the curve of 0.987 at a minimal computational cost (0.289 GFLOPs), substantially outperforming the average diagnostic accuracy of expert sonographers. Our approach demonstrates that combining cross-modal synthetic augmentation with knowledge-driven efficient modeling can democratize expert-level, real-time cancer screening for resource-constrained primary care settings.
  •  

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

arXiv:2510.10689v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.
  •  
❌