❌

Reading view

CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks

arXiv:2508.11360v2 Announce Type: replace Abstract: As autonomous agents become adept at understanding and interacting with graphical user interface (GUI) environments, a new era of automated task execution is emerging. Recent studies have demonstrated that Reinforcement Learning (RL) can effectively enhance agents' performance in dynamic interactive GUI environments. However, these methods face two key limitations: (1) they overlook the significant variation in difficulty across different GUI tasks by treating the entire training data as a uniform set, which hampers the agent's ability to adapt its learning process; and (2) most approaches collapse task-specific nuances into a single, coarse reward, leaving the agent with a uniform signal that yields inefficient policy updates. To address these limitations, we propose CRAFT-GUI, a curriculum learning framework based on Group Relative Policy Optimization (GRPO) that explicitly accounts for the varying difficulty across trajectories. To enable more fine-grained policy optimization, we design a reward function that combines simple rule-based signals with model-judged evaluation, providing richer and more nuanced feedback during training. Experimental results demonstrate that our method achieves significant improvements over previous state-of-the-art approaches, outperforming them by 5.6% on public benchmarks Android Control and 10.3% on our internal online benchmarks, respectively. These findings empirically validate the effectiveness of integrating reinforcement learning with curriculum learning in GUI interaction tasks.
  •  

FNDC1 Competitively Binds Gbeta2 to Suppress the beta-Catenin-Destruction Complex and Promote Gastric Cancer Malignancy

FASEB J. 2026 Mar 31;40(6):e71634. doi: 10.1096/fj.202503587R.

ABSTRACT

Gastric cancer (GC) is a leading cause of cancer-related deaths and has high recurrence rate. Although fibronectin domain-containing protein 1 (FNDC1) is implicated in GC progression, its molecular mechanisms remain unclear. Multi-omics analyses (TCGA, GEO datasets) were used to assess FNDC1 expression and clinical correlation. In vitro (cell proliferation, invasion, EMT markers) and in vivo (xenograft) experiments, combined with molecular assays (Co-IP, WB, ChIP), explored FNDC1's function and mechanism. FNDC1 was significantly upregulated in GC, correlating with advanced clinicopathological features and poor prognosis. Knockdown of FNDC1 suppressed GC cell proliferation, invasion, and metastasis by inhibiting EMT and Wnt/Ξ²-catenin signaling. Mechanistically, FNDC1 competitively bound the WD5 domain (residues 224-254) of GΞ²2, disrupting GΞ²Ξ³-Dvl1 interaction. This prevented Dvl1 degradation, promoted Axin1 ubiquitination, and destabilized the Ξ²-catenin-destruction complex (GSK3 Ξ²-APC-Axin1), leading to Ξ²-catenin accumulation and Wnt pathway activation. FNDC1 drives GC malignancy by targeting the GΞ²2-Dvl1 axis to activate Wnt/Ξ²-catenin signaling, suggesting FNDC1 as a novel prognostic biomarker and therapeutic target.

PMID:41808415 | PMC:PMC12976582 | DOI:10.1096/fj.202503587R

  •  

Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single-Image Denoising

arXiv:2502.06432v3 Announce Type: replace-cross Abstract: Many studies have concentrated on constructing supervised models utilizing paired datasets for image denoising, which proves to be expensive and time-consuming. Current self-supervised and unsupervised approaches typically rely on blind-spot networks or sub-image pairs sampling, resulting in pixel information loss and destruction of detailed structural information, thereby significantly constraining the efficacy of such methods. In this paper, we introduce Prompt-SID, a prompt-learning-based single image denoising framework that emphasizes preserving of structural details. This approach is trained in a self-supervised manner using downsampled image pairs. It captures original-scale image information through structural encoding and integrates this prompt into the denoiser. To achieve this, we propose a structural representation generation model based on the latent diffusion process and design a structural attention module within the transformer-based denoiser architecture to decode the prompt. Additionally, we introduce a scale replay training mechanism, which effectively mitigates the scale gap from images of different resolutions. We conduct comprehensive experiments on synthetic, real-world, and fluorescence imaging datasets, showcasing the remarkable effectiveness of Prompt-SID. Our code will be released at https://github.com/huaqlili/Prompt-SID.
  •  

From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning

arXiv:2603.03825v1 Announce Type: cross Abstract: The cold-start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attention Score (VAS), an attention-based metric that quantifies how much a model attends to visual tokens. We find that reasoning performance is strongly correlated with VAS (r=0.9616): models with higher VAS achieve substantially stronger multimodal reasoning. Surprisingly, multimodal cold-start fails to elevate VAS, resulting in attention distributions close to the base model, whereas text-only cold-start leads to a clear increase. We term this counter-intuitive phenomenon Lazy Attention Localization. To validate its causal role, we design training-free interventions that directly modulate attention allocation during inference, performance gains of 1$-$2% without any retraining. Building on these insights, we further propose Attention-Guided Visual Anchoring and Reflection (AVAR), a comprehensive cold-start framework that integrates visual-anchored data synthesis, attention-guided objectives, and visual-anchored reward shaping. Applied to Qwen2.5-VL-7B, AVAR achieves an average gain of 7.0% across 7 multimodal reasoning benchmarks. Ablation studies further confirm that each component of AVAR contributes step-wise to the overall gains. The code, data, and models are available at https://github.com/lrlbbzl/Qwen-AVAR.
  •  
❌