❌

Normal view

ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents

arXiv:2603.20340v2 Announce Type: replace-cross Abstract: Self-generated skills for web agents are often unstable and can even hurt performance relative to direct acting. We argue that the key bottleneck is not only skill generation quality, but the fact that web skills remain implicit and therefore cannot be checked or locally repaired. To address this, we present ContractSkill, a framework that converts a draft skill into an executable artifact with explicit procedural structure, enabling deterministic verifica tion, fault localization, and minimal local repair. This turns skill refinement from full rewriting into localized editing of a single skill artifact. Experiments on VisualWebArena show that Contract Skill is effective in realistic web environments, while MiniWoB provides a controlled test of the mechanism behind the gain. Under matched transfer layers, repaired artifacts also remain reusable after removing the source model from the loop, providing evi dence of portability within the same benchmark family rather than full-benchmark generalization. These results suggest that the central challenge is not merely generating skills, but mak ing them explicit, executable, and repairable. Code is available at https://github.com/underfitting-lu/contractskill.git.

Proposed Role of Circadian Clock Genes in Pathogenesis of HCC: Molecular Subtyping and Characterization

28 March 2026 at 18:00

Biomedicines. 2026 Mar 12;14(3):645. doi: 10.3390/biomedicines14030645.

ABSTRACT

Background: Hepatocellular carcinoma (HCC) stands as a prevalent global health issue with increasing incidence and mortality rates. Hepatocellular carcinoma (HCC) exhibits profound molecular and clinical heterogeneity, which limits the effectiveness of current therapeutic strategies. Circadian rhythm disruption has been implicated in metabolic reprogramming, proliferation, and immune modulation in cancer, but its role in shaping HCC heterogeneity remains poorly defined. Methods: Four public HCC transcriptomic cohorts (TCGA-LIHC, CHCC, LIRI, LICA) were integrated using RMA normalization and ComBat for batch correction. Consensus clustering based on 31 core circadian clock genes (CCGs) identified robust molecular subtypes. Multi-omics characterization-including genomic alterations, pathway activity (GSEA/GSVA), immune microenvironment profiling (CIBERSORT, EPIC, MCP-counter, xCell), and drug-sensitivity prediction (pRRophetic/oncoPredict)-was performed to delineate subtype-specific biological properties. A nine-gene CCG-based RiskScore model was constructed using LASSO Cox regression to internally validate subtype robustness and intra-subtype risk stratification. Results: Using consensus clustering of 31 core CCGs in TCGA-LIHC and three independent validation cohorts (CHCC, LIRI, LICA), we identified three reproducible subtypes-Cluster-1 (metabolic-quiescent), Cluster-2 (transition-intermediate), and Cluster-3 (proliferation-inflammatory)-which were recapitulated across cohorts and showed distinct overall survival (Cluster-3 worst; log-rank p values significant across datasets). Multi-omic characterization revealed that Cluster-3 exhibits the highest tumor mutational burden and CNV burden with enrichment of TP53/AXIN1/TERT alterations, strong activation of cell-cycle, E2F, and G2M programs, and an immune-hot yet immunosuppressed microenvironment enriched for TAMs, Tregs and MDSCs. By contrast, Cluster-1 shows relative genomic stability, dominant hepatic metabolic signatures (fatty-acid oxidation, bile-acid and xenobiotic metabolism) and an immune-cold phenotype. Single-cell mapping linked ALAS1 expression to malignant hepatocytes predominating in Cluster-1, whereas NONO and CSNK1D localized to stromal (CAFs/TECs) and both malignant/immune compartments respectively in Cluster-3, providing a cellular mechanism for subtype-specific metabolism, angiogenesis and immune modulation. Finally, a nine-gene CCG-based RiskScore validated prognostic stratification and drug-sensitivity predictions indicated subtype-specific therapeutic vulnerabilities (notably increased predicted TKI sensitivity in Cluster-3). Conclusion: In conclusion, this study proposes a robust circadian rhythm-based molecular classification of hepatocellular carcinoma, revealing three biologically and clinically distinct subtypes characterized by divergent genomic alterations, metabolic programs, immune microenvironment states, and prognostic patterns. By integrating bulk and single-cell transcriptomic data, we identify subtype-specific roles of key circadian regulators-including ALAS1, NONO, and CSNK1D-in shaping tumor metabolism, proliferation, stromal remodeling, and immune suppression. These findings highlight circadian dysregulation as a potential upstream factor associated with HCC heterogeneity and provide a conceptual framework for developing subtype-tailored mechanistic studies and circadian-informed therapeutic strategies.

PMID:41898292 | PMC:PMC13024568 | DOI:10.3390/biomedicines14030645

TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward

arXiv:2603.07700v1 Announce Type: cross Abstract: While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models remain an unsolved problem. Existing RL approaches for few-step diffusion models strongly rely on back-propagating through differentiable reward models, thereby excluding the majority of important real-world reward signals, e.g., non-differentiable rewards such as humans' binary likeness, object counts, etc. To properly incorporate non-differentiable rewards to improve few-step generative models, we introduce TDM-R1, a novel reinforcement learning paradigm built upon a leading few-step model, Trajectory Distribution Matching (TDM). TDM-R1 decouples the learning process into surrogate reward learning and generator learning. Furthermore, we developed practical methods to obtain per-step reward signals along the deterministic generation trajectory of TDM, resulting in a unified RL post-training method that significantly improves few-step models' ability with generic rewards. We conduct extensive experiments ranging from text-rendering, visual quality, and preference alignment. All results demonstrate that TDM-R1 is a powerful reinforcement learning paradigm for few-step text-to-image models, achieving state-of-the-art reinforcement learning performances on both in-domain and out-of-domain metrics. Furthermore, TDM-R1 also scales effectively to the recent strong Z-Image model, consistently outperforming both its 100-NFE and few-step variants with only 4 NFEs. Project page: https://github.com/Luo-Yihong/TDM-R1

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

arXiv:2508.05612v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients diminishes over time. These issues lead to suboptimal gradient updates and hinder long-term learning efficiency. To address these issues, we propose Shuffle-R1, a simple yet principled framework that improves RL fine-tuning efficiency by dynamically restructuring trajectory sampling and batch composition. It introduces (1) Pairwise Trajectory Sampling, which selects high-contrast trajectories with large advantages to improve gradient signal quality, and (2) Advantage-based Trajectory Shuffle, which increases exposure of valuable rollouts through informed batch reshuffling. Experiments across multiple reasoning benchmarks show that our framework consistently outperforms strong RL baselines with minimal overhead. These results highlight the importance of data-centric adaptations for more efficient RL training in MLLM.
❌