❌

Reading view

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.
  •  

CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test

arXiv:2605.23491v2 Announce Type: replace-cross Abstract: Recently, Reinforcement Learning with Verifiable Rewards (RLVR) and Test-Time Scaling (TTS) have advanced LLM code generation through executable verification. Yet Ground-Truth Unit Tests (GT UTs) remain a bottleneck: SOTA RLVR methods require them for costly training, while existing TTS methods lose competitiveness without them. This motivates GT-free TTS, where existing methods directly use self-generated UTs to refine and select code candidates. Yet such UTs are often noisy or spuriously coupled with wrong code, and UT quality in turn cannot be validated without reliable code. The key challenge is therefore to jointly improve both. To this end, we present CoSPlay, a GT-free, training-free framework that jointly improves codes and UTs through cooperative self-play. It first explores diverse solution ideas and identifies their potential failure modes to produce discriminative UT ideas. It then uses bidirectional pass-count signals from the Code-UT execution matrix to iteratively prune or fix weak codes and refresh or replace unreliable UTs, letting the two pools co-evolve. Finally, when multiple codes remain tied at the highest pass count, it picks the final code from the largest output-consensus cluster, since correct codes agree on the same inputs while wrong codes diverge. Experiments on four challenging benchmarks show that CoSPlay on Qwen2.5-7B-Instruct improves average BoN from 22.1% to 33.2% and UT accuracy from 14.6% to 78.3%, matching or surpassing the RLVR model CURE-7B. When applied to CURE-7B, it further improves BoN by 5.7%. CoSPlay also generalizes across diverse backbones and outperforms GT-free TTS baselines under comparable token budgets, with continued gains as the budget scales up. These results suggest a scalable inference strategy for competitive code generation without any GT data.
  •  

PRXL2B facilitates the progression of hepatocellular carcinoma and the therapeutic efficacy of oncolytic adenovirus H101 through the PI3K/AKT/PD-L1 axis

Biosci Trends. 2026 May 21. doi: 10.5582/bst.2026.01000. Online ahead of print.

ABSTRACT

Oncolytic adenovirus H101 has shown antitumor activity in hepatocellular carcinoma (HCC), but the molecular determinants of treatment response remain unclear. In this study, a Hepa1-6 subcutaneous tumor model was established in C57BL/6 mice and treated with intratumoral H101, followed by integrated transcriptomic and proteomic analyses to identify candidate genes associated with H101 response. PRXL2B was selected for further investigation using public multi-omics datasets, tissue microarray-based immunohistochemistry, in vitro functional assays, mechanistic analyses, and in vivo validation experiments. Integrated multi-omics analyses identified PRXL2B as a candidate gene downregulated after H101 treatment. Public datasets and tissue-based validation further showed that PRXL2B was upregulated in HCC tissues. In MHCC97H and HCCLM3 cells, PRXL2B knockdown inhibited proliferation, migration, and invasion, promoted apoptosis and cell-cycle arrest, and enhanced the antitumor effect of H101. Mechanistically, PRXL2B silencing reduced AKT phosphorylation and PD-L1 expression. In vivo, PRXL2B knockdown suppressed tumor growth, and the combination of PRXL2B knockdown and H101 produced the strongest antitumor effect. These findings indicate that PRXL2B promotes malignant phenotypes in HCC and may modulate H101 efficacy through the PI3K/AKT/PD-L1 axis. Targeting PRXL2B may therefore represent a potential strategy to enhance the therapeutic efficacy of oncolytic virus therapy in HCC.

PMID:42161529 | DOI:10.5582/bst.2026.01000

  •  
❌