❌

Normal view

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

By: Yu Bai Β· Yukai Miao Β· Dawei Wang Β· Li Chen Β· Yanyu Ren Β· Yuqian Shi Β· Dan Li Β· Ying Xiong Β· Chengqiu Tan Β· Run Zhou Β· Li Li
14 September 2026 at 12:00
arXiv:2609.12404v1 Announce Type: new Abstract: Learning from trial and error is a promising way to improve language agents on complex tasks such as computer control. Reflexion introduced verbal reinforcement learning, which turns failed trials into text that guides later attempts without updating model parameters. We introduce VRL-Bench, a harness for fair evaluation of trial-and-error learning under finite trial budgets. Across three models on MiniWoB and WebShop, we evaluate updates from several prominent verbal-memory methods spanning Reflexion and later work: each improves observed success over memory-free retry in some settings but reduces it in others. Replay experiments show that using reflection can reduce success rates, revealing a trade-off between exploiting experience and continued exploration. We propose VEX$^2$, a verbal exploration--exploitation scheduler that uses a language model to jointly select policies and allocate the remaining trial budget. VEX$^2$ is the only evaluated update to achieve positive observed success-rate gains over retry in all six settings.

Confidence-Gated Transductive Test Generation for Code Reranking

arXiv:2609.12489v1 Announce Type: new Abstract: Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult to obtain. We propose Confidence-Gated Transductive Test Generation (CoTT), which first uses an efficient inductive procedure and invokes transductive generation only when inductive confidence is low. This adaptive design improves output reliability while allocating extra computation only when needed. On code reranking benchmarks, CoTT outperforms prior baselines across the reported metrics while reducing cost relative to applying transductive generation to every input. These results show that confidence-based allocation of test-time computation provides a favorable efficiency-effectiveness trade-off with a single efficient LLM.

Hepatic Usp2 orchestrates de novo lipogenesis through G3bp2 stabilization and Ξ²-catenin activation

Cell Death Discovery, Published online: 14 September 2026; doi:10.1038/s41420-026-03333-2

Hepatic Usp2 orchestrates de novo lipogenesis through G3bp2 stabilization and Ξ²-catenin activation
❌