❌

Reading view

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

arXiv:2602.12670v3 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurations over 7,308 trajectories. Curated Skills raise average pass rate by 16.2 percentage points(pp), but effects vary widely by domain (+4.5pp for Software Engineering to +51.9pp for Healthcare) and 16 of 84 tasks show negative deltas. Self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming. Focused Skills with 2--3 modules outperform comprehensive documentation, and smaller models with Skills can match larger models without them.
  •  

Longitudinal Effects of a Smartphone Game (Tumaini) for HIV Prevention Among Kenyan Adolescents: 45-Month Trajectories of Condom Use–Related Proximal Outcomes From a Randomized Controlled Trial

Background: African adolescents and young adults account for a disproportionate number of new HIV infections. There is an urgent need to identify scalable and cost-effective behavioral HIV prevention strategies for this population. Using a condom at first sex is associated with a higher likelihood of consistent use later. Tumaini (“Hope for the Future” in Swahili; Emory University) is a choose-your-own-adventure smartphone game that has been shown to reduce the risk of unprotected first sex by end line in a 45-month randomized controlled trial in western Kenya. Objective: This study aimed to assess the impact of Tumaini on proximal outcomes related to condom use at first sex (specifically, behavioral intentions, self-efficacy, attitudes, and knowledge) longitudinally across mid-adolescence in the above trial. Methods: Adolescent participants (n=996, mean baseline age 14, SD 0.56 years) were randomized 1:1 to receive either a smartphone loaded with Tumaini or an attention-control math game for 5 to 7 weeks at 3 time points (mean age 14.0, SD 0.56; 15.3, SD 0.55; and 16.0, SD 0.56 years, respectively). They completed a behavioral survey at 13 time points, through mean age 17.7 (SD 0.56) years. Using generalized estimating equations and controlling for age at baseline, we modeled mean scores (overall and stratified by gender) on a range of condom-related survey items over time to assess mean differences at specific time points. We applied appropriate Bonferroni corrections to inferences about cross-arm differences in mean changes relative to baseline at 4 time points (after each intervention period and at end line; α=.05/4) and within-arm mean changes relative to baseline at each of the 12 post-baseline time points (α=.05/12). Analyses were conducted as intent-to-treat. Results: At end line, 97.8% (n=974) of the sample had been retained. Participants in both arms dedicated a mean total of >30 hours to their assigned game. There was significant improvement across all condom-related proximal outcomes in the intervention arm relative to the control arm immediately after initial intervention exposure. For almost all outcomes, a significant cross-arm difference was also present at end line and for most outcomes at the 2 intervening comparison time points. Some outcomes saw stronger intervention effects on female participants (eg, self-efficacy to refuse unprotected sex) or male participants (eg, knowledge that condoms are an effective way to prevent HIV). In each arm, intention to use a condom at first sex was consistently higher among male participants; however, female intervention-arm scores overtook male control-arm scores following initial intervention exposure. Conclusions: Tumaini significantly improved theory-based proximal outcomes related to condom use, with effects sustained 45 months post initial exposure and 16 months post most recent exposure. Adolescents benefited from even short-term exposure, though repeated exposure generally sustained and reinforced intervention effects. As access to smartphones increases, Tumaini has potential for high scalability and impact on condom-related outcomes. Trial Registration: ClinicalTrials.gov NCT04437667; https://clinicaltrials.gov/study/NCT04437667
  •  

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

arXiv:2602.12670v2 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurations over 7,308 trajectories. Curated Skills raise average pass rate by 16.2 percentage points(pp), but effects vary widely by domain (+4.5pp for Software Engineering to +51.9pp for Healthcare) and 16 of 84 tasks show negative deltas. Self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming. Focused Skills with 2--3 modules outperform comprehensive documentation, and smaller models with Skills can match larger models without them.
  •  

CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think

arXiv:2603.02547v1 Announce Type: cross Abstract: We study why continuous diffusion language models (DLMs) have lagged behind discrete diffusion approaches despite their appealing continuous generative dynamics. Under a controlled token--recovery study, we identify token rounding, the final projection from denoised embeddings to tokens, as a primary bottleneck. Building on these insights, we propose CoDAR (Continuous Diffusion with Contextual AutoRegressive Decoder), a two--stage framework that keeps diffusion entirely continuous in an embedding space while learning a strong, context--conditional discretizer: an autoregressive Transformer decoder that cross--attends to the denoised embedding sequence and performs contextualized rounding to tokens. Experiments on LM1B and OpenWebText demonstrate that CoDAR substantially improves generation quality over latent diffusion and becomes competitive with strong discrete DLMs, while exposing a simple decoder--temperature knob to navigate the fluency--diversity trade off.
  •  

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

arXiv:2602.13954v1 Announce Type: cross Abstract: We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.
  •  

High Precision Audience Expansion via Extreme Classification in a Two-Sided Marketplace

arXiv:2602.14358v1 Announce Type: cross Abstract: Airbnb search must balance a worldwide, highly varied supply of homes with guests whose location, amenity, style, and price expectations differ widely. Meeting those expectations hinges on an efficient retrieval stage that surfaces only the listings a guest might realistically book, before resource intensive ranking models are applied to determine the best results. Unlike many recommendation engines, our system faces a distinctive challenge, location retrieval, that sits upstream of ranking and determines which geographic areas are queried in order to filter inventory to a candidate set. The preexisting approach employs a deep bayesian bandit based system to predict a rectangular retrieval bounds area that can be used for filtering. The purpose of this paper is to demonstrate the methodology, challenges, and impact of rearchitecting search to retrieve from the subset of most bookable high precision rectangular map cells defined by dividing the world into 25M uniform cells.
  •  
❌