❌

Reading view

ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

arXiv:2604.23853v2 Announce Type: replace Abstract: Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step cost, a pipeline cannot distinguish adding a missing step to fix a bug from removing an expensive step that never affected the outcome. We use the cost-attribution gap to ask whether the rule types inside a distilled skill transfer the same way to new tasks. ClawTrace records cost-attributed agent traces and compiles each session into a TraceCard; CostCraft reads TraceCards and writes three kinds of skill patches: preserve, prune, and repair. We find a pattern aggregate metrics hide. On 30 held-out SpreadsheetBench tasks across two seeds, removing prune patches roughly tripled the quality-regression count without lowering median cost. Across the full 84-task SkillsBench transfer, CostCraft saves no aggregate cost. All three quality regressions trace to the preserve lane, and both quality wins trace to the prune lane: prune patches act as quality guardrails while preserve patches drive regressions. We argue that reusable agent skills should be evaluated at the rule-type level, not as monolithic instruction packages. To support this, we release ClawTrace, the TraceCard schema, and the full set of typed skills.
  •  

A Gamified Mobile Health Intervention to Promote Physical Activity, Executive Function, and Mental Health in College Students: Randomized Controlled Trial

Background: College students commonly experience suboptimal health conditions, including insufficient physical activity (PA), excessive body weight, and declining physical fitness. Traditional interventions face low adherence, while gamified mobile health (mHealth) programs may improve engagement and outcomes. Objective: This study aimed to evaluate the feasibility and effectiveness of a novel gamified, incentive-based mHealth intervention on primary outcomes (PA and adherence) and secondary outcomes (physical fitness, body composition, executive function [EF], and mental health). Methods: A 2-arm parallel-group randomized controlled trial (RCT) was conducted in 2025 at Yantai University with 160 college students (18‐25 years; BMI 18.5‐30.0) who were randomized 1:1 (computer-generated, sex-stratified blocks of 4; concealed allocation) to the intervention group (IG) or control group (CG; n=80 each); major exclusions were contraindications to exercise, severe physical/mental illness, recent PA interventions, or psychotropic medication use. Both used the same fitness watch–app system and identical PA targets (β‰₯150 min moderate-to-vigorous physical activity [MVPA] per week or β‰₯900 metabolic equivalent-minutes [MET-min] per week); IG additionally received team-based gamification (competition, points/leaderboards, feedback, and rewards), while CG received monitoring only. PA and adherence were monitored throughout the 8-week intervention; other outcomes were assessed at baseline and 8 weeks (fitness, body composition, EF, and mental health). Open-label with blinded outcome assessors/analysts; intention-to-treat (ITT) with multiple imputation. Results: At 8 weeks, data were available for 154 participants (IG 78; CG 76); all 160 were analyzed per ITT. Compared to the CG, the IG demonstrated significantly higher mean levels in all primary PA outcomes over 8 weeks (daily steps: mean 10,356, SD 1245 versus 8242, SD 1087; Ξ”=2114; =1.81, 95% CI 1.44‐2.18;
  •  
❌