❌

Normal view

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.

Action with Visual Primitives

arXiv:2605.22183v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions and visual observations to actions in a single forward pass. While conceptually simple, this formulation entangles instruction comprehension, spatial scene understanding, and motor control within a single learning objective. As a result, the action expert must implicitly relearn cognitive and perceptual capabilities already present in the pretrained VLM, which can limit both learning efficiency and generalization. We introduce AVP (Action with Visual Primitives), an end-to-end architecture that implements this visual-primitive-centric interface: the VLM infers the next-stage target and emits visual-primitive tokens that condition a flow-matching action expert, with supervision derived from end-effector kinematics. Real-robot experiments on general pick-and-place tasks show that AVP improves the success rate by 27.61% over pi_0.5 and outperforms other recent methods, with consistent gains in data efficiency, spatial-compositional generalization, and object-level transfer.

Predicting Neuromodulation Outcome for Parkinson's Disease with Generative Virtual Brain Model

arXiv:2603.29176v1 Announce Type: new Abstract: Parkinson's disease (PD) affects over ten million people worldwide. Although temporal interference (TI) and deep brain stimulation (DBS) are promising therapies, inter-individual variability limits empirical treatment selection, increasing non-negligible surgical risk and cost. Previous explorations either resort to limited statistical biomarkers that are insufficient to characterize variability, or employ AI-driven methods which is prone to overfitting and opacity. We bridge this gap with a pretraining-finetuning framework to predict outcomes directly from resting-state fMRI. Critically, a generative virtual brain foundation model, pretrained on a collective dataset (2707 subjects, 5621 sessions) to capture universal disorder patterns, was finetuned on PD cohorts receiving TI (n=51) or DBS (n=55) to yield individualized virtual brains with high fidelity to empirical functional connectivity (r=0.935). By constructing counterfactual estimations between pathological and healthy neural states within these personalized models, we predicted clinical responses (TI: AUPR=0.853; DBS: AUPR=0.915), substantially outperforming baselines. External and prospective validations (n=14, n=11) highlight the feasibility of clinical translation. Moreover, our framework provides state-dependent regional patterns linked to response, offering hypothesis-generating mechanistic insights.

Associations Between Short-Video Platform Use and Health Across Health Distribution and Usage Behaviors in China: Cross-Sectional Questionnaire Study

Background: Short-video platforms, characterized by algorithmic curation and passive consumption, have emerged as dominant components of digital life. However, the associations between short-video platform use and health across different groups and usage behaviors remain understudied. Objective: This study investigates associations between short-video platform use and health, examining whether these relationships vary across health status, usage behaviors, and socioeconomic status. Methods: A cross-sectional study was conducted using multistage stratified sampling across eastern, central, and western China from July to September 2024. The inclusion criteria were age 18 years or older, ability to communicate effectively, and no cognitive disorders or mental disturbance. Of 7725 participants enrolled, 46.96% (n=3628) were male, and the average age was 65.49 (SD 8.39) years. The data were collected via face-to-face interviews using a structured questionnaire. Self-rated health and relative health deprivation (Kakwani index) were used to measure health. Quantile regression explored associations between whether using short-video platform and health varies across the health distribution, while linear regression examined associations of years, frequency, daily duration, and purpose diversity of short-video platform use with health. Moderating effect analysis explored the role of socioeconomic status in the relationship between the daily duration of use and health. Results: Coefficients were tested using 2-tailed tests, and statistical significance was defined as a 2-sided value less than .05. Quantile regression revealed heterogeneous associations. Compared to nonusers, short-video platform users had better self-rated health at the 70th to 90th quantiles and lower relative health deprivation at the 10th to 30th quantiles. However, the users at the 10th quantile of self-rated health had worse self-rated health (=−2.224, 95% CI −3.835 to −0.613). Longer engagement (≥3 y) correlated with lower relative health deprivation (=1.970, 95% CI 0.308-3.632), while daily use of 1‐4 hours was associated with poorer self-rated health (=−3.385, 95% CI −4.872 to −1.898; =−3.038, 95% CI −5.054 to −1.022) and higher relative health deprivation (=0.035, 95% CI 0.021-0.050;
❌