❌

Normal view

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

25 May 2026 at 18:00

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

The role of growth heterogeneity in solid nodular non-small cell lung cancer in clinical practice: a narrative review

25 May 2026 at 18:00

J Thorac Dis. 2026 Apr 30;18(4):417. doi: 10.21037/jtd-2025-1-2697. Epub 2026 Mar 26.

ABSTRACT

BACKGROUND AND OBJECTIVE: Lung cancer remains the leading cause of cancer related mortality worldwide, and early detection and precise stratified management are crucial for improving patient outcomes. Tumor growth kinetics, as a characterization of its proliferation and malignant differentiation, is a key decision-making factor and research hotspot in clinical practice today. This study aimed to elucidate the growth kinetics of solid nodular non-small cell lung cancer (NSCLC) as a critical determinant of early diagnosis, prognostic evaluation, and treatment strategy selection, and to address the challenge that significant heterogeneity in tumor growth poses to risk stratification and clinical decision-making.

METHODS: We conducted a retrospective search of PubMed, Embase, Web of Science, and Scopus databases, focusing on the current research status of solid nodular NSCLC, particularly in terms of molecular mechanisms, prognosis, modeling prediction, and management strategies related to its growth heterogeneity, with the aim of exploring future research directions.

KEY CONTENT AND FINDINGS: Volume doubling time (VDT) serves as a key metric for evaluating nodule dynamics. While earlier studies suggested a generally rapid growth pattern (VDT <400 days) in solid nodular NSCLC, recent evidence reveals considerable heterogeneity, with some tumors demonstrating indolent growth pattern (VDT >40-600 days). The prognosis of rapidly growing nodules is usually poor, so nodule management recommendations should be personalized based on growth dynamics and patient characteristics. Traditional radiological features, and deep learning models show promise for growth risk stratification but require large-scale external validation and refinement. Molecular and pathological studies suggest that the tumor microenvironment and immune cell infiltration may contribute to growth heterogeneity, though direct mechanistic evidence remains limited. Artificial intelligence (AI) based approaches exhibit significant potential in predicting individual tumor growth behavior.

CONCLUSIONS: Growth heterogeneity in solid nodular NSCLC carries substantial clinical significance but remains insufficiently studied. Future research should prioritize imaging based modeling to predict individualized growth dynamics. Integrating multi-omics analyses may help elucidate the molecular factors underlying growth heterogeneity. AI driven risk stratification based on large-scale multi center sequence data can achieve truly personalized and growth oriented management strategies.

PMID:42182806 | PMC:PMC13190150 | DOI:10.21037/jtd-2025-1-2697

Nonlinear atomic tunnelling boosted by bright squeezed vacuum

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10485-9

Bright squeezed vacuum light boosts nonlinear atomic tunnelling ionization more than 20-fold compared with coherent light, enabling quantum control of strong-field processes without increasing classical intensity.

Targeting cysteinyl leukotriene receptor 1 reprograms tumor-promoting myelopoiesis and overcomes immune checkpoint therapy resistance

Nature Cancer, Published online: 19 May 2026; doi:10.1038/s43018-026-01174-7

Tang et al. identify cysteinyl leukotriene receptor 1 (CysLTR1) as a critical regulator of tumor-induced myelopoiesis, suggesting CysLTR1 targeting to sensitize tumors to immune checkpoint blockade.

Engineered immunosuppressive dendritic cells protect against cardiac remodelling

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10346-5

Lesion-targeted immune modulation is a feasible strategy to control cardiac fibrosis, and engineered dendritic cells are a promising therapeutic platform for treating cardiac remodelling and heart failure.

Superconductivity and electronic structures of nickelate thin film superstructures

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10352-7

Engineered Ruddlesden–Popper nickelate superstructures show that specific Fermi surface features enable ambient-pressure superconductivity, linking structural configuration, electronic structure and superconducting behaviour. .

Can LLMs Learn to Reason Robustly under Noisy Supervision?

arXiv:2604.03993v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels due to expert scarcity remains critically underexplored. In this work, we take the first step toward a systematic analysis of noisy label mechanisms in RLVR. In contrast to supervised classification, most RLVR algorithms incorporate a rollout-based condition: a label's influence on training is contingent on whether the current policy can generate rollouts that realize it, a property that naturally extends to noisy labels. Based on this observation, we distinguish two types of noise: inactive noisy labels, which reduce data efficiency, and active noisy labels, which are reinforced and risk skewing the model toward incorrect distributions. From experiments on training with noisy samples, we identify an Early Correctness Coherence phenomenon: although noisy samples begin to lag behind in later stages, accuracy on both clean and noisy samples increases similarly in early training. Motivated by this dynamic, we propose Online Label Refinement (OLR), which progressively corrects potentially noisy labels with majority-voted answers when two conditions hold: a positive slope in the majority answer's rollout pass rate and stable historical consistency across updates, enabling gradual self-correction as the policy improves. We evaluate OLR on six in-distribution mathematical reasoning benchmarks (AIME24/25, AMC, MATH-500, Minerva, and Olympiad) and three out-of-distribution tasks (ARC-c, GPQA-diamond, and MMLU-pro). Across noise ratios from 0.1 to 0.9, OLR consistently improves robustness under both inactive and active noisy-label settings, achieving average gains of 3.6% to 3.9% on in-distribution benchmarks and 3.3% to 4.6% on out-of-distribution evaluations.

Vero: An Open RL Recipe for General Visual Reasoning

arXiv:2604.04917v2 Announce Type: cross Abstract: What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) show such broad visual reasoning is within reach, but the recipe behind them remains unclear, locked behind proprietary reinforcement learning (RL) pipelines with non-public data. We introduce Vero, a family of fully open VLMs that matches or exceeds existing open-weight models across diverse visual reasoning tasks. We scale RL data and rewards across six broad task categories, constructing Vero-600K, a 600K-sample dataset from 59 datasets, and designing task-routed rewards that handle heterogeneous answer formats. Vero achieves state-of-the-art performance, improving over four base models by 3.6-5.3 points on average across VeroEval, our suite of 30 challenging benchmarks. Starting from Qwen3-VL-8B-Instruct, Vero outperforms Qwen3-VL-8B-Thinking on 23 of 30 benchmarks without additional proprietary thinking data. When trained from the same base model, Vero-600K exceeds existing RL datasets across task categories. Systematic ablations reveal that different task categories elicit qualitatively distinct reasoning patterns that transfer poorly in isolation, suggesting that broad data coverage is the primary driver of strong RL scaling. All data, code, and models are released.

RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning

arXiv:2504.18594v3 Announce Type: replace-cross Abstract: Compared to untargeted attacks, targeted transfer-based attack is still suffering from much lower Attack Success Rates (ASRs), although significant improvements have been achieved by kinds of methods, such as diversifying input, stabilizing the gradient, and re-training surrogate models. In this paper, we find that adversarial examples generated by existing methods rely heavily on a small subset of surrogate model parameters, which in turn limits their transferability to unseen target models. Inspired by this, we propose the Random Parameter Pruning Attack (RaPA), which introduces parameter-level randomization during the attack process. At each optimization step, RaPA randomly prunes model parameters to generate diverse yet semantically consistent surrogate variants.We show this parameter-level randomization is equivalent to adding an importance-equalization regularizer, thereby alleviating the over-reliance issue. Extensive experiments across both CNN and Transformer architectures demonstrate that RaPA substantially enhances transferability. In the challenging case of transferring from CNN-based to Transformer-based models, RaPA achieves up to 11.7% higher average ASRs than state-of-the-art baselines(with 33.3% ASRs), while being training-free, cross-architecture efficient, and easily integrated into existing attack frameworks. Code is available in https://github.com/molarsu/RaPA.

NSUN2/ALYREF-mediated RNA m5c modification promotes anoikis resistance of prostate cancer through activating autophagy

Oncogene, Published online: 07 April 2026; doi:10.1038/s41388-026-03762-4

NSUN2/ALYREF-mediated RNA m5c modification promotes anoikis resistance of prostate cancer through activating autophagy

Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain

The development of genetically encoded fluorescent reporters, along with their corresponding knock-in mouse lines for labeling α-Syn inclusions, enables diverse applications in studying the propagation and pathological effects of α-Syn inclusions in the live brain.

LLM-Meta-SR: In-Context Learning for Evolving Selection Operators in Symbolic Regression

arXiv:2505.18602v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have revolutionized algorithm development, yet their application in symbolic regression, where algorithms automatically discover symbolic expressions from data, remains limited. In this paper, we propose a meta-learning framework that enables LLMs to automatically design selection operators for evolutionary symbolic regression algorithms. We first identify two key limitations in existing LLM-based algorithm evolution techniques: lack of semantic guidance and code bloat. The absence of semantic awareness can lead to ineffective exchange of useful code components, while bloat results in unnecessarily complex components; both can hinder evolutionary learning progress or reduce the interpretability of the designed algorithm. To address these issues, we enhance the LLM-based evolution framework for meta-symbolic regression with two key innovations: a complementary, semantics-aware selection operator and bloat control. Additionally, we embed domain knowledge into the prompt, enabling the LLM to generate more effective and contextually relevant selection operators. Our experimental results on symbolic regression benchmarks show that LLMs can devise selection operators that outperform nine expert-designed baselines, achieving state-of-the-art performance. Moreover, the evolved operator can further improve a state-of-the-art symbolic regression algorithm, achieving the best performance among 28 symbolic regression and other machine learning algorithms across 116 regression datasets. This demonstrates that LLMs can exceed expert-level algorithm design for symbolic regression.

OpenSage: Self-programming Agent Generation Engine

arXiv:2602.16891v2 Announce Type: replace Abstract: Agent development kits (ADKs) provide effective platforms and tooling for constructing agents, and their designs are critical to the constructed agents' performance, especially the functionality for agent topology, tools, and memory. However, current ADKs either lack sufficient functional support or rely on humans to manually design these components, limiting agents' generalizability and overall performance. We propose OpenSage, the first ADK that enables LLMs to automatically create agents with self-generated topology and toolsets while providing comprehensive and structured memory support. OpenSage offers effective functionality for agents to create and manage their own sub-agents and toolkits. It also features a hierarchical, graph-based memory system for efficient management and a specialized toolkit tailored to software engineering tasks. Extensive experiments across three state-of-the-art benchmarks with various backbone models demonstrate the advantages of OpenSage over existing ADKs. We also conduct rigorous ablation studies to demonstrate the effectiveness of our design for each component. We believe OpenSage can pave the way for the next generation of agent development, shifting the focus from human-centered to AI-centered paradigms.

OTESGN: Optimal Transport-Enhanced Syntactic-Semantic Graph Networks for Aspect-Based Sentiment Analysis

arXiv:2509.08612v3 Announce Type: replace-cross Abstract: Aspect-based sentiment analysis (ABSA) aims to identify aspect terms and determine their sentiment polarity. While dependency trees combined with contextual semantics provide structural cues, existing approaches often rely on dot-product similarity and fixed graphs, which limit their ability to capture nonlinear associations and adapt to noisy contexts. To address these limitations, we propose the Optimal Transport-Enhanced Syntactic-Semantic Graph Network (OTESGN), a model that jointly integrates structural and distributional signals. Specifically, a Syntactic Graph-Aware Attention module models global dependencies with syntax-guided masking, while a Semantic Optimal Transport Attention module formulates aspect-opinion association as a distribution matching problem solved via the Sinkhorn algorithm. An Adaptive Attention Fusion mechanism balances heterogeneous features, and contrastive regularization enhances robustness. Extensive experiments on three benchmark datasets (Rest14, Laptop14, and Twitter) demonstrate that OTESGN delivers state-of-the-art performance. Notably, it surpasses competitive baselines by up to +1.30 Macro-F1 on Laptop14 and +1.01 on Twitter. Ablation studies and visualization analyses further highlight OTESGN's ability to capture fine-grained sentiment associations and suppress noise from irrelevant context.

MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models

arXiv:2511.20629v4 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) with reward models has advanced alignment of generative models to human aesthetic and perceptual preferences. However, jointly optimizing multiple rewards often incurs an alignment tax, improving one dimension while degrading others. To address this, we introduce two complementary methods: MapReduce LoRA and Reward-aware Token Embedding (RaTE). MapReduce LoRA trains preference-specific LoRA experts in parallel and iteratively merges them to refine a shared base model; RaTE learns reward-specific token embeddings that compose at inference for flexible preference control. Experiments on Text-to-Image generation (Stable Diffusion 3.5 Medium and FLUX.1-dev) show improvements of 36.1%, 4.6%, and 55.7%, and 32.7%, 4.3%, and 67.1% on GenEval, PickScore, and OCR, respectively. On Text-to-Video generation (HunyuanVideo), visual and motion quality improve by 48.1% and 90.0%, respectively. On the language task, Helpful Assistant, with Llama-2 7B, helpful and harmless improve by 43.4% and 136.7%, respectively. Our framework sets a new state-of-the-art multi-preference alignment recipe across modalities.
❌