❌

Reading view

Test-Time Deep Thinking to Explore Implicit Rules

arXiv:2605.24828v1 Announce Type: new Abstract: With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in environments governed by implicit rules--hidden constraints that cannot be observed directly and must be inferred through interaction. This causes agents to fall into repetitive trial-and-error loops, ultimately leading to task failure. To address this challenge, we propose Test-Time Exploration (TTExplore), a framework where a thinker component analyzes interaction history to infer these implicit rules and guide an actor. Effective exploration in this setting critically depends on the reasoning ability of the thinker. However, evaluating deep reasoning trajectories is inherently unstable and difficult, which poses a major obstacle to effective training. To overcome this issue, we introduce a novel and stable reinforcement learning pipeline. The core idea is to use accurate task-level scores as indirect rewards to bypass the difficulty of evaluating intermediate reasoning, and to retain only a single thinking node per trajectory to alleviate reward sparsity. Using this pipeline, we train a specialized 7B model, Exp-Thinker. Experiments on five text-based embodied tasks show that TTExplore equipped with Exp-Thinker improves baseline agent performance by an average of $14$-$19$ points, demonstrating the effectiveness of explicitly reasoning about implicit rules.
  •  

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundation models. However, current evaluations are limited, typically focusing on one or two cities and tasks and relying on random splits that introduce spatial leakage, leading to inflated performance and weak support for cross-location generalization and fair comparison. To address this, we propose CityRep, a unified benchmark that evaluates urban representations across data modalities, cities, and tasks using spatially structured splits. CityRep consists of three key components: (1) a spatial unit-agnostic evaluation framework that supports heterogeneous urban representations through a standardized alignment module; (2) a unified evaluation protocol using block-based spatial splits to mitigate spatial leakage and enable rigorous model comparison; and (3) an extensible multi-city, multi-task benchmark suite spanning 8 cities and 8 tasks across regression, classification, and distribution prediction. We evaluate 11 representative urban representation models. Results show that performance is highly sensitive to the split protocol, with random splits inflating scores and altering model rankings. We also observe substantial variability across cities and tasks, underscoring the need for generalization-aware evaluation. CityRep is released as a reproducible benchmark with datasets, evaluation pipelines, and diagnostic tools to facilitate fair comparison and support future research in urban representation learning towards urban foundation models.
  •  

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate over only narrow slices of that world, limiting context-sensitive reasoning and effective assistance. Existing benchmarks similarly provide only partial user state and therefore fail to capture performance in such a broad, always-on setting. To address this gap, we introduce Claw-Anything, a benchmark that expands agent context along three dimensions: long-horizon activity histories, interdependent backend services, and integrated GUI and CLI interaction across multiple devices. To instantiate this setting, we simulate months of user activity through multi-round event injection, producing complex world states and realistic noise, including irrelevant events and conflicting signals. Agents must reason over rich contextual environments while remaining robust to such noise. This expanded scope also enables the evaluation of proactive assistance, requiring agents to anticipate user needs and deliver timely recommendations. Experiments show that GPT-5.5 achieves only 34.5% pass@1, substantially below prior benchmarks, underscoring a gap between current agent capabilities and the demands of always-on personal assistance. Alongside the benchmark, we release an automated data-generation pipeline that yields 2,000 training environments and improves the base model by 23.7%, demonstrating its utility of scalable data infrastructure.
  •  

MuNet: A Mutualistic Network for Joint 3D Human Mesh Recovery and 3D Clothed Human Reconstruction from Single Images

arXiv:2605.25861v2 Announce Type: cross Abstract: 3D human mesh recovery and 3D clothed human reconstruction are inherently related, yet they have long been studied in isolation, thereby overlooking the potential gains of joint optimization. To overcome this limitation, we propose to address these two tasks within a unified framework, which allows their mutual dependencies to be effectively exploited. Building on this idea, we propose MuNet, a mutualistic network for joint 3D human mesh recovery and 3D clothed human reconstruction from single images. First, we adopt 2-manifold graphs as a unified representation for all 3D models, enabling consistent modeling across 3D human mesh recovery and clothed human reconstruction. Second, we design an end-to-end graph convolutional network that progressively deforms an initial graph into a 3D human mesh and refines it into a detailed 3D clothed human model. Third, we introduce a mutualistic mechanism that allows reciprocal interaction between the two tasks {during training}, where 3D human mesh recovery provides guidance for 3D clothed human reconstruction, and reconstruction feedback refines the 3D human mesh recovery. We extensively evaluate MuNet on six benchmark datasets for 3D human mesh recovery and 3D clothed human reconstruction, including Human3.6M, 3DPW, MPI-INF-3DHP, THuman2.0, CAPE, and RenderPeople. Experimental results demonstrate that MuNet achieves state-of-the-art performance on both tasks across all datasets. The code of MuNet is released for research purposes at https://github.com/starVisionTeam/MuNet.
  •  

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning

arXiv:2605.12906v2 Announce Type: replace-cross Abstract: Data selection during supervised fine-tuning (SFT) can critically change the behavior of large language models (LLMs). Although existing work has studied the effect of selecting data based on heuristics such as perplexity, difficulty, or length, the reported findings are often inconsistent or context-dependent. In this work, we systematically study the role of data difficulty in fine-tuning from both empirical and theoretical perspectives, and find that there is no universally optimal difficulty level; rather, its effectiveness depends on the dataset size. We show that for a fixed data budget, there exists an optimal data difficulty for SFT, and that this optimal difficulty shifts toward harder data as the data budget increases. To explain this phenomenon, we conduct controlled synthetic experiments that reveal a simple underlying mechanism: the interplay between the (in-distribution) generalization gap and the extrapolation gap. We further support this mechanism through a theoretical analysis using PAC-Bayesian generalization bounds. Overall, our results clarify how data size and difficulty jointly affect the trade-off between generalization and extrapolation in SFT, providing guidance for difficulty-based data selection under certain model and data conditions.
  •  

TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting

arXiv:2605.22365v2 Announce Type: replace-cross Abstract: Time Series Forecasting (TSF) is highly vulnerable to backdoor attacks, yet effective defenses remain underexplored due to challenges arising from data entanglement and shifts in task formulation. To fill this gap, we conduct a systematic evaluation of thirteen representative backdoor defenses across the TSF life cycle and analyze their failure modes. Our results reveal two fundamental issues: (1) data entanglement induces channel-level signal dilution, rendering sample-filtering and trigger-synthesis defenses ineffective at localizing backdoors; and (2) task-formulation shift leads to training-loss degeneration, causing poisoned and clean windows to become indistinguishable at training stages. Based on these findings, we propose a training-time backdoor defense for TSF, termed TimeGuard. Our method adopts channel-wise pool training as the core paradigm and initializes a high-confidence pool using time-aware criteria to mitigate signal dilution. Moreover, we introduce distance-regularized loss selection to progressively expand the reliable pool during training and ease loss degeneration. Extensive experiments across multiple datasets, forecasting architectures, and TSF backdoor attacks demonstrate that TimeGuard substantially improves robustness, boosting $\mathrm{MAE}_\mathrm{P}$ by $1.96\times$ over the leading baseline, while preserving clean performance within 5% $\mathrm{MAE}_\mathrm{C}$.
  •  

FGFR1 Promotes Malignant Progression in Lung Squamous Cell Carcinoma Through Activation of Wnt/beta-Catenin Signaling

Cancer Med. 2026 Apr;15(4):e71833. doi: 10.1002/cam4.71833.

ABSTRACT

OBJECTIVES: This study aims to elucidate the role of FGFR1 in activating the Wnt/β-catenin signaling pathway and the underlying mechanisms by which it promotes malignant progression in lung squamous cell carcinoma (LUSC). By integrating multi-omics analysis with functional experiments, the clinical heterogeneity of FGFR1 amplification, signaling crosstalk, and their regulatory networks governing tumor phenotypes were revealed.

METHODS: Using TCGA data (n = 490), we analyzed the relationship between FGFR1 copy number variation (CNV) and mRNA expression in LUSC, and validated the correlation with protein expression in a clinical cohort (n = 38). GSEA and single-gene GSEA were performed to identify signaling pathways associated with high FGFR1 expression. The interaction between FGFR1 and the Wnt/β-catenin pathway was investigated by immunohistochemistry, immunofluorescence, stable cell lines, Western blot, qPCR, and functional assays.

RESULTS: FGFR1 amplification correlated with increased mRNA and protein expression. The top 25% FGFR1 high-expression group enriched Wnt/β-catenin, PI3K-Akt, and cAMP pathways. Mechanistically, FGFR1 promoted β-catenin nuclear accumulation and enhanced β-catenin signaling through PKA-associated phosphorylation and Akt/GSK3β-related regulation of β-catenin stability, and these effects were attenuated by AKT inhibition. CTNNB1 knockdown significantly inhibited proliferation, migration, invasion, and tumor growth of LUSC cells.

CONCLUSIONS: Our findings indicate that FGFR1 activates Wnt/β-catenin signaling through coordinated regulation of β-catenin phosphorylation, stability, and subcellular localization, thereby promoting malignant progression in LUSC. These results provide a rationale for targeting the FGFR1-Wnt/β-catenin axis as a potential therapeutic strategy.

PMID:41998829 | DOI:10.1002/cam4.71833

  •  

Characterization and regulatory mechanism evaluation of C8orf33 in hepatocellular carcinoma through multiomics profiling

Discov Oncol. 2026 Apr 11. doi: 10.1007/s12672-026-04951-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) is a major cause of cancer-related mortality. Chromosome 8 open reading frame 33 (C8orf33) has been noted as a potential oncogenic factor in several cancers, but its biological roles and regulatory mechanism in HCC microenvironment remain unknown.

METHODS: We integrated bulk RNA sequencing, single-cell RNA sequencing (scRNA-seq), and spatial transcriptomics (ST) to characterize the expression landscape of C8orf33. We then performed C8orf33 loss-of-function studies in HCC cell lines, including in vitro phenotypic assays and subcutaneous xenografts.

RESULTS: C8orf33 was broadly overexpressed and associated with unfavorable prognosis across multiple Cancers. In HCC, higher C8orf33 aligned with advanced stage and shorter overall survival. C8orf33 knockdown reduced proliferation and migration, impaired tumorigenic capacity, and increased apoptosis. ScRNA-seq analyses identified a malignant population of Epi3 with high C8orf33 expression. Cell-cell communication analysis suggested that C8orf33-high Epi3 state was associated with an enriched MIF-CD74/CXCR4/CD44 signaling program toward macrophage populations with M2-like features. ST analyses further confirmed the colocalization of C8orf33 with malignant features in tumor cores. In Huh7 cells, C8orf33 knockdown was accompanied by reduced mRNA and protein levels of MIF and its receptor components. Consistently, xenografts derived from C8orf33-silenced cells showed lower expression of these MIF-axis components and reduced infiltration of CD163 and CD206-positive macrophages.

CONCLUSION: These results support a tumor-promoting association of C8orf33 in HCC and suggest a potential link to macrophage-associated immunomodulatory features, nominating C8orf33 as a candidate biomarker and therapeutic target.

PMID:41965457 | DOI:10.1007/s12672-026-04951-z

  •  

Characterization and regulatory mechanism evaluation of C8orf33 in hepatocellular carcinoma through multiomics profiling

Discov Oncol. 2026 Apr 11. doi: 10.1007/s12672-026-04951-z. Online ahead of print.

ABSTRACT

BACKGROUND: Hepatocellular carcinoma (HCC) is a major cause of cancer-related mortality. Chromosome 8 open reading frame 33 (C8orf33) has been noted as a potential oncogenic factor in several cancers, but its biological roles and regulatory mechanism in HCC microenvironment remain unknown.

METHODS: We integrated bulk RNA sequencing, single-cell RNA sequencing (scRNA-seq), and spatial transcriptomics (ST) to characterize the expression landscape of C8orf33. We then performed C8orf33 loss-of-function studies in HCC cell lines, including in vitro phenotypic assays and subcutaneous xenografts.

RESULTS: C8orf33 was broadly overexpressed and associated with unfavorable prognosis across multiple Cancers. In HCC, higher C8orf33 aligned with advanced stage and shorter overall survival. C8orf33 knockdown reduced proliferation and migration, impaired tumorigenic capacity, and increased apoptosis. ScRNA-seq analyses identified a malignant population of Epi3 with high C8orf33 expression. Cell-cell communication analysis suggested that C8orf33-high Epi3 state was associated with an enriched MIF-CD74/CXCR4/CD44 signaling program toward macrophage populations with M2-like features. ST analyses further confirmed the colocalization of C8orf33 with malignant features in tumor cores. In Huh7 cells, C8orf33 knockdown was accompanied by reduced mRNA and protein levels of MIF and its receptor components. Consistently, xenografts derived from C8orf33-silenced cells showed lower expression of these MIF-axis components and reduced infiltration of CD163 and CD206-positive macrophages.

CONCLUSION: These results support a tumor-promoting association of C8orf33 in HCC and suggest a potential link to macrophage-associated immunomodulatory features, nominating C8orf33 as a candidate biomarker and therapeutic target.

PMID:41965457 | DOI:10.1007/s12672-026-04951-z

  •  

Graphicalized vision-language modeling for comprehensive lung nodule analysis and risk stratification

npj Digital Medicine, Published online: 11 April 2026; doi:10.1038/s41746-026-02602-9

Graphicalized vision-language modeling for comprehensive lung nodule analysis and risk stratification
  •  

Social Media Intervention Based on the Information-Motivation-Behavioral Skills Model Promotes HIV Testing and Reduces High-Risk Behaviors Among Men Who Have Sex With Men in Resource-Limited Settings in China: Randomized Controlled Trial

Background: Social media intervention may enhance HIV prevention among men who have sex with men, but the effect of this intervention in resource-limited settings remains unclear. Objective: This randomized controlled trial evaluated whether a social media intervention grounded in the information-motivation-behavioral skills (IMB) model could be beneficial for HIV prevention among men who have sex with men in resource-limited settings. Methods: Participants were recruited in Nanning, China, between April 2023 and April 2024. Eligible participants were randomly assigned to either the social media intervention group or the routine HIV prevention services control group. Participants in the intervention group received a 3-month social media intervention, which included completing video-based tasks. Baseline surveys were conducted, followed by follow-up surveys every 3 months, for a total of 2 follow-ups. Outcomes included HIV testing uptake, high-risk behavior, AIDS-related knowledge, safe sex self-efficacy, and attitude. Results: A total of 180 eligible men who have sex with men were enrolled (90 per group). Follow-up rates were 97.8% (88/90) and 95.5% (86/90) for the intervention and control groups, respectively. At the follow-ups, the intervention group demonstrated significantly higher uptake of HIV testing, a lower proportion of participants reporting high-risk sexual behaviors, and higher condom use self-efficacy compared to the control group (all
  •  

VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing

arXiv:2604.03322v1 Announce Type: cross Abstract: Quality inspection in smart manufacturing requires identifying intrinsic material and surface properties beyond visible geometry, yet vision-only methods remain vulnerable to occlusion and reflection. We propose VitaTouch, a property-aware vision-tactile-language model for material-property inference and natural-language attribute description. VitaTouch uses modality-specific encoders and a dual Q-Former to extract language-relevant visual and tactile features, which are compressed into prefix tokens for a large language model. We align each modality with text and explicitly couple vision and touch through contrastive learning. We also construct VitaSet, a multimodal dataset with 186 objects, 52k images, and 5.1k human-verified instruction-answer pairs. VitaTouch achieves the best performance on HCT and the overall TVL benchmark, while remaining competitive on SSVTP. On VitaSet, it reaches 88.89% hardness accuracy, 75.13% roughness accuracy, and 54.81% descriptor recall; the material-description task further achieves a peak semantic similarity of 0.9009. With LoRA-based fine-tuning, VitaTouch attains 100.0%, 96.0%, and 92.0% accuracy for 2-, 3-, and 5-category defect recognition, respectively, and delivers 94.0% closed-loop recognition accuracy and 94.0% end-to-end sorting success in 100 laboratory robotic trials. More details are available at the project page: https://vitatouch.github.io/
  •  

CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks

arXiv:2604.04060v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in complex applications, their vulnerability to adversarial attacks raises urgent safety concerns, especially those evolving over multi-round interactions. Existing defenses are largely reactive and struggle to adapt as adversaries refine strategies across rounds. In this work, we propose CoopGuard , a stateful multi-round LLM defense framework based on cooperative agents that maintains and updates an internal defense state to counter evolving attacks. It employs three specialized agents (Deferring Agent, Tempting Agent, and Forensic Agent) for complementary round-level strategies, coordinated by System Agent, which conditions decisions on the evolving defense state (interaction history) and orchestrates agents over time. To evaluate evolving threats, we introduce the EMRA benchmark with 5,200 adversarial samples across 8 attack types, simulating progressively LLM multi-round attacks. Experiments show that CoopGuard reduces attack success rate by 78.9% over state-of-the-art defenses, while improving deceptive rate by 186% and reducing attack efficiency by 167.9%, offering a more comprehensive assessment of multi-round defense. These results demonstrate that CoopGuard provides robust protection for LLMs in multi-round adversarial scenarios.
  •  

Reflection of Episodes: Learning to Play Game from Expert and Self Experiences

arXiv:2502.13388v3 Announce Type: replace Abstract: StarCraft II is a complex and dynamic real-time strategy (RTS) game environment, which is very suitable for artificial intelligence and reinforcement learning research. To address the problem of Large Language Model(LLM) learning in complex environments through self-reflection, we propose a Reflection of Episodes(ROE) framework based on expert experience and self-experience. This framework first obtains key information in the game through a keyframe selection method, then makes decisions based on expert experience and self-experience. After a game is completed, it reflects on the previous experience to obtain new self-experience. Finally, in the experiment, our method beat the robot under the Very Hard difficulty in TextStarCraft II. We analyze the data of the LLM in the process of the game in detail, verified its effectiveness.
  •  

ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction

arXiv:2604.01561v1 Announce Type: cross Abstract: We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic regions, leading to unstable reconstruction and motion estimation, which often resorts to external dense motion guidance such as pre-computed optical flow to further stabilize and constrain the reconstruction of dynamic components. However, this introduces additional complexity and potential error propagation. To address these issues, ReFlow integrates a Complete Canonical Space Construction module for enhanced initialization of both static and dynamic regions, and a Separation-Based Dynamic Scene Modeling module that decouples static and dynamic components for targeted motion supervision. The core of ReFlow is a novel self-correction flow matching mechanism, consisting of Full Flow Matching to align 3D scene flow with time-varying 2D observations, and Camera Flow Matching to enforce multi-view consistency for static objects. Together, these modules enable robust and accurate dynamic scene reconstruction. Extensive experiments across diverse scenarios demonstrate that ReFlow achieves superior reconstruction quality and robustness, establishing a novel self-correction paradigm for monocular 4D reconstruction.
  •  

Reflection of Episodes: Learning to Play Game from Expert and Self Experiences

arXiv:2502.13388v2 Announce Type: replace Abstract: StarCraft II is a complex and dynamic real-time strategy (RTS) game environment, which is very suitable for artificial intelligence and reinforcement learning research. To address the problem of Large Language Model(LLM) learning in complex environments through self-reflection, we propose a Reflection of Episodes(ROE) framework based on expert experience and self-experience. This framework first obtains key information in the game through a keyframe selection method, then makes decisions based on expert experience and self-experience. After a game is completed, it reflects on the previous experience to obtain new self-experience. Finally, in the experiment, our method beat the robot under the Very Hard difficulty in TextStarCraft II. We analyze the data of the LLM in the process of the game in detail, verified its effectiveness.
  •  

Generation Is Compression: Zero-Shot Video Coding via Stochastic Rectified Flow

arXiv:2603.26571v2 Announce Type: replace-cross Abstract: Recent advances in generative modeling have enabled perceptual video compression at ultra-low bitrates, yet existing methods predominantly treat the generative model as a refinement or reconstruction module attached to a separately designed codec backbone. We propose \emph{Generative Video Codebook Codec} (GVCC), a zero-shot framework that turns a pretrained video generative model into the codec itself: the transmitted bitstream directly specifies the generative decoding trajectory, with no retraining required. To enable this, we convert the deterministic rectified-flow ODE of modern video foundation models into an equivalent SDE at inference time, unlocking per-step stochastic injection points for codebook-driven compression. Building on this unified backbone, we instantiate three complementary conditioning strategies -- \emph{Image-to-Video} (I2V) with autoregressive GOP chaining, tail latent residual correction, and adaptive atom allocation, \emph{Text-to-Video} (T2V) operating at near-zero side information as a pure generative prior, and \emph{First-Last-Frame-to-Video} (FLF2V) with boundary-sharing GOP chaining for dual-anchor temporal control. Together, these variants span a principled trade-off space between spatial fidelity, temporal coherence, and compression efficiency. Experiments on standard benchmarks show that GVCC achieves high-quality reconstruction below 0.002\,bpp while supporting flexible bitrate control through a single hyperparameter.
  •  

Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain

The development of genetically encoded fluorescent reporters, along with their corresponding knock-in mouse lines for labeling α-Syn inclusions, enables diverse applications in studying the propagation and pathological effects of α-Syn inclusions in the live brain.
  •  
❌