❌

Reading view

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

arXiv:2605.25488v1 Announce Type: cross Abstract: Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, most existing approaches rely on a single static reference image to condition the entire video generation process at inference stage. This static conditioning paradigm often creates a mismatch between fixed identity features and dynamically evolving facial motion, leading to identity drift, temporal inconsistency, and degraded perceptual quality. We introduce Test-Time Self-Adaptive Conditioning (TT-SAC), a parameter-free inference framework that enables pretrained talking-head generators to adapt their conditioning representations during inference without retraining, gradient updates, or additional supervision. Instead of treating the reference portrait as immutable, TT-SAC composes the generator with its encoder in a feedback loop: the generator's own outputs are re-encoded to construct a refined conditioning representation that better aligns with the temporal dynamics of the synthesized sequence. A single adaptation step approximates a self-consistent equilibrium of the generative process, stabilizing identity and motion across time. We further provide theoretical analysis showing that test-time conditioning adaptation reduces feature variance and improves generative stability under mild Lipschitz assumptions, while exhibiting a principled bias-variance tradeoff that governs the optimal strength of adaptation. Extensive experiments on state-of-the-art talking-head generators and benchmark datasets demonstrate consistent improvements in lip-sync accuracy, temporal coherence, identity preservation, and perceptual fidelity. TT-SAC offers a model-agnostic and training-free strategy for enhancing generative video models, establishing test-time conditioning adaptation as an effective mechanism for stabilizing audio-driven portrait animation.
  •  

Dynamic Dual-Granularity Skill Bank for Agentic RL

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often lack principled mechanisms for maintaining an evolving skill memory. We propose D2Skill, a dynamic dual-granularity skill bank for agentic RL that organizes reusable experience into task skills for high-level guidance and step skills for fine-grained decision support and error correction. D2Skill jointly trains the policy and skill bank through paired baseline and skill-injected rollouts under the same policy, using their performance gap to derive hindsight utility signals for both skill updating and policy optimization. Built entirely from training-time experience, the skill bank is continuously expanded through reflection and maintained with utility-aware retrieval and pruning. Experiments on ALFWorld, WebShop, and Search-Augmented QA tasks show that D2Skill substantially improves performance over skill-free baselines across models of different scales. Further ablations and analyses show that both dual-granularity skill modeling and dynamic skill maintenance are critical to these gains, while the learned skills exhibit higher utility, transfer across evaluation settings, and introduce only modest training overhead.
  •  

TreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding

arXiv:2604.03309v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and capture whole-part relationships in complex scenes. Moreover, dense pairwise comparisons and inconsistent hierarchical labels from 2D priors hinder feature learning, resulting in suboptimal segmentation. To address these limitations, we introduce TreeGaussian, a tree-guided cascaded contrastive learning framework that explicitly models hierarchical semantic relationships and reduces redundancy in contrastive supervision. By constructing a multi-level object tree, TreeGaussian enables structured learning across object-part hierarchies. In addition, we propose a two-stage cascaded contrastive learning strategy that progressively refines feature representations from global to local, mitigating saturation and stabilizing training. A Consistent Segmentation Detection (CSD) mechanism and a graph-based denoising module are further introduced to align segmentation modes across views while suppressing unstable Gaussian points, enhancing segmentation consistency and quality. Extensive experiments, including open-vocabulary 3D object selection, 3D point cloud understanding, and ablation studies, demonstrate the effectiveness and robustness of our approach.
  •  

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos

arXiv:2512.03666v2 Announce Type: replace-cross Abstract: A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG). Despite recent progress, existing STVG studies remain largely confined to object-centric and descriptive instructions, neglecting the task-oriented reasoning that is crucial for embodied agents to accomplish goal-directed interactions. To bridge this gap, we introduce \textbf{ToG-Bench}, the first task-oriented spatio-temporal video grounding benchmark for egocentric videos. ToG-Bench is characterized by three key features: (1) \textbf{Task-oriented Grounding}, which requires identifying and localizing objects based on intended tasks rather than straightforward descriptions; (2) \textbf{Explicit-Implicit Dual Grounding}, where target objects can be either explicitly mentioned or implicitly inferred by contextual reasoning; (3) \textbf{One-to-Many Grounding}, where a single instruction may correspond to multiple objects involved in task execution. Built upon videos sourced from ScanNet, ToG-Bench comprises 100 annotated clips with 2,704 task-oriented grounding instructions, constructed via a semi-automated pipeline that combines foundation model annotation and human refinement. In addition, we introduce a set of task-level evaluation metrics tailored for multi-object and explicit-implicit object grounding, and systematically benchmark seven state-of-the-art MLLMs. Extensive experiments reveal the intrinsic challenges of task-oriented STVG and substantial performance gaps across explicit-implicit and multi-object grounding, highlighting the difficulty of bridging perception and interaction in embodied scenarios. Data and code will be released at: \href{https://github.com/qaxuDev/ToG-Bench}{https://github.com/qaxuDev/ToG-Bench}..
  •  

Predicting Neuromodulation Outcome for Parkinson's Disease with Generative Virtual Brain Model

arXiv:2603.29176v1 Announce Type: new Abstract: Parkinson's disease (PD) affects over ten million people worldwide. Although temporal interference (TI) and deep brain stimulation (DBS) are promising therapies, inter-individual variability limits empirical treatment selection, increasing non-negligible surgical risk and cost. Previous explorations either resort to limited statistical biomarkers that are insufficient to characterize variability, or employ AI-driven methods which is prone to overfitting and opacity. We bridge this gap with a pretraining-finetuning framework to predict outcomes directly from resting-state fMRI. Critically, a generative virtual brain foundation model, pretrained on a collective dataset (2707 subjects, 5621 sessions) to capture universal disorder patterns, was finetuned on PD cohorts receiving TI (n=51) or DBS (n=55) to yield individualized virtual brains with high fidelity to empirical functional connectivity (r=0.935). By constructing counterfactual estimations between pathological and healthy neural states within these personalized models, we predicted clinical responses (TI: AUPR=0.853; DBS: AUPR=0.915), substantially outperforming baselines. External and prospective validations (n=14, n=11) highlight the feasibility of clinical translation. Moreover, our framework provides state-dependent regional patterns linked to response, offering hypothesis-generating mechanistic insights.
  •  

Robust transcriptomic hallmarks targeting intratumor heterogeneity in intrahepatic cholangiocarcinoma

Cell Rep Med. 2026 Mar 30:102708. doi: 10.1016/j.xcrm.2026.102708. Online ahead of print.

ABSTRACT

Intratumor heterogeneity (ITH) undermines transcriptome-based stratification in intrahepatic cholangiocarcinoma (iCCA). Here, we integrate multi-omics data from multi-region, single-region, and single-cell RNA sequencing cohorts to systematically characterize gene expression ITH. We uncover that immune and stromal heterogeneity are primary drivers of ITH, leading to misclassification of a median 27.8% of tumors by existing subtyping systems. To overcome this, we identify a low-intratumor-heterogeneity/high-intertumor-variability (LIHV) gene set and develop an ITH-insensitive classification system defining five subgroups: inflammatory (SI), metabolic (SII), atypical (SIII-1), immune-silent (SIII-2), and neurodegenerative (SIII-3). These subgroups exhibit distinct clinical outcomes, molecular features, immune landscapes, and therapeutic vulnerabilities. GPRC5A and VTCN1 serve as robust immunohistochemical biomarkers for SI and SIII tumors, while serum CEA and CA19-9 identify inflammatory iCCA. Therapeutically, HSP90 inhibition synergizes with anti-PD1 in inflammatory iCCA, whereas combined anti-PD1 and anti-TIM3 suppresses neurodegenerative iCCA. Collectively, our study provides a robust molecular framework and actionable therapeutic strategies for iCCA.

PMID:41916296 | DOI:10.1016/j.xcrm.2026.102708

  •  

Dynamic Targetable Extracellular Vesicle Surface Proteins Monitor Depth of Response to CAR T Therapy

Res Sq [Preprint]. 2026 Mar 18:rs.3.rs-8913641. doi: 10.21203/rs.3.rs-8913641/v1.

ABSTRACT

Extracellular vesicles (EVs) represent a promising liquid biopsy platform in multiple myeloma (MM). We developed an MM EV Surface Protein Assay to quantify and dynamically monitor four MM EV subpopulations defined by targetable MM surface proteins (BCMA, CD38, GPRC5D, and CD319) across 336 serial blood samples from 45 relapsed/refractory MM (RRMM) patients treated with anti-BCMA chimeric antigen receptor (CAR) T-cell therapy. All four MM EV subpopulations significantly decreased in 43 patients with initial response, while BCMA+, GPRC5D+, and CD319+ MM EVs increased in 19 patients with progression, and antigen escape was detected by BCMA+ MM EVs. MM EV subpopulations differentiated minimal residual disease (MRD) status and complemented MRD for detecting early relapse before clinical progression. Notably, CD319+ MM EVs were early predictors of progression-free and overall survival in MRD-negative patients. This assay enables noninvasive monitoring of deep response, progression, and antigen escape, and stratifies survival in MRD-negative patients with RRMM.

PMID:41890853 | PMC:PMC13015583 | DOI:10.21203/rs.3.rs-8913641/v1

  •  

Unraveling the Link Between Azathioprine and Acute Pancreatitis: Integrating Network Toxicology, Machine Learning, and Mendelian Randomization

CPT Pharmacometrics Syst Pharmacol. 2026 Mar;15(3):e70178. doi: 10.1002/psp4.70178.

ABSTRACT

Azathioprine (AZA), a widely used immunosuppressant, can induce acute pancreatitis (AP), yet the underlying molecular mechanisms remain unclear. This study employed an integrative multiomics strategy-combining network toxicology, machine learning, Mendelian randomization (MR), and molecular docking-to elucidate the biological basis of AZA-induced AP. AZA-associated genes were first identified through bioinformatics databases and analyzed using protein-protein interaction networks and GO/KEGG functional enrichment. Least absolute shrinkage and selection operator (LASSO) regression and support vector machine recursive feature elimination (SVM-RFE) were applied to prioritize key differentially expressed genes for diagnostic modeling. MR was then used to examine potential causal links between gene expression and AP risk, followed by molecular docking to assess AZA-protein interactions. Sixty-eight candidate genes related to AZA-induced AP were identified. Enrichment analyses indicated involvement in lipid metabolic regulation, inflammatory pathways, and energy homeostasis. Machine learning highlighted seven key genes-CES1, CTSK, JAK1, NR3C2, PLIN5, WEE1, and RORA-as central to AP development. MR analysis further demonstrated that decreased expression of CES1 and CTSK may mediate AZA-related AP susceptibility. Docking simulations revealed strong, specific binding between AZA and both CES1 and CTSK. Overall, this study identifies CES1 and CTSK as genetically protective factors and mechanistic mediators in AZA-triggered AP. These findings offer new molecular insights into the genomic and biochemical pathways underlying this adverse drug reaction.

PMID:41832938 | DOI:10.1002/psp4.70178

  •  

LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma

Oncogene, Published online: 14 March 2026; doi:10.1038/s41388-026-03714-y

LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma
  •  

CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

arXiv:2601.09923v2 Announce Type: replace Abstract: AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior to steal credentials or cause financial loss. The only known robust defense is architectural isolation that strictly separates trusted task planning from untrusted environment observations. However, applying this design to Computer Use Agents (CUAs) -- systems that automate tasks by viewing screens and executing actions -- presents a fundamental challenge: current agents require continuous observation of UI state to determine each action, conflicting with the isolation required for security. We resolve this tension by demonstrating that UI workflows, while dynamic, are structurally predictable. We introduce Single-Shot Planning for CUAs, where a trusted planner generates a complete execution graph with conditional branches before any observation of potentially malicious content, providing provable control flow integrity guarantees against arbitrary instruction injections. Although this architectural isolation successfully prevents instruction injections, we show that additional measures are needed to prevent Branch Steering attacks, which manipulate UI elements to trigger unintended valid paths within the plan. We evaluate our design on OSWorld, and retain up to 57% of the performance of frontier models while improving performance for smaller open-source models by up to 19%, demonstrating that rigorous security and utility can coexist in CUAs.
  •  

Unified Medical Image Segmentation with State Space Modeling Snake

arXiv:2507.12760v2 Announce Type: replace-cross Abstract: Unified Medical Image Segmentation (UMIS) is critical for comprehensive anatomical assessment but faces challenges due to multi-scale structural heterogeneity. Conventional pixel-based approaches, lacking object-level anatomical insight and inter-organ relational modeling, struggle with morphological complexity and feature conflicts, limiting their efficacy in UMIS. We propose Mamba Snake, a novel deep snake framework enhanced by state space modeling for UMIS. Mamba Snake frames multi-contour evolution as a hierarchical state space atlas, effectively modeling macroscopic inter-organ topological relationships and microscopic contour refinements. We introduce a snake-specific vision state space module, the Mamba Evolution Block (MEB), which leverages effective spatiotemporal information aggregation for adaptive refinement of complex morphologies. Energy map shape priors further ensure robust long-range contour evolution in heterogeneous data. Additionally, a dual-classification synergy mechanism is incorporated to concurrently optimize detection and segmentation, mitigating under-segmentation of microstructures in UMIS. Extensive evaluations across five clinical datasets reveal Mamba Snake's superior performance, with an average Dice improvement of 3\% over state-of-the-art methods.
  •  

MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning

arXiv:2603.03379v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly used for long-duration tasks, maintaining effective long-term memory has become a critical challenge. Current methods often face a trade-off between cost and accuracy. Simple storage methods often fail to retrieve relevant information, while complex indexing methods (such as memory graphs) require heavy computation and can cause information loss. Furthermore, relying on the working LLM to process all memories is computationally expensive and slow. To address these limitations, we propose MemSifter, a novel framework that offloads the memory retrieval process to a small-scale proxy model. Instead of increasing the burden on the primary working LLM, MemSifter uses a smaller model to reason about the task before retrieving the necessary information. This approach requires no heavy computation during the indexing phase and adds minimal overhead during inference. To optimize the proxy model, we introduce a memory-specific Reinforcement Learning (RL) training paradigm. We design a task-outcome-oriented reward based on the working LLM's actual performance in completing the task. The reward measures the actual contribution of retrieved memories by mutiple interactions with the working LLM, and discriminates retrieved rankings by stepped decreasing contributions. Additionally, we employ training techniques such as Curriculum Learning and Model Merging to improve performance. We evaluated MemSifter on eight LLM memory benchmarks, including Deep Research tasks. The results demonstrate that our method meets or exceeds the performance of existing state-of-the-art approaches in both retrieval accuracy and final task completion. MemSifter offers an efficient and scalable solution for long-term LLM memory. We have open-sourced the model weights, code, and training data to support further research.
  •  

Chain of World: World Model Thinking in Latent Motion

arXiv:2603.03195v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.
  •  

FineRef: Fine-Grained Error Reflection and Correction for Long-Form Generation with Citations

arXiv:2602.18437v1 Announce Type: cross Abstract: Generating with citations is crucial for trustworthy Large Language Models (LLMs), yet even advanced LLMs often produce mismatched or irrelevant citations. Existing methods over-optimize citation fidelity while overlooking relevance to the user query, which degrades answer quality and robustness in real-world settings with noisy or irrelevant retrieved content. Moreover, the prevailing single-pass paradigm struggles to deliver optimal answers in long-form generation that requiring multiple citations. To address these limitations, we propose FineRef, a framework based on Fine-grained error Reflection, which explicitly teaches the model to self-identify and correct two key citation errors, mismatch and irrelevance, on a per-citation basis. FineRef follows a two-stage training strategy. The first stage instills an "attempt-reflect-correct" behavioral pattern via supervised fine-tuning, using fine-grained and controllable reflection data constructed by specialized lightweight models. An online self-reflective bootstrapping strategy is designed to improve generalization by iteratively enriching training data with verified, self-improving examples. To further enhance the self-reflection and correction capability, the second stage applies process-level reinforcement learning with a multi-dimensional reward scheme that promotes reflection accuracy, answer quality, and correction gain. Experiments on the ALCE benchmark demonstrate that FineRef significantly improves both citation performance and answer accuracy. Our 7B model outperforms GPT-4 by up to 18% in Citation F1 and 4% in EM Recall, while also surpassing the state-of-the-art model across key evaluation metrics. FineRef also exhibits strong generalization and robustness in domain transfer settings and noisy retrieval scenarios.
  •  

Are We Measuring Oversmoothing in Graph Neural Networks Correctly?

arXiv:2502.04591v4 Announce Type: replace-cross Abstract: Oversmoothing is a fundamental challenge in graph neural networks (GNNs): as the number of layers increases, node embeddings become increasingly similar, and model performance drops sharply. Traditionally, oversmoothing has been quantified using metrics that measure the similarity of neighbouring node features, such as the Dirichlet energy. We argue that these metrics have critical limitations and fail to reliably capture oversmoothing in realistic scenarios. For instance, they provide meaningful insights only for very deep networks, while typical GNNs show a performance drop already with as few as 10 layers. As an alternative, we propose measuring oversmoothing by examining the numerical or effective rank of the feature representations. We provide extensive numerical evaluation across diverse graph architectures and datasets to show that rank-based metrics consistently capture oversmoothing, whereas energy-based metrics often fail. Notably, we reveal that drops in the rank align closely with performance degradation, even in scenarios where energy metrics remain unchanged. Along with the experimental evaluation, we provide theoretical support for this approach, clarifying why Dirichlet-like measures may fail to capture performance drop and proving that the numerical rank of feature representations collapses to one for a broad family of GNN architectures.
  •  
❌