❌

Normal view

MCLR: Improving Conditional Modeling in Visual Generative Models via Inter-Class Likelihood-Ratio Maximization and Establishing the Equivalence between Classifier-Free Guidance and Alignment Objectives

arXiv:2603.22364v1 Announce Type: cross Abstract: Diffusion models have achieved state-of-the-art performance in generative modeling, but their success often relies heavily on classifier-free guidance (CFG), an inference-time heuristic that modifies the sampling trajectory. From a theoretical perspective, diffusion models trained with standard denoising score matching (DSM) are expected to recover the target data distribution, raising the question of why inference-time guidance is necessary in practice. In this work, we ask whether the DSM training objective can be modified in a principled manner such that standard reverse-time sampling, without inference-time guidance, yields effects comparable to CFG. We identify insufficient inter-class separation as a key limitation of standard diffusion models. To address this, we propose MCLR, a principled alignment objective that explicitly maximizes inter-class likelihood-ratios during training. Models fine-tuned with MCLR exhibit CFG-like improvements under standard sampling, achieving comparable qualitative and quantitative gains without requiring inference-time guidance. Beyond empirical benefits, we provide a theoretical result showing that the CFG-guided score is exactly the optimal solution to a weighted MCLR objective. This establishes a formal equivalence between classifier-free guidance and alignment-based objectives, offering a mechanistic interpretation of CFG.

Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs

arXiv:2511.05919v3 Announce Type: replace-cross Abstract: LLMs are now an integral part of information retrieval. As such, their role as question answering chatbots raises significant concerns due to their shown vulnerability to adversarial man-in-the-middle (MitM) attacks. Here, we propose the first principled attack evaluation on LLM factual memory under prompt injection via Xmera, our novel, theory-grounded MitM framework. By perturbing the input given to "victim" LLMs in three closed-book and fact-based QA settings, we undermine the correctness of the responses and assess the uncertainty of their generation process. Surprisingly, trivial instruction-based attacks report the highest success rate (up to ~85.3%) while simultaneously having a high uncertainty for incorrectly answered questions. To provide a simple defense mechanism against Xmera, we train Random Forest classifiers on the response uncertainty levels to distinguish between attacked and unattacked queries (average AUC of up to ~94.8%). We believe that signaling users to be cautious about the answers they receive from black-box and potentially corrupt LLMs is a first checkpoint toward user cyberspace safety.

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

arXiv:2601.13719v2 Announce Type: replace-cross Abstract: Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We present HAVEN, a unified framework for long-video understanding that enables coherent and comprehensive reasoning by integrating audiovisual entity cohesion and hierarchical video indexing with agentic search. First, we preserve semantic consistency by integrating entity-level representations across visual and auditory streams, while organizing content into a structured hierarchy spanning global summary, scene, segment, and entity levels. Then we employ an agentic search mechanism to enable dynamic retrieval and reasoning across these layers, facilitating coherent narrative reconstruction and fine-grained entity tracking. Extensive experiments demonstrate that our method achieves good temporal coherence, entity consistency, and retrieval efficiency, establishing a new state-of-the-art with an overall accuracy of 84.1% on LVBench. Notably, it achieves outstanding performance in the challenging reasoning category, reaching 80.1%. These results highlight the effectiveness of structured, multimodal reasoning for comprehensive and context-consistent understanding of long-form videos.

19-Hydroxybufalin Inhibits Gastric Cancer Cell Proliferation by Modulating Metabolic Reprogramming

17 March 2026 at 18:00

J Proteome Res. 2026 Apr 3;25(4):2014-2023. doi: 10.1021/acs.jproteome.5c00983. Epub 2026 Mar 17.

ABSTRACT

OBJECTIVE: 19-Hydroxybufalin (19-H) is a natural bioactive compound with anticancer potential, but its molecular target and mechanism of action remain unclear. This study aimed to systematically evaluate its antigastric cancer activity and identify potential molecular targets.

METHODS: The antitumor effect of 19-H was evaluated in both in vitro and in vivo models. Multiomics analysis, thermal proteome profiling, molecular docking, and molecular dynamics simulations were employed to elucidate the mechanism of action. Functional assays were further conducted to validate the key target.

RESULTS: 19-H exhibited nanomolar-level inhibitory activity against various gastric cancer cell lines, significantly suppressing tumor growth in subcutaneous xenograft and patient-derived xenograft models. Multiomics analysis revealed that 19-H reshaped metabolic pathways in gastric cancer. TPP screening identified PLPP2 as a potential target with significantly increased thermal stability upon 19-H treatment. Molecular simulations further revealed that 19-H binds stably to the Ξ±-helical region of PLPP2.

CONCLUSIONS: 19-H exerts its antigastric cancer effect by targeting PLPP2 and remodeling the metabolic network. PLPP2 may represent a novel therapeutic target for gastric cancer.

PMID:41842934 | DOI:10.1021/acs.jproteome.5c00983

❌