❌

Normal view

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection

arXiv:2604.04815v1 Announce Type: cross Abstract: The rapid development of Large Language Models (LLMs) has transformed fake news detection and fact-checking tasks from simple classification to complex reasoning. However, evaluation frameworks have not kept pace. Current benchmarks are static, making them vulnerable to benchmark data contamination (BDC) and ineffective at assessing reasoning under temporal uncertainty. To address this, we introduce LiveFact a continuously updated benchmark that simulates the real-world "fog of war" in misinformation detection. LiveFact uses dynamic, temporal evidence sets to evaluate models on their ability to reason with evolving, incomplete information rather than on memorized knowledge. We propose a dual-mode evaluation: Classification Mode for final verification and Inference Mode for evidence-based reasoning, along with a component to monitor BDC explicitly. Tests with 22 LLMs show that open-source Mixture-of-Experts models, such as Qwen3-235B-A22B, now match or outperform proprietary state-of-the-art systems. More importantly, our analysis finds a significant "reasoning gap." Capable models exhibit epistemic humility by recognizing unverifiable claims in early data slices-an aspect traditional static benchmarks overlook. LiveFact sets a sustainable standard for evaluating robust, temporally aware AI verification.

HAG: Hierarchical Demographic Tree-based Agent Generation for Topic-Adaptive Simulation

arXiv:2601.05656v3 Announce Type: replace Abstract: High-fidelity agent initialization is crucial for credible Agent-Based Modeling across diverse domains. A robust framework should be Topic-Adaptive, capturing macro-level joint distributions while ensuring micro-level individual rationality. Existing approaches fall into two categories: static data-based retrieval methods that fail to adapt to unseen topics absent from the data, and LLM-based generation methods that lack macro-level distribution awareness, resulting in inconsistencies between micro-level persona attributes and reality. To address these problems, we propose HAG, a Hierarchical Agent Generation framework that formalizes population generation as a two-stage decision process. Firstly, utilizing a World Knowledge Model to infer hierarchical conditional probabilities to construct the Topic-Adaptive Tree, achieving macro-level distribution alignment. Then, grounded real-world data, instantiation and agentic augmentation are carried out to ensure micro-level consistency. Given the lack of specialized evaluation, we establish a multi-domain benchmark and a comprehensive PACE evaluation framework. Extensive experiments show that HAG significantly outperforms representative baselines, reducing population alignment errors by an average of 37.7% and enhancing sociological consistency by 18.8%.

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

arXiv:2603.22844v2 Announce Type: new Abstract: Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-based desmoking approaches rely on scarce paired supervision and deterministic restoration pipelines, making it difficult to perform exploration or reinforcement-driven refinement under real surgical conditions. We propose PhySe-RPO, a diffusion restoration framework optimized through Physics- and Semantics-Guided Relative Policy Optimization. The core idea is to transform deterministic restoration into a stochastic policy, enabling trajectory-level exploration and critic-free updates via group-relative optimization. A physics-guided reward imposes illumination and color consistency, while a visual-concept semantic reward learned from CLIP-based surgical concepts promotes smoke-free and anatomically coherent restoration. Together with a reference-free perceptual constraint, PhySe-RPO produces results that are physically consistent, semantically faithful, and clinically interpretable across synthetic and real robotic surgical datasets, providing a principled route to robust diffusion-based restoration under limited paired supervision.

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

Risk-adaptive therapy guided by dynamic ctDNA in nasopharyngeal carcinoma

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10244-w

A clinical trial testing whether monitoring ctDNA clearance during treatment for nasopharyngeal cancer could be used to inform decisions about an individual’s subsequent therapeutic programme shows promising results.

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

arXiv:2507.01352v3 Announce Type: replace-cross Abstract: Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced human preferences. We hypothesize that this brittleness stems primarily from limitations in preference datasets, which are often narrowly scoped, synthetically labeled, or lack rigorous quality control. To address these challenges, we present SynPref-40M, a large-scale preference dataset comprising 40 million preference pairs. To enable data curation at scale, we design a human-AI synergistic two-stage pipeline that leverages the complementary strengths of human annotation quality and AI scalability. In this pipeline, humans provide verified annotations, while LLMs perform automatic curation based on human guidance. Training on this preference mixture, we introduce Skywork-Reward-V2, a suite of eight reward models ranging from 0.6B to 8B parameters, trained on a carefully curated subset of 26 million preference pairs from SynPref-40M. We demonstrate that Skywork-Reward-V2 is versatile across a wide range of capabilities, including alignment with human preferences, objective correctness, safety, resistance to stylistic biases, and best-of-N scaling. These reward models achieve state-of-the-art performance across seven major reward model benchmarks, outperform generative reward models, and demonstrate strong downstream performance. Ablation studies confirm that effectiveness stems not only from data scale but also from high-quality curation. The Skywork-Reward-V2 series represents substantial progress in open reward models, demonstrating how human-AI curation synergy can unlock significantly higher data quality.

A Very Big Video Reasoning Suite

arXiv:2602.20159v1 Announce Type: cross Abstract: Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/ .
❌