❌

Normal view

SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy

arXiv:2609.12749v1 Announce Type: new Abstract: Offline-to-online reinforcement learning reduces interaction cost for real-world robot learning but suffers from persistent value estimation instability. Existing methods address this through pessimistic regularization, lower-bound calibration, and architectural normalization, but an overlooked source of instability lies in the entropy formulation: the standard log-entropy term can become negative, destabilizing policy updates. We introduce SCQ (Sigmoid-Bounded Conservative Q-Learning), which replaces this term with a sigmoid-bounded formulation that stays strictly positive. SCQ retains conservative Q regularization and return-based lower-bound calibration, stabilizing policy optimization without sacrificing exploration. We evaluate SCQ on D4RL (Minari) benchmarks under both single-demonstration and standard dataset settings, as well as on simulation and real-world visual tasks. SCQ matches or exceeds baseline performance while exhibiting more stable training dynamics across state-based and visual benchmarks, and transfers to four real-robot platforms including manipulation, wheeled, quadruped, and humanoid systems. A direct clipping intervention that removes negative log-probability contributions, together with gradient-matched positive-score controls, indicates that positivity rather than a particular score shape alone drives much of the improvement. Project website: https://scq-rl.github.io.

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

arXiv:2507.01663v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.

CAR T cells secreting anti-EpCAM bispecific T cell engagers overcome tumor heterogeneity in targeting epithelial-originated carcinomas

CAR T cells engineered to secrete tumor-localized anti-EpCAM bispecific T cell engagers (BTCEs) overcome antigen escape and heterogeneity across multiple epithelial carcinomas in preclinical models, achieving complete tumor eradication where conventional single-target CAR T cell therapies often failed, supporting broad translational potential for solid tumor immunotherapy.

In vivo-directed evolution identifies AAV-WM04 as a next-generation vector for potent and sustained hearing restoration in DFNB9

AAV-WM04, an AAV vector identified through in-vivo-directed screening in the adult cochlea, enables highly efficient and selective inner hair cell transduction. Dual-AAV delivery of OTOF using AAV-WM04 restores hearing in a DFNA9 deafness mouse model at low doses, highlighting its translational potential for gene therapy.

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

arXiv:2605.25446v1 Announce Type: new Abstract: Electrocardiography (ECG) is central to cardiovascular care, but conventional AI models are often restricted to common arrhythmias and may generalize poorly across populations or clinically subtle diseases. We developed ECG Contrastive Language-Image Pre-training (ECGCLIP), a signal-language contrastive learning framework that aligns ECG waveforms with expert diagnostic reports. ECGCLIP was pre-trained on 2,837,962 ECG studies from 1,324,856 patients and evaluated on a held-out internal test set plus nine independent external cohorts comprising about 1.5 million ECGs. Evaluation covered 89 downstream tasks, including 45 ECG diagnoses, 39 echocardiographic targets, and 5 rare cardiac diseases, using PRAUC as the primary metric. ECGCLIP consistently improved performance over random initialization and Merl-R18 baselines. On the internal test set, ECGCLIP-R34 achieved strong performance for atrial fibrillation (PRAUC 0.900) and ST-segment elevation myocardial infarction (PRAUC 0.383), with robust generalization across all external cohorts. It also improved low-prevalence and diagnostically elusive diseases, including Ebstein anomaly, constrictive pericarditis, dextrocardia, and cardiac amyloidosis, with internal PRAUC values of 0.253, 0.175, 0.121, and 0.201, respectively. ECGCLIP was data efficient, matching or exceeding full-dataset baseline performance with only 10% of training data. Feature visualization and saliency analysis suggested clinically meaningful representations aligned with established electrocardiographic criteria. These findings indicate that large-scale ECG-report contrastive pre-training can expand routine ECG interpretation beyond common arrhythmias toward broad cardiovascular assessment and opportunistic screening of echocardiographic and rare conditions.

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

arXiv:2605.26074v1 Announce Type: cross Abstract: Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market. We introduce StakeBench, an evaluation framework for language understanding grounded in market commitment. StakeBench links 560,876 comments from 2,261 resolved markets to verified position, action, and market-odds records across Polymarket and Manifold. Supervision is derived from observable market behavior. Position sides, post-comment trading actions, and market-odds trajectories replace human annotation. Four diagnostic tasks test whether models detect market commitment, identify the revealed side, anticipate future action, and perform collective odds projection. Three commitment-aware metrics measure alignment with revealed preferences rather than perceived sentiment. Validity audits and explicit interpretation boundaries help distinguish observable commitment signals from latent belief and causal market-odds impact. Across 15 LLMs and 18 topics and platform settings, models partially recover position-side signals, with Directed Accuracy from 0.506 to 0.599, but show structural failures on later tasks. Ten of the fifteen models collapse to one or two action labels in future action anticipation, and no model consistently improves on the naive odds-direction baseline in collective odds projection. Model scale is not correlated with performance, finance-domain tuning does not improve revealed-side identification, and platform incentives strongly shape higher-order results. StakeBench is packaged with evaluation code and dataset under CC-BY 4.0.

SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

arXiv:2605.10989v3 Announce Type: replace-cross Abstract: The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that suffer from gradient mismatch problem and information loss induced by fixed-range gradient clipping. To address this, we propose SURrogate GradiEnt Adaptation (SURGE), a novel learnable gradient compensation framework with theoretical grounding. SURGE mitigates gradient mismatch through auxiliary backpropagation. Specifically, we design a Dual-Path Gradient Compensator (DPGC) that constructs a parallel full-precision auxiliary branch for each binarized layer, decoupling gradient flow via output decomposition during backpropagation. DPGC enables bias-reduced gradient estimation by leveraging the full-precision branch to estimate components beyond STE's first-order approximation. To further enhance training stability, we introduce an Adaptive Gradient Scaler (AGS) based on an optimal scale factor to dynamically balance inter-branch gradient contributions via norm-based scaling. Experiments on image classification, object detection, and language understanding tasks demonstrate that SURGE performs best over state-of-the-art methods.

Mummified early Permian reptile reveals ancient amniote breathing apparatus

Nature, Published online: 08 April 2026; doi:10.1038/s41586-026-10307-y

A mummified fossil of the early Permian reptile Captorhinus reveals the potential ancestral amniote breathing mechanism and its impact on terrestrial vertebrate evolution.

IC3-Evolve: Proof-/Witness-Gated Offline LLM-Driven Heuristic Evolution for IC3 Hardware Model Checking

arXiv:2604.03232v1 Announce Type: new Abstract: IC3, also known as property-directed reachability (PDR), is a commonly-used algorithm for hardware safety model checking. It checks if a state transition system complies with a given safety property. IC3 either returns UNSAFE (indicating property violation) with a counterexample trace, or SAFE with a checkable inductive invariant as the proof to safety. In practice, the performance of IC3 is dominated by a large web of interacting heuristics and implementation choices, making manual tuning costly, brittle, and hard to reproduce. This paper presents IC3-Evolve, an automated offline code-evolution framework that utilizes an LLM to propose small, slot-restricted and auditable patches to an IC3 implementation. Crucially, every candidate patch is admitted only through proof- /witness-gated validation: SAFE runs must emit a certificate that is independently checked, and UNSAFE runs must emit a replayable counterexample trace, preventing unsound edits from being deployed. Since the LLM is used only offline, the deployed artifact is a standalone evolved checker with zero ML/LLM inference overhead and no runtime model dependency. We evolve on the public hardware model checking competition (HWMCC) benchmark and evaluate the generalizability on unseen public and industrial model checking benchmarks, showing that IC3-Evolve can reliably discover practical heuristic improvements under strict correctness gates.

Multi-omics integration and machine learning reveal gut-immune signatures in idiopathic pulmonary fibrosis: insights from bulk RNA-seq, single-cell profiles, spatial transcriptomics, and experimental validation

Front Immunol. 2026 Mar 19;17:1730289. doi: 10.3389/fimmu.2026.1730289. eCollection 2026.

ABSTRACT

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive, fatal lung disease with limited treatment options and a poor prognosis. Recent studies suggest a critical role for the gut-immune-lung axis in IPF, yet the underlying molecular mechanisms remain unclear.

METHODS: The current study performed in silico multi-omics integration of publicly available datasets, including bulk RNA-seq, single-cell and spatial transcriptomics, as well as peripheral blood multi-omics data to uncover key molecular signatures in IPF. Furthermore, machine learning techniques were utilized to identify core genes, whereas functional analyses and Mendelian randomization were conducted to evaluate the causal relationships among gut microbiota, immune cells, and IPF. Additionally, experimental validation using qPCR and ELISA assays was conducted in vitro, in vivo, and in patient plasma to confirm the expression patterns of key genes.

RESULTS: Across integrated public bulk, single-cell, spatial, and blood multi-omics, CXCL13, IL33, TLR4, and IGF1 were identified as core IPF genes consistently linked to immune infiltration and fibrotic remodeling. Deconvolution, scRNA-seq, and spatial mapping localized their dysregulation to fibroblasts and immune compartments (notably B-cell, macrophage, and mast-cell axes), highlighting fibroblast-immune crosstalk in fibrotic foci. A four-gene model robustly distinguished IPF from controls across cohorts. Mendelian randomization supported a gut-immune-lung axis, indicating causal effects of specific gut taxa on IPF risk via immune phenotypes. qPCR/ELISA in TGF-β1-stimulated fibroblasts, bleomycin mouse lungs, and patient plasma corroborated upregulation of IL33, CXCL13, IGF1 and downregulation of TLR4. Drug-signature reversal nominated cucurbitacin I and temsirolimus; molecular docking was performed as a preliminary in silico, computer-simulation-based assessment of potential ligand-protein interactions between these compounds and the four core targets.

CONCLUSION: This study provides new insights into the importance of gut-immune-lung axis in IPF and identifies CXCL13, IL33, TLR4, and IGF1 as diagnostic signatures and therapeutic targets. By integrating public multi-omics resources with experimental validation, our findings offer a foundation for future diagnostic and treatment strategies aimed at modulating the gut microbiota and immune system in IPF.

PMID:41939867 | PMC:PMC13043422 | DOI:10.3389/fimmu.2026.1730289

Multi-omics integration and machine learning reveal gut-immune signatures in idiopathic pulmonary fibrosis: insights from bulk RNA-seq, single-cell profiles, spatial transcriptomics, and experimental validation

Front Immunol. 2026 Mar 19;17:1730289. doi: 10.3389/fimmu.2026.1730289. eCollection 2026.

ABSTRACT

BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive, fatal lung disease with limited treatment options and a poor prognosis. Recent studies suggest a critical role for the gut-immune-lung axis in IPF, yet the underlying molecular mechanisms remain unclear.

METHODS: The current study performed in silico multi-omics integration of publicly available datasets, including bulk RNA-seq, single-cell and spatial transcriptomics, as well as peripheral blood multi-omics data to uncover key molecular signatures in IPF. Furthermore, machine learning techniques were utilized to identify core genes, whereas functional analyses and Mendelian randomization were conducted to evaluate the causal relationships among gut microbiota, immune cells, and IPF. Additionally, experimental validation using qPCR and ELISA assays was conducted in vitro, in vivo, and in patient plasma to confirm the expression patterns of key genes.

RESULTS: Across integrated public bulk, single-cell, spatial, and blood multi-omics, CXCL13, IL33, TLR4, and IGF1 were identified as core IPF genes consistently linked to immune infiltration and fibrotic remodeling. Deconvolution, scRNA-seq, and spatial mapping localized their dysregulation to fibroblasts and immune compartments (notably B-cell, macrophage, and mast-cell axes), highlighting fibroblast-immune crosstalk in fibrotic foci. A four-gene model robustly distinguished IPF from controls across cohorts. Mendelian randomization supported a gut-immune-lung axis, indicating causal effects of specific gut taxa on IPF risk via immune phenotypes. qPCR/ELISA in TGF-β1-stimulated fibroblasts, bleomycin mouse lungs, and patient plasma corroborated upregulation of IL33, CXCL13, IGF1 and downregulation of TLR4. Drug-signature reversal nominated cucurbitacin I and temsirolimus; molecular docking was performed as a preliminary in silico, computer-simulation-based assessment of potential ligand-protein interactions between these compounds and the four core targets.

CONCLUSION: This study provides new insights into the importance of gut-immune-lung axis in IPF and identifies CXCL13, IL33, TLR4, and IGF1 as diagnostic signatures and therapeutic targets. By integrating public multi-omics resources with experimental validation, our findings offer a foundation for future diagnostic and treatment strategies aimed at modulating the gut microbiota and immune system in IPF.

PMID:41939867 | PMC:PMC13043422 | DOI:10.3389/fimmu.2026.1730289

Accuracy of Radiomics-Based Machine Learning for Predicting Risk of Recurrence in Non–Small Cell Lung Cancer: Systematic Review and Meta-Analysis

Background: During the diagnosis and treatment of non–small cell lung cancer (NSCLC), detecting the risk of its recurrence in an early phase is still challenging. Recent studies have investigated the radiomics-based machine learning (ML) models for detecting the risk of recurrence in NSCLC. However, there is still insufficient systematic evidence to prove its efficiency. Objective: This study is designed to systematically evaluate the effectiveness of radiomics-based ML in predicting the risk of recurrence in NSCLC, aiming to provide evidence-based support for the subsequent development of scoring tools to forecast recurrence risk. Methods: For acquiring research on radiomics-based models for forecasting the risk of recurrence in NSCLC, Cochrane Library, Web of Science, PubMed, and Embase were systematically retrieved, up to October 24, 2025. Studies on analyzing the recurrence of NSCLC using radiomics-based ML were included, while those in which only texture analysis was conducted or radiomics-based ML was not constructed were excluded. The Radiomics Quality Score (RQS) was used to appraise the eligible studies. Subgroup analyses were conducted according to the variables of the model, the background of treatment, the stage of lung cancer, and the pathological type. Results: Ultimately, 30 eligible studies in total were included, covering 7964 patients with NSCLC. According to the meta-analysis, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.850 (95% CI 0.834‐0.866, 95% prediction interval [PI] 0.623‐1.004) in the training set. Specifically, the pooled c-index was 0.876 (95% CI 0.853‐0.900) among the patients receiving the stereotactic body radiation therapy and 0.825 (95% CI 0.804‐0.848) among those who received surgeries combined with other adjuvant treatment regimens. The c-index of the radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.833 (95% CI 0.822‐0.854, 95% PI 0.717‐0.945) in the training set. In contrast, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.878 (95% CI 0.854‐0.902, 95% PI 0.681‐1.000) in the validation set. The c-index of radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.854 (95% CI 0.830‐0.878, 95% PI 0.655‐0.992) in the validation set. The average RQS across the included studies was 27.4%, revealing methodological limitations and an absence of standardization. Conclusions: This study is the first to confirm that radiomics-based ML models effectively predict the risk of recurrence in NSCLC. This study provides evidence-based support for the subsequent development or updating of radiomics-based ML models. However, the current methodological application of radiomics remains concerning. Therefore, in the future, research should standardize the workflow for implementing radiomics-based ML and incorporate multicenter imaging data to enhance its generalizability. Trial Registration: PROSPERO CRD42025631191; https://www.crd.york.ac.uk/PROSPERO/view/CRD42025631191

WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment

arXiv:2603.22690v1 Announce Type: cross Abstract: Privacy-preserving semantic understanding of human activities is important for indoor sensing, yet existing Wi-Fi CSI-based systems mainly focus on pose estimation or predefined action classification rather than fine-grained language generation. Mapping CSI to natural-language descriptions remains challenging because of the semantic gap between wireless signals and language and direction-sensitive ambiguities such as left/right limb confusion. We propose WiFi2Cap, a three-stage framework for generating action captions directly from Wi-Fi CSI. A vision-language teacher learns transferable supervision from synchronized video-text pairs, and a CSI student is aligned to the teacher's visual space and text embeddings. To improve direction-sensitive captioning, we introduce a Mirror-Consistency Loss that reduces mirrored-action and left-right ambiguities during cross-modal alignment. A prefix-tuned language model then generates action descriptions from CSI embeddings. We also introduce the WiFi2Cap Dataset, a synchronized CSI-RGB-sentence benchmark for semantic captioning from Wi-Fi signals. Experimental results show that WiFi2Cap consistently outperforms baseline methods on BLEU-4, METEOR, ROUGE-L, CIDEr, and SPICE, demonstrating effective privacy-friendly semantic sensing.

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

arXiv:2603.21597v2 Announce Type: replace Abstract: Modern clinical practice increasingly depends on reasoning over heterogeneous, evolving, and incomplete patient data. Although recent advances in multimodal foundation models have improved performance on various clinical tasks, most existing models remain static, opaque, and poorly aligned with real-world clinical workflows. We present Cerebra, an interactive multi-agent AI team that coordinates specialized agents for EHR, clinical notes, and medical imaging analysis. These outputs are synthesized into a clinician-facing dashboard that combines visual analytics with a conversational interface, enabling clinicians to interrogate predictions and contextualize risk at the point of care. Cerebra supports privacy-preserving deployment by operating on structured representations and remains robust when modalities are incomplete. We evaluated Cerebra using a massive multi-institutional dataset spanning 3 million patients from four independent healthcare systems. Cerebra consistently outperformed both state-of-the-art single-modality models and large multimodal language model baselines. In dementia risk prediction, it achieved AUROCs up to 0.80, compared with 0.74 for the strongest single-modality model and 0.68 for language model baselines. For dementia diagnosis, it achieved an AUROC of 0.86, and for survival prediction, a C-index of 0.81. In a reader study with experienced physicians, Cerebra significantly improved expert performance, increasing accuracy by 17.5 percentage points in prospective dementia risk estimation. These results demonstrate Cerebra's potential for interpretable, robust decision support in clinical care.

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

arXiv:2603.08090v1 Announce Type: cross Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit critical limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across different subject difficulty levels and prompt scenarios, and 3) a profound lack of actionable insights and diagnostic guidance for subsequent model refinement. To address these limitations, we propose DSH-Bench, a comprehensive benchmark that enables systematic multi-perspective analysis of subject-driven T2I models through four principal innovations: 1) a hierarchical taxonomy sampling mechanism ensuring comprehensive subject representation across 58 fine-grained categories, 2) an innovative classification scheme categorizing both subject difficulty level and prompt scenario for granular capability assessment, 3) a novel Subject Identity Consistency Score (SICS) metric demonstrating a 9.4\% higher correlation with human evaluation compared to existing measures in quantifying subject preservation, and 4) a comprehensive set of diagnostic insights derived from the benchmark, offering critical guidance for optimizing future model training paradigms and data construction strategies. Through an extensive empirical evaluation of 19 leading models, DSH-Bench uncovers previously obscured limitations in current approaches, establishing concrete directions for future research and development.

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

arXiv:2512.10416v3 Announce Type: replace-cross Abstract: Deep learning has advanced vectorized road extraction in urban settings, yet off-road environments remain underexplored and challenging. A significant domain gap causes advanced models to fail in wild terrains due to two key issues: lack of large-scale vectorized datasets and structural weakness in prevailing methods. Models such as SAM-Road employ a node-centric paradigm that reasons at sparse endpoints, making them fragile to occlusions and ambiguous junctions in off-road scenes, leading to topological errors. This work addresses these limitations in two complementary ways. First, we release WildRoad, a global off-road road network dataset constructed efficiently with a dedicated interactive annotation tool tailored for road-network labeling. Second, we introduce MaGRoad (Mask-aware Geodesic Road network extractor), a path-centric framework that aggregates multi-scale visual evidence along candidate paths to infer connectivity robustly. Extensive experiments show that MaGRoad achieves state-of-the-art performance on our challenging WildRoad benchmark while generalizing well to urban datasets. An efficient vertex extraction strategy also yields roughly 2.5X faster inference, improving practical applicability. Together, the dataset and path-centric paradigm provide a stronger foundation for mapping roads in the wild. We release both the dataset and code at this repository. We release both the dataset and code at https://github.com/xiaofei-guan/MaGRoad.

When Relevance Meets Novelty: Dual-Stable Periodic Optimization for Serendipitous Recommendation

arXiv:2508.00450v3 Announce Type: replace-cross Abstract: Traditional recommendation systems tend to trap users in strong feedback loops by excessively pushing content aligned with their historical preferences, thereby limiting exploration opportunities and causing content fatigue. Although large language models (LLMs) demonstrate potential with their diverse content generation capabilities, existing LLM-enhanced dual-model frameworks face two major limitations: first, they overlook long-term preferences driven by group identity, leading to biased interest modeling; second, they suffer from static optimization flaws, as a one-time alignment process fails to leverage incremental user data for closed-loop optimization. To address these challenges, we propose the Co-Evolutionary Alignment (CoEA) method. For interest modeling bias, we introduce Dual-Stable Interest Exploration (DSIE) module, jointly modeling long-term group identity and short-term individual interests through parallel processing of behavioral sequences. For static optimization limitations, we design a Periodic Collaborative Optimization (PCO) mechanism. This mechanism regularly conducts preference verification on incremental data using the Relevance LLM, then guides the Novelty LLM to perform fine-tuning based on the verification results, and subsequently feeds back the output of the continually fine-tuned Novelty LLM to the Relevance LLM for re-evaluation, thereby achieving a dynamic closed-loop optimization. Extensive online and offline experiments verify the effectiveness of the CoEA model in serendipitous recommendation.

LEDOM: Reverse Language Model

arXiv:2507.01335v3 Announce Type: replace-cross Abstract: Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at scale, and ask what reasoning patterns emerge when a model conditions on future context to predict the past. We train LEDOM, an open-source purely reverse autoregressive language model (2B/7B parameters, 435B tokens), and find it develops capabilities distinct from forward models, including abductive inference, question synthesis, and natural resolution of the reversal curse. We then explore one application of the reverse model: combining forward likelihood $P(y \mid x)$ with reverse posterior $P(x \mid y)$ through noisy channel duality. We propose Reverse Reward, which reranks forward outputs using reverse posterior estimates, and prove that bidirectional scoring penalizes hallucinated reasoning chains whose backward reconstruction degrades. Reverse Reward yields gains of up to 6.6\% on AIME 2024 and 15\% on AMC 2023 across multiple strong baselines. We release all models, code, and data here.

DEFNet: Multitasks-based Deep Evidential Fusion Network for Blind Image Quality Assessment

arXiv:2507.19418v1 Announce Type: cross Abstract: Blind image quality assessment (BIQA) methods often incorporate auxiliary tasks to improve performance. However, existing approaches face limitations due to insufficient integration and a lack of flexible uncertainty estimation, leading to suboptimal performance. To address these challenges, we propose a multitasks-based Deep Evidential Fusion Network (DEFNet) for BIQA, which performs multitask optimization with the assistance of scene and distortion type classification tasks. To achieve a more robust and reliable representation, we design a novel trustworthy information fusion strategy. It first combines diverse features and patterns across sub-regions to enhance information richness, and then performs local-global information fusion by balancing fine-grained details with coarse-grained context. Moreover, DEFNet exploits advanced uncertainty estimation technique inspired by evidential learning with the help of normal-inverse gamma distribution mixture. Extensive experiments on both synthetic and authentic distortion datasets demonstrate the effectiveness and robustness of the proposed framework. Additional evaluation and analysis are carried out to highlight its strong generalization capability and adaptability to previously unseen scenarios.
❌