❌

Normal view

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

arXiv:2609.12446v1 Announce Type: new Abstract: Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.

Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings

arXiv:2609.11983v1 Announce Type: cross Abstract: An OpenReview bug in November 2025 broke anonymity at several conferences and prompted calls for open review, which motivate us to ask what shifting from blind to open would mean for authors. Analyzing over 18,000 reviewed submissions to ICLR 2026, split into de facto open and blind groups by arXiv preprint timing, we find that ratings rise with author reputation under both mechanisms, with a steeper slope under open review that is statistically significant, and that the open-blind difference is concentrated at the borderline ratings. The pattern holds across five reputation proxies (including institution, h-index, and citation count), three author-aggregation rules, and five definitions of the open window. A controlled simulation with five AI models as reviewers, holding the manuscript fixed and varying the author reputation, reproduces the effect. With claude-opus-5 as the reviewer, for example, rating rises by 0.5 points as the author moves from low to high reputation.

MInTRL: Off-policy Intervention can boost On-policy RL

arXiv:2609.12419v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts. During generation, a judge-intervention policy periodically reviews the current policy's output, replaces erroneous suffixes with short corrections, and immediately returns control to the policy. During training, MInTRL adopts a sequence-level advantage-regression objective that eliminates the need for importance sampling. We show that sparse, local interventions can substantially improve coverage beyond finite-budget on-policy sampling while preserving the overall on-policy nature of the resulting trajectories. Across math and code benchmarks, MInTRL consistently outperforms standard on-policy and off-policy baselines. Ablations show that MInTRL remains effective with self-intervention and across different judge policies, while performance peaks at moderate intervention intensity, highlighting the importance of intervening minimally. These results establish minimal intervention as an effective paradigm for enhancing on-policy RL.

Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

arXiv:2504.04903v3 Announce Type: replace-cross Abstract: We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.

Advanced and underlying therapeutic strategies in transformed small cell lung cancer

Front Med (Lausanne). 2026 Aug 27;13:1865050. doi: 10.3389/fmed.2026.1865050. eCollection 2026.

ABSTRACT

Transformed small-cell lung cancer (T-SCLC) is a clinically important form of histologic transformation and a mechanism of acquired resistance in non-small-cell lung cancer (NSCLC). It is associated with poor prognosis, with a median overall survival of only about 9-13 months. This review summarizes recent advances in the mechanisms, diagnosis, monitoring, and treatment of T-SCLC. Repeat biopsy remains the gold standard for confirming histologic transformation, whereas molecular profiling and liquid biopsy may facilitate early detection and longitudinal disease monitoring. Platinum-etoposide remains the most commonly used clinical standard after transformation, but its benefit is typically transient and durable disease control remains uncommon. Continuation of EGFR tyrosine kinase inhibitors combined with chemotherapy may prolong progression-free survival in selected patients but has not consistently improved overall survival. Anti-angiogenic therapy, particularly anlotinib, and chemo-immunotherapy have shown encouraging activity in selected patients, while emerging strategies targeting DLL3, MYC, SOX2, and epigenetic regulators may broaden the therapeutic landscape. Prospective studies integrating repeat tissue sampling, comprehensive genomic profiling, biomarker-guided patient stratification, pharmacogenomics, functional drug-sensitivity testing where feasible, and integrated multi-omics approaches are needed to advance molecularly guided and individualized treatment for T-SCLC.

PMID:42724635 | PMC:PMC13560167 | DOI:10.3389/fmed.2026.1865050

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

❌