❌

Normal view

GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting

arXiv:2609.12165v1 Announce Type: new Abstract: Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.

LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory

arXiv:2609.12436v1 Announce Type: new Abstract: Long-running LLM agents require memory mechanisms that maintain coherent internal states across interactions. We study a lifecycle-labeled memory setting in which write episodes provide lifecycle metadata during training, and phase-aware readout is used during evaluation. This setting reflects the need to distinguish information that should remain influential across future interactions from information that should affect only the current context. A mismatch between these lifecycles can cause temporary information to overwrite durable knowledge, leading to behavioral drift in persistent agents. Within this setting, we introduce \textbf{LifeFuse-Mem}, a lifecycle-aware neural memory framework that separates information according to its temporal commitment. LifeFuse-Mem uses dedicated memory components and lifecycle-aware updates to allow stable and transient knowledge to evolve locally without converting temporary context into durable state. On the controlled anti-overwrite benchmark, LifeFuse-Mem improves acquisition-controlled retention and reduces temporary overwrite; on two public long-memory benchmarks, it remains broadly competitive. These results suggest that explicit lifecycle signals can help diagnose and mitigate overwrite in compact online memory.

Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving

arXiv:2609.12606v1 Announce Type: new Abstract: While multimodal reasoning has advanced rapidly, solving complex geometry problems critically hinges on active visual assistance, such as constructing auxiliary lines, spurring the rise of Visual Chain-of-Thought (VCoT). However, existing evaluations typically assess visual generation quality and final answer accuracy in isolation, failing to examine whether intermediate visual aids are geometrically valid, effectively utilized in subsequent reasoning, or causally responsible for task success. To bridge this gap, we introduce GeoVAD-Bench, a diagnostic benchmark that pairs a fine-grained five-dimensional trajectory diagnosis covering perception, auxiliary quality, utilization, deductive reasoning, and final correctness with controlled No-Aux, Auto-Aux, and GT-Aux intervention settings to systematically isolate intermediate error modes, the causal gains of visual aids, and the resulting autonomy gap. Our findings reveal that while high-quality auxiliary aids offer substantial theoretical gains for geometric problem solving, autonomous generation is frequently hampered by compounding errors across geometric perception, faithful visual manipulation, visual-state grounding, and deductive reasoning. Guided by these diagnostic insights, we establish a specialized data construction pipeline encompassing geometric perception, diagram editing, and interleaved visual-textual reasoning trajectories, and develop a progressive SFT and multimodal RL training framework. The resulting model, GeoWeave-8B, outperforms the base model by +25.3% in final geometric accuracy and achieves a +30.4% gain in process average across the four intermediate diagnostic dimensions.

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

arXiv:2603.01131v4 Announce Type: replace-cross Abstract: Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations, competing hypotheses, disease relations, and treatment decisions. Large language models have advanced medical text understanding and generation. Yet their clinical use remains limited by weak evidence grounding, opaque reasoning, and inconsistent links among differential diagnosis, final diagnosis, diagnostic basis, and treatment planning. We introduce MedCollab, a multi-agent framework for full-cycle clinical diagnosis and report generation. MedCollab coordinates specialist and examination agents according to patient records. It structures agent deliberation with an Issue-Based Information System (IBIS) protocol, so that each diagnostic position is supported by patient-specific evidence and medical knowledge. It also builds Hierarchical Disease Relation Chains (HDRC) to connect accepted hypotheses through progression, complication, and comorbidity relations. During multi-round deliberation, a verifier-guided consensus module evaluates evidence support, medical plausibility, and logical conflicts. It then adjusts agent contributions and filters unsupported reasoning. Experiments on ClinicalBench and MIMIC-IV show that MedCollab outperforms leading LLMs and medical multi-agent baselines in diagnostic accuracy, evidence consistency, and clinical reasoning quality. These results indicate that structured and auditable collaboration can produce more faithful and clinically coherent diagnostic reports.

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

❌