❌

Reading view

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

arXiv:2609.09783v1 Announce Type: cross Abstract: Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.
  •  

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

arXiv:2609.09793v1 Announce Type: cross Abstract: Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply it to GLM-5.3-Flash (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8). The attack survives the architecture, but what it reaches is no longer where a reader of the original recipe would look for it. Editing the attention, dense and routed-expert writers on their own removes 0.039, 0.016 and 0.148 of refusal respectively; editing all three together removes 0.776. As a result, 74% of the effect exists only under the joint intervention. The part the conventional recipe reaches by module-name matching accounts for 0.066 of that 0.776, which is why it fails silently on an MoE. The effect does not follow from removing just any direction: ablating a random direction orthogonal to it leaves refusal unchanged. A category-concentrated residue survives every edit we tried: subspaces fitted on violence, sexual content and hate leave measurable refusal at every rank from 1 to 12. We report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.
  •  

GANDR: Claim Auditing for Verifiable Legal Answer Generation

arXiv:2609.10293v1 Announce Type: cross Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as a whole, so a correct conclusion can rest on fabricated or loosely matched citations and still score well. Closing this gap requires both a system built for per-claim verification and an evaluation that measures it. We introduce GANDR (Grounded ANswer DRafter), a two-agent system in which a Drafter writes an answer in a structured legal-reasoning format and a separate Critic, with the same view as a human verifier, audits each claim against its cited source and emits a per-claim audit trace on every round. We pair it with a strict correctness criterion requiring every citation to resolve to a passage the retriever returned. On a 185-item legal benchmark where all six systems share one backbone, one retrieval surface, and one citation instruction, GANDR ranks first on every primary metric, reaching 70.8% strict accuracy and leading the strongest baseline by 11.3 points (p
  •  

MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production

arXiv:2609.10385v1 Announce Type: cross Abstract: Animation and VFX pre-production review requires teams to translate loosely specified creative intent--briefs, evolving specifications, heterogeneous references, and verbal decisions--into revisions that junior artists can execute without repeated clarification. In practice, criteria drift across iterations, review judgments lose their evidential basis, and the reasoning behind a request rarely survives the senior-junior handoff. We contribute a design framework for intent-evidence-action alignment: intent is articulated into a shared project record, judgments are anchored to grounded evidence, and authorized decisions are converted into clear revision tasks tied directly to reference notes. We instantiate this framework in MOONWALK, a professional pre-production review system comprising a shared intent record, reference/specification anchoring, structured work-in-progress comparison, and supervisor-authorized action planning. In this workflow, AI handles administrative coordination--flagging missing context and organizing notes--while artists retain full creative direction. An in-studio study with professional practitioners compares MOONWALK with a chat-only (chatbot) interface using matched production materials, while participants' existing workflows provide a retrospective ecological baseline. Results indicate stronger intent alignment, decision traceability, and checklist executability, while also showing that aesthetic authority and final prioritization must remain with practitioners. The evaluation establishes the value of the integrated structured workflow over unstructured conversational AI chatbot. Code: https://github.com/Akinesia112/Moonwalk/tree/english-version
  •  

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

  •  

S1P-TREM2 axis protects immunosuppressive neutrophils from ferroptosis to promote tumour progression in hepatocellular carcinoma

Gut. 2026 Sep 7:gutjnl-2025-337414. doi: 10.1136/gutjnl-2025-337414. Online ahead of print.

ABSTRACT

BACKGROUND: Neutrophils are increasingly recognised as immunosuppressive drivers of hepatocellular carcinoma (HCC), yet their persistence in the oxidative, lipid-rich tumour microenvironment remains poorly understood.

OBJECTIVE: To elucidate the metabolic and molecular programmes that enable tumour-associated neutrophils (TANs) to resist ferroptosis and sustain immunosuppression in HCC.

DESIGN: We employed human HCC samples, multiple murine HCC models, transcriptomic and lipidomic profiling, genetic loss-of-function systems and therapeutic interventions. Ferroptosis sensitivity, lipid metabolic rewiring and immunological consequences of TANs were systematically evaluated across models and validated in patient datasets and biospecimens.

RESULTS: TANs in human HCC and mouse models exhibit pronounced lipid accumulation and oxidative stress compared with peripheral neutrophils. Multi-omic profiling revealed that TANs are enriched for lipid-binding gene programmes and undergo rewiring towards sphingolipid and unsaturated fatty acid metabolism. We identified triggering receptor expressed on myeloid cells 2 (TREM2) as a key lipid-sensing receptor selectively expressed in TANs. Functional deletion of TREM2 reprogrammed the tumour immune microenvironment, restoring CD8+ T cell activity and suppressing HCC progression. Mechanistically, tumour-derived sphingosine-1-phosphate (S1P) activates TREM2, triggering nuclear factor erythroid 2-related factor 2 (NRF2)-mediated transcription of glutathione peroxidase 4 (GPX4) and solute carrier family 7 member 11 (SLC7A11), thereby promoting ferroptosis resistance. TREM2 expression is transcriptionally induced by granulocyte-macrophage colony-stimulating factor-signal transducer and activator of transcription 3 (GM-CSF-STAT3) signalling. Genetic deletion of TREM2, clustered regularly interspaced short palindromic repeats/CRISPR-associated protein 9 (CRISPR/Cas9)-mediated knockout of sphingosine kinase 1/2 (SPHK1/2) in tumour cells, or pharmacological inhibition of S1P synthesis disrupts this protective lipid-immune circuit, sensitises TANs to ferroptosis and restricts tumour growth. Therapeutically, a peptide-based TREM2 inhibitor reprogrammes TANs, restores CD8+ T cell function and enhances anti-programmed cell death protein 1 (PD-1) immunotherapy efficacy. Clinically, TREM2+ polymorphonuclear myeloid-derived suppressor cells (PMN-MDSCs) are enriched in HCC tumours, correlate with SPHK1/2 expression and T cell dysfunction and associate with poor patient prognosis.

CONCLUSION: Our study uncovers the S1P-TREM2-NRF2 axis as a critical metabolic-immune circuit that preserves neutrophil survival and immunosuppressive function in HCC. Targeting this lipid-dependent ferroptosis resistance pathway offers a promising therapeutic strategy to overcome immunotherapy resistance in liver cancer.

PMID:42705697 | DOI:10.1136/gutjnl-2025-337414

  •  
❌