❌

Reading view

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.
  •  

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.
  •  

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

arXiv:2609.12808v1 Announce Type: new Abstract: Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent. We introduce K-Bench, a benchmark that scores LLM unlearning under agentic deployment. K-Bench inspects all six channels a ReAct agent exposes, including its chain-of-thought (CoT), tool calls and tool observations, and elicited summary. A query counts as leaked if the secret appears in any of them. Each experiment places the secret in exactly one of the agent's three sources (the weights, the prompt, or the retrieval store). The K-Score is computed separately for each source and credits forgetting only when the agent remains usable. Clearing the answer channel does not make the secret unrecoverable. On structured retrieval, the secret stays verbatim in the tool-observation channel and the aggregate leak rate is unchanged. When the secret lives in the prompt or the retrieval store, TOFU and MUSE report no leakage, while the deployed agent still leaks it on 22--86\% of queries. When the secret is in the weights, none of the twenty evaluated published methods demonstrably removes it, and only an input-corruption intervention reaches selective forgetting under the evaluated observer. The top-ranked method changes across base models. A refusal-tuning method resists the evaluated extraction without verified knowledge removal.
  •  

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

arXiv:2609.13076v1 Announce Type: cross Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.
  •  

HyQuant: Hybrid-Precision Quantization for LLM Attention

arXiv:2608.27875v2 Announce Type: replace Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .
  •  

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

arXiv:2507.01663v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.
  •  

Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization

arXiv:2606.23989v5 Announce Type: replace-cross Abstract: Large language models produce fluent multi-document summaries, but their attributions are typically coarse---whole documents or passages---and generated post hoc, leaving each statement hard to verify. We argue that attribution should be a structural property of generation rather than a downstream prediction. We present CAMS, a Claim-Anchored Multi-document Summarization framework that decomposes every source document into atomic claims whose provenance is resolved deterministically from verbatim quotes to token spans, clusters equivalent claims across documents while flagging inter-source conflicts, selects a support-aware and salient subset, and rewrites it so that every summary sentence terminates in claim identifiers resolving back to source spans. This yields a separation we make explicit: provenance is an invariant holding for every emitted sentence independently of model accuracy, whereas faithfulness is an objective that selection, constrained rewriting, and verification only encourage---a distinction end-to-end and post-hoc systems conflate. We evaluate on MultiNews, DiverseSumm, and zero-shot on WCEP under a two-regime protocol separating reference-free citation quality from gold-aligned localization, audited by a support model never used for selection or verification. CAMSmatches strong end-to-end and span-attribution baselines on summary quality while improving faithfulness and citation precision, raising multi-source attribution accuracy from 38% to 64% without inflating the number of cited sources, and cutting human verification time per claim by $3.4\times$. We release code and ${\sim}320$K claim--quote--span annotations over MultiNews as a reusable fine-grained attribution resource.
  •  

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

arXiv:2608.18234v3 Announce Type: replace-cross Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.
  •  

A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

arXiv:2608.21099v2 Announce Type: replace-cross Abstract: Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong representation capabilities, yet adapting them to multi-modal scenarios remains challenging. Existing dense cross-modal fusion strategies often force heterogeneous modalities to interact indiscriminately, which may introduce redundant information and disrupt the valuable pre-trained representations. To address this issue, we revisit multi-modal fusion from the perspective of socialized learning and propose adapter to DINOv3 (A2DINOv3), a multi-expert collaboration framework with a Socialized Collaboration Protocol (SCP). Specifically, RGB and infrared branches are modeled as heterogeneous experts that independently preserve their specialized knowledge while exchanging complementary information through selective and constrained interactions. This design mitigates harmful cross-modal interference and prevents degradation of pre-trained priors during adaptation. Furthermore, a zero-initialization strategy is introduced to gradually activate cross-modal collaboration, enabling a smooth transition from modality-specific learning to cooperative representation learning. Extensive experiments on four multi-modal benchmarks, including aerial detection (GAIIC), autonomous driving (FLIR), low-light surveillance (LLVIP), and diverse real-world scenarios (M3FD), demonstrate that A2DINOv3 consistently achieves state-of-the-art performance in multi-modal object detection.
  •  

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

  •  

MMSDH facilitates ACSL4 propionylation to counteract ferroptosis upon hypoxia and impairs PDAC chemotherapy efficacy

Nature Cancer, Published online: 11 September 2026; doi:10.1038/s43018-026-01236-w

Zheng et al. describe how hypoxia-induced methylmalonate semialdehyde dehydrogenase lactylation promotes acyl-CoA synthetase long-chain family member 4 propionylation and degradation, thereby suppressing ferroptosis induced by chemotherapy, and develop a blocking peptide that increased chemotherapy efficacy in pancreatic ductal adenocarcinoma.
  •  

Narrative review of the staging classification controversy in stage N3 small cell lung cancer: from the perspective of overlapping Veterans Administration Lung Study Group and International Association for the Study of Lung Cancer definitions

J Thorac Dis. 2026 Aug 31;18(8):950. doi: 10.21037/jtd-2026-1704. Epub 2026 Aug 28.

ABSTRACT

BACKGROUND AND OBJECTIVE: Traditionally, two primary systems have been employed for staging small cell lung cancer (SCLC): the Veterans Administration Lung Study Group (VALG) system and the International Association for the Study of Lung Cancer (IASLC) tumor, node, metastasis (TNM) system. The term "limited disease" is defined differently: VALG characterizes it as disease encompassed within a single tolerable radiation field, while IASLC defines it as the lack of distant metastases (M0). Patients with N3 disease frequently satisfy VALG extensive-stage (ES) criteria while meeting IASLC limited-stage (LS) criteria, resulting in a notable staging discrepancy. Therefore, this review aims to clarify the clinical challenges posed by this staging overlap and provide insights for standardizing staging terminology and optimizing therapeutic decision-making in N3 SCLC.

METHODS: A narrative review utilizing a systematized search strategy was conducted. While strict adherence to PRISMA guidelines was not pursued because the extensive heterogeneity of the literature precluded a formal meta-analysis, rigorous search criteria were applied to minimize selection bias. Databases including PubMed, Web of Science, Embase, the Cochrane Library, and China National Knowledge Infrastructure (CNKI) were searched for literature from January 2000 to March 2026. Studies examining stage N3 SCLC, spatial metastatic burden, and definitional inconsistencies between the VALG and IASLC staging systems were analyzed to assess their effects on treatment dosimetry, systemic therapy, and survival outcomes.

KEY CONTENT AND FINDINGS: The staging overlap in N3 SCLC leads to heterogeneous clinical management depending on its spatial metastatic burden, and this highly variable cohort can be stratified into distinct prognostic subgroups based on the anatomical distribution (single-region vs. multi-region) of the involved lymph nodes.

CONCLUSIONS: These findings should guide clinical trial design and terminology. Clinical decision-making must transcend historical paradigms and technical constraints. Future strategies must incorporate spatial evaluations of metastatic burden alongside innovative multimodal tools, such as artificial intelligence (AI) and multi-omics, to facilitate tailored therapy for SCLC.

PMID:42724560 | PMC:PMC13559235 | DOI:10.21037/jtd-2026-1704

  •  

A global digital navigator of human health for precision medicine

Nature Medicine, Published online: 11 September 2026; doi:10.1038/s41591-026-04621-1

The International Consortium of Digital Twins in Healthcare and Medicine was established to advance medical digital twin technology as a new infrastructure for precision health.
  •  

Joint impact of pathological burden and cognitive resilience on Alzheimer’s disease risk

Nature Medicine, Published online: 11 September 2026; doi:10.1038/s41591-026-04635-9

A 15-year cohort study shows that Alzheimer’s dementia risk is jointly shaped by Alzheimer’s pathology and cognitive resilience, with high resilience linked to lower risk, even under greater pathology.
  •  

Viral gene replication enhances AAV vector quality and reduces manufacturing costs

Liu and colleagues developed a robust in cellulo plasmid DNA replication system in human cells for replicating plasmid-borne adeno-associated virus (AAV) Rep/Cap genes during recombinant AAV (rAAV) production. This new approach not only enables a 10- to 20-fold plasmid reduction to significantly lower manufacturing costs but also substantially enhances rAAV potency, titer, and purity.
  •  

Lineage-specific pulmonary transcriptome landscape of coronavirus infection unveils universal immunotherapy for viral pneumonia

In the infection courses of different SARS-CoV-2 variants, disease outcomes and signatures were delineated by physiological changes, viral load, pathology, and pulmonary transcriptome analysis. This multi-dimensional landscape of disease outcomes and underlying mechanisms might provide important clues for immunotherapy of SARS-CoV-2 infection and pneumonia caused by other respiratory viruses.
  •  

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

arXiv:2609.09815v1 Announce Type: new Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decisions in an opaque, order-sensitive model call. We ask whether the manager needs to be generative at all. UnitBoost replaces that model with a defined meta-level operator: a task-given unit map turns worker outputs into slot-value proposals, a constrained argmax assembles the output, and the slots left unfilled or unsupported become an explicit residual for the next round. The operator is order-free, records unit provenance, and gives a simple guarantee: without coupling constraints, unit-wise maximization under the same admission score dominates selection of any complete candidate. On three held-out benchmarks, it exceeds the best single candidate chosen with gold labels by 0.060-0.195 absolute task-score points and input-matched generative managers by 0.048-0.076. Replacing only the management step improves six compound-system configurations by 0.013-0.182. Residual-directed rounds raise FanOutQA cell F1 from 0.4778 to 0.5524; matched controls show that the true residual outperforms random targets and ordinary rereading, while a label-free supply signal flags exhaustion after one unproductive round. The same analysis measures three conditions in which no such gain is available (one indivisible unit, unavailable unit identity, and an endpoint that charges for every emitted unit) and quantifies cross-unit coupling as a repair cost. The manager gives up semantic freedom and gains order invariance, unit provenance, and testable failure conditions.
  •  

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

arXiv:2609.09783v1 Announce Type: cross Abstract: Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.
  •  

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

arXiv:2609.10142v1 Announce Type: cross Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes them at evaluation time, yet the mechanism behind its lasting protection remains unclear. Analyzing its temporal optimization dynamics, we find that the defense emerges from an early compensatory adaptation phase followed by a steady-state phase where the corrective signal decays; in parameter space, attention output projections emerge as the dominant residual-write route for defensive updates. Through Intervention Delta Preservation (IDP) and IDP Continuation experiments, we further show that preserving or reinjecting the weight offset fails to maintain protection, indicating that preventative steering relies on active adaptation rather than a static defense. Motivated by this finding, we propose Progressive Intensity Scheduling (PIS), which starts with a moderate injection strength and increases it after static-strength alignment begins to decay. Across the evaluated Qwen2.5 and Gemma-3 models, PIS improves safety robustness over static-strength steering while reducing harmful trait expression.
  •  

KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?

arXiv:2607.27231v3 Announce Type: replace Abstract: Modern AI systems depend on specialized accelerator kernels, whose development is complicated by increasingly diverse operators and hardware. LLMs and agentic systems promise to automate this work, but existing evaluations do not show whether their performance transfers across operator sources and hardware platforms, or what such transfer costs. We present KernelGenBench, the first unified multi-source and multi-chip infrastructure for evaluating LLM- and agent-generated Triton kernels. With a common Triton target spanning six hardware platforms, it provides the broadest cross-vendor hardware coverage among existing kernel-generation benchmarks. We report two controlled analytical views: KernelGenBench-MS (Multi-Source) covers 210 operators from PyTorch ATen, production vLLM operators, and proprietary cuBLAS routines, while KernelGenBench-MC (Multi-Chip) evaluates a semantically stable 110-operator subset across six hardware platforms. Our evaluation consumed over 15 billion tokens. Agentic execution improved correctness, but no method dominated across sources and platforms: vLLM posed the strongest correctness challenge, cuBLAS set the highest performance ceiling, and AutoKernel accuracy fell from 87% on NVIDIA to 25% on Iluvatar CoreX. These improvements were costly: specialized agents averaged 4.99 million tokens per successful operator, rising to 6.25 million for CUDA Optimized Skill. The results establish operator source, hardware platform, and agentic scaffold as distinct dimensions of kernel-generation capability, and show that success in a familiar source-hardware setting is not a reliable proxy for deployment readiness.
  •  
❌