❌

Reading view

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.
  •  

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

arXiv:2507.01663v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.
  •  

Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

arXiv:2511.15015v4 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has become a practical architecture for scaling LLM capacity while keeping per-token compute modest, but deploying MoE models on a single, memory-limited GPU remains difficult because expert weights dominate the HBM footprint. Existing expert offloading and prefetching systems reduce the resident set, yet they often pay expert-loading costs on the critical path when activation becomes dense. Post-training quantization (PTQ) lowers the footprint without transfers, but prevailing pipelines fix expert bit-widths offline and assume routing remains stable, even though MoE expert utilization is heavy-tailed and the hot set can shift across workloads. We present DynaExq, a runtime-aware mixed-precision serving system that treats single-GPU MoE inference under a hard HBM envelope as an online, budget-constrained precision allocation problem. The key insight is to keep the experts that dominate runtime traffic resident at higher precision, while maintaining a low-precision fallback for the remaining experts, so the system can reduce transfer volume and avoid the waiting latency that limits offloading and prefetching under dense activation. DynaExq estimates long-horizon expert hotness from router traces, selects a per-layer high-precision resident set via a budget-feasible top-$n$ rule, and applies promotions and demotions asynchronously through stable expert handles so the forward pass always executes on a fully materialized expert version. Across Qwen3-MoE-30B/80B and six benchmarks, DynaExq improves accuracy over static PTQ on Qwen3-80B (73.09% to 77.57%) under comparable device-memory budgets and achieves up to 2.73x higher throughput than offloading/prefetch baselines at batch size 32.
  •  

An engineered nanopore identifies saccharides, amino acids, peptides and ribonucleotides

Nature Biotechnology, Published online: 14 September 2026; doi:10.1038/s41587-026-03308-9

Modified nanopore simultaneously identifies diverse biomolecules and their modifications.
  •  

Advanced and underlying therapeutic strategies in transformed small cell lung cancer

Front Med (Lausanne). 2026 Aug 27;13:1865050. doi: 10.3389/fmed.2026.1865050. eCollection 2026.

ABSTRACT

Transformed small-cell lung cancer (T-SCLC) is a clinically important form of histologic transformation and a mechanism of acquired resistance in non-small-cell lung cancer (NSCLC). It is associated with poor prognosis, with a median overall survival of only about 9-13 months. This review summarizes recent advances in the mechanisms, diagnosis, monitoring, and treatment of T-SCLC. Repeat biopsy remains the gold standard for confirming histologic transformation, whereas molecular profiling and liquid biopsy may facilitate early detection and longitudinal disease monitoring. Platinum-etoposide remains the most commonly used clinical standard after transformation, but its benefit is typically transient and durable disease control remains uncommon. Continuation of EGFR tyrosine kinase inhibitors combined with chemotherapy may prolong progression-free survival in selected patients but has not consistently improved overall survival. Anti-angiogenic therapy, particularly anlotinib, and chemo-immunotherapy have shown encouraging activity in selected patients, while emerging strategies targeting DLL3, MYC, SOX2, and epigenetic regulators may broaden the therapeutic landscape. Prospective studies integrating repeat tissue sampling, comprehensive genomic profiling, biomarker-guided patient stratification, pharmacogenomics, functional drug-sensitivity testing where feasible, and integrated multi-omics approaches are needed to advance molecularly guided and individualized treatment for T-SCLC.

PMID:42724635 | PMC:PMC13560167 | DOI:10.3389/fmed.2026.1865050

  •  

Advanced and underlying therapeutic strategies in transformed small cell lung cancer

Front Med (Lausanne). 2026 Aug 27;13:1865050. doi: 10.3389/fmed.2026.1865050. eCollection 2026.

ABSTRACT

Transformed small-cell lung cancer (T-SCLC) is a clinically important form of histologic transformation and a mechanism of acquired resistance in non-small-cell lung cancer (NSCLC). It is associated with poor prognosis, with a median overall survival of only about 9-13 months. This review summarizes recent advances in the mechanisms, diagnosis, monitoring, and treatment of T-SCLC. Repeat biopsy remains the gold standard for confirming histologic transformation, whereas molecular profiling and liquid biopsy may facilitate early detection and longitudinal disease monitoring. Platinum-etoposide remains the most commonly used clinical standard after transformation, but its benefit is typically transient and durable disease control remains uncommon. Continuation of EGFR tyrosine kinase inhibitors combined with chemotherapy may prolong progression-free survival in selected patients but has not consistently improved overall survival. Anti-angiogenic therapy, particularly anlotinib, and chemo-immunotherapy have shown encouraging activity in selected patients, while emerging strategies targeting DLL3, MYC, SOX2, and epigenetic regulators may broaden the therapeutic landscape. Prospective studies integrating repeat tissue sampling, comprehensive genomic profiling, biomarker-guided patient stratification, pharmacogenomics, functional drug-sensitivity testing where feasible, and integrated multi-omics approaches are needed to advance molecularly guided and individualized treatment for T-SCLC.

PMID:42724635 | PMC:PMC13560167 | DOI:10.3389/fmed.2026.1865050

  •  

Baseline cellular state shapes the molecular impact of mutant KRAS alleles in reconstituted pancreatic cancer cells

Mol Omics. 2026 Sep 10:aaiag022. doi: 10.1093/molecular-omics/aaiag022. Online ahead of print.

ABSTRACT

KRAS is mutated in over 90% of pancreatic ductal adenocarcinomas (PDAC), where hotspot alterations in codons 12, 13, and 61 drive tumor initiation and progression. Although distinct biochemical properties have been described for individual KRAS mutants, whether they generate unique allele-specific signaling programs in PDAC cells remains unresolved. Here, we systematically interrogated the molecular consequences of seven common KRAS mutant variants in reconstituted isogenic, KRAS-deficient PDAC cell lines by integrated transcriptomic, proteomic, and phosphoproteomic profiling. We found that baseline cellular state, rather than allele identity, was the predominant driver of molecular variation. Comparisons with established KRAS reference signatures revealed significant but moderate overlap at the mRNA level and less so at the proteome level. Pathway analyses highlighted interferon response and mitochondrial translation-related proteins as recurrently altered across mutant alleles, while phosphoproteomic data confirmed robust ERK1/2 activity and suppression of DYRK kinase substrates by mutant KRAS expression. Importantly, no robust mutant allele-specific molecular programs were identified in our KRAS-reconstituted cell lines. Together, our study establishes a comprehensive multi-omics resource for KRAS signaling in PDAC and demonstrates that cellular context exerts a stronger influence than allele identity in shaping molecular profiles, with implications for interpreting putative allele-specific signaling dependencies.

PMID:42720273 | DOI:10.1093/molecular-omics/aaiag022

  •  

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review

npj Digital Medicine, Published online: 10 September 2026; doi:10.1038/s41746-026-03228-7

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review
  •  

Targeting the MNK1-MYH9 axis blocks YAP1 recruitment to prevent thrombosis and platelet activation-induced NETosis

MNK1 acts as a structural shield on MYH9, preventing YAP1-mediated platelet activation. Developing MD2 to lock this MNK1-MYH9 complex introduces a safe antithrombotic strategy, shifting the therapeutic paradigm from kinase inhibition to stabilizing protein-protein interactions against immunothrombosis.
  •  

Switchable single-atom catalysts for highly selective C–C coupling in direct methane oxidation

Nature Nanotechnology, Published online: 07 September 2026; doi:10.1038/s41565-026-02271-5

Single copper atoms on boron nanosheets dynamically and reversibly switch to clusters, enabling the direct conversion of methane to acetic acid with 97% selectivity and high activity without the requirement for carbon monoxide.
  •  

PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving

arXiv:2609.10372v1 Announce Type: cross Abstract: We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived Time-to-First-Response (PTFR) as a QoE objective and minimizes it under quality/cost constraints. Unlike prior work on cascaded routing, semantic caching, or adaptive retrieval, PACE jointly controls which answer source composes the response and what fills the waiting window. Deployed on a humanoid-robot sales service, it combines three mechanisms: a load-adaptive cascading router, a joint path-filler controller, and volatility-aware cache admission. On 75k CarQA requests, the cascade halves pure-LLM PTFR at P95 (0.29 vs 0.53s at c16). The adaptive controller reaches 0.41s P95, outperforming RAG by 2.4 times at high load with equal quality. The filler controller cuts calls by 94% with zero conflict. Volatility-aware admission reduces stale answers from 86% to 0%. A gating rule ensures the controller never worse than the baseline, with exposure bounded by one hold period. This is the first quantification of filler-answer conflict risk in deployed services.
  •  

EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering

arXiv:2609.08435v2 Announce Type: replace Abstract: In persistent interactions, long contexts may encode an evolving process rather than a fixed record: later events can revise or revoke earlier information, changing what remains valid and what conclusions follow. We call this setting information evolution (IE). Solving IE requires identifying valid records, applying updates in order, and reconstructing the query-relevant state from the event history. Existing text-first synthesis pipelines make such data difficult to verify because state transitions and answer logic remain implicit. We introduce EvolveScaler, a code-driven framework that defines information evolution before rendering it as natural language. Human-authored operational specifications define state transitions, record validity, difficulty controls, and executable answer logic; a strong LLM then synthesizes a self-contained simulator from each specification. Executing validated simulators produces natural-language multi-turn event histories, while deterministic replay computes reference answers and atomic checklists. We instantiate EvolveScaler with 117 task prototypes and 159 final-question operators across five difficulty levels spanning approximately 7 to 1,200 events per instance, yielding about 35,100 training examples and 585 validated evaluation instances. On the very_long tier, the strongest model reaches 59.3% avg@5, while six models score below 10%. Training an internal A3B model on 6,000 EvolveScaler examples improves performance over its base checkpoint on all eight independently constructed out-of-distribution benchmarks, with a 5.25-point average gain. These results show that code-driven IE synthesis provides both challenging evaluation and transferable training supervision.
  •  

The RNA-binding protein La/SSB is associated with HNSCC progression and TFAP2C/FSCN1-linked transcriptional regulation

Oncogene, Published online: 03 September 2026; doi:10.1038/s41388-026-03959-7

The RNA-binding protein La/SSB is associated with HNSCC progression and TFAP2C/FSCN1-linked transcriptional regulation
  •  

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

arXiv:2605.23945v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) has become a key post-training paradigm for improving model quality. However, the synchronous three-stage RLHF pipeline is often bottlenecked by the generation stage, where response-length skew causes the effective batch size to shrink rapidly during decoding, leaving GPUs underutilized while a few long responses remain unfinished. Mainstream frameworks employ a static tensor parallelism (TP) configuration that cannot adapt to changing batch characteristics, leaving substantial performance headroom unexplored. We propose PAT, an adaptive TP method that dynamically reconfigures TP during the generation stage of each RLHF iteration. PAT introduces two key techniques. First, a predictor-guided online reconfiguration method decides both the reconfiguration point and the target TP configuration based on offline profiling, triggering reconfiguration only when the predicted latency benefit outweighs the reconfiguration overhead. Second, a lightweight online reconfiguration mechanism updates only the states and layouts affected by TP changes: it adapts unfinished decoding states through a cost-model-based choice between KV-cache migration and recomputation, performs in-place weight resharding, and reuses cached communication groups. We implement PAT on top of SGLang and integrate it with the VeRL framework. Evaluations on LLaMA3.1-8B and Qwen3-14B using DeepScaleR show that PAT reduces generation latency by up to 34.6% and end-to-end RLHF training iteration latency by up to 27.2% compared to the original VeRL setup.
  •  

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

arXiv:2605.24468v1 Announce Type: new Abstract: Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and partial conclusions. The challenge is not merely that these histories grow long, but that information needed for the current decision may be scattered across distant steps and only become relevant later. Existing approaches address this difficulty by truncating the interaction history, compressing it into shorter surrogates, or retrieving selected parts of it for reuse, but they do not explicitly model how access to past interaction should adapt to the agent's evolving state. We instead cast long-horizon reasoning as a problem of state-adaptive memory. To this end, we propose State-Adaptive Memory~(SAM), a standalone framework that consolidates ongoing interaction into compact memory cues while preserving raw trajectory pages for intent-driven recall. These cues are not treated as replacements for history; rather, they serve as lightweight handles that allow the agent to reconstruct temporally distant information according to its current needs, without retraining the underlying backbone. We further optimize the memory module through expert-guided supervision and reinforcement learning, aligning it with trajectory-level utility. Across BrowseComp, BrowseComp-ZH, WideSearch, and HLE, SAM consistently outperforms strong baselines over diverse agent backbones. Our results suggest that explicit memory modeling provides a simple and effective foundation for long-horizon agentic reasoning.
  •  

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

arXiv:2605.24486v1 Announce Type: new Abstract: Recent progress on long-horizon agentic tasks has been driven largely by scaling up individual agents through stronger models, better tools, and more effective scaffolding. In contrast, much less is understood about scaling out: whether multiple peer agents, all targeting the same task, can become an additional source of capability without relying on explicit role specialization or workflow orchestration. We study this question and propose AgentFugue, a collective reasoning framework built around a shared reasoning hub. As peer agents explore the same task in parallel, the hub records concise notes on what each agent has established, attempted, or ruled out, and enables each agent to selectively access what other agents have discovered in a form useful for its current search. This design turns otherwise isolated trajectories into a connected ecology of reusable intermediate reasoning without requiring centralized planning. We instantiate the hub as a plug-in communication layer, trained with supervised fine-tuning and end-to-end reinforcement learning. Across the challenging long-horizon settings we study, AgentFugue improves over strong baselines. Our results suggest that collective reasoning can turn scaling out peer agent systems into a distinct source of capability gains, rather than merely a way of spending more compute.
  •  

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-world deployment faces a fundamental device--cloud dilemma: on-device models are efficient but often brittle, while cloud models are stronger but costly in computation. State-of-the-art LLM device--cloud routers usually make coarse task-level decisions, which cannot adapt to the changing difficulty of multi-step agent interactions. To address this issue, we present Hera, a step-level device--cloud LLM agent coordinator for long-horizon tasks achieving a strong performance--cost Pareto frontier. Hera adopts a novel two-stage training paradigm: (1) imitation learning for cold-start, followed by (2) reinforcement learning that jointly optimizes task success and cloud usage efficiency. The first stage casts step-level routing as a supervised classification problem: the device agent is replayed on cloud trajectories, with each state labeled by the agreement between device and cloud actions. In the second stage, we perform cost-aware reinforcement learning by grouping identical states across trajectories and updating Hera with labels favoring higher expected return and fewer future cloud calls. We evaluate Hera on ALFWorld, WebShop, and AppWorld, where it consistently outperforms prior methods, achieving 92.5% of the cloud-only success rate with cloud use in only 46.3% of steps.
  •  

Clustering as Reasoning: A $k$-Means Interpretation of Chain-of-Thought Graph Learning

arXiv:2605.24867v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has shown promise in enhancing the reasoning capabilities of large language models (LLMs) on text-attributed graphs (TAGs). This work reframes CoT-based graph learning through the principle of clustering as reasoning, offering a $k$-means interpretation of how iterative reasoning operates over graph-structured data. We observe that existing graph CoT methods rely on disjoint architectures and fixed graph representations, limiting step-by-step semantic-topological interaction and interpretability. To overcome this limitation, we propose a unified framework named KCoT that integrates CoT reasoning with graph representation learning. Our key theoretical result reveals a formal mathematical correspondence between a Transformer block and the $k$-means algorithm, allowing reasoning to be interpreted as iterative assignment and update steps. Based on this insight, we introduce a Semantic Discriminating Prompt that explicitly formulates these steps as structured CoT reasoning, together with a structure-grounded alignment strategy to fuse topological priors with evolving thought-conditioned representations. Experiments on standard benchmarks demonstrate consistent improvements over state-of-the-art methods, validating clustering as a principled mechanism for CoT-based graph learning.
  •  

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for CUAs requires consistent task instruction, executable environment, and verifiable reward. However, hand-curated benchmarks achieve high reward fidelity but cover few applications and LLM-as-judge-based datasets scale broadly but lack reliable verification. We present CUA-Gym, a scalable pipeline that co-generates task instructions, environment states, and reward functions. Concretely, a Generator agent constructs the initial and golden environment states, and a separate Discriminator agent writes the reward function from the task specification. An orchestrator agent drives the two through iterative rounds upon execution. Generated tuples then pass a final filter combining LLM majority voting and agent rollouts, ensuring quality beyond the per-task adversarial loop. To address the scarcity of training environments, we further synthesize CUA-Gym-Hub, a broad suite of high-fidelity mock web applications grounded in real-world software-use distributions, expanding the scale of CUA RLVR data by magnitude. Using this pipeline, we construct CUA-Gym, a dataset of 32,112 verified RLVR training tuples grounded in 110 environments. Trained with GSPO on CUA-Gym, our CUA-Gym-A3B and CUA-Gym-A17B achieve 62.1% and 72.6% on OSWorld-Verified, outperforming prior open-source CUAs at comparable scales, with performance scaling smoothly in both data volume and environment diversity. The same checkpoints also improve on the held-out WebArena benchmark, indicating transfer beyond the training environments. We will open-source the full synthesis pipeline, dataset, CUA-Gym-Hub environments, and models.
  •  

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

arXiv:2605.25475v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows linearly with sequence length, quickly becoming the bottleneck for long context inference. A practical remedy is to evict less important KV entries; however, existing eviction policies are largely heuristic and struggle to capture the rich, input-dependent distribution of token importance. In this work, we introduce a learnable indexer that predicts KV importance, enabling more accurate retention of critical tokens. Meanwhile, naively evicting tokens permanently discards their information, leading to irreversible forgetting and degraded retrieval over long ranges. To address this, we propose a lightweight latent memory module that compresses evicted tokens into a compact, online-updated state and provides residual readouts to compensate for the attention contributions lost through KV eviction. Collectively, our method enables accurate long-context inference under a bounded KV budget, delivering consistent improvements on RULER (4K/16K) across Qwen, Mistral, and Llama models (up to 25 points under aggressive eviction), markedly more stable Needle-in-a-Haystack retrieval, and superior LongBench scores and compression curves compared to existing eviction policies.
  •  
❌