❌

Normal view

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.

Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

arXiv:2609.12373v1 Announce Type: new Abstract: Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents

arXiv:2609.12394v1 Announce Type: new Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a distribution mismatch with production environments; expensive real-device failures remain underutilized; and fixed benchmarks saturate, losing the power to guide iteration. We present BlueLM-GUI, a 35B-A3B mobile GUI agent built as a real-device-centric flywheel that closes these gaps through three principles. Every Sample Matters: a dual-track pipeline with Heterogeneous Triple-System Consensus evaluation and an Error Correction \& Derivation Module salvages every trajectory into usable supervision. Every Rollout Is Real: a three-stage recipe---continual pre-training, supervised fine-tuning, and agentic reinforcement learning on hundreds of real phones---grounds every rollout in real production environments, so the capability the model learns transfers directly to deployment. Every Query Evolves: a quota-driven benchmark methodology with three orthogonal axes enables precise attribution and allows the benchmark to be systematically upgraded as the model improves. BlueLM-GUI achieves 87.4 on MobileGUI-VBench, surpassing the best closed-source model by 5.1 points, and 84.9 on AndroidWorld, the best result among open-source models and competitive with closed-source models. These results demonstrate that grounding model training and iterative improvement in both real devices and the three Every principles yields strong, robust, and transferable mobile GUI capability.

Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?

arXiv:2609.12437v1 Announce Type: cross Abstract: Adaptive RAG often uses retrieval-time signals to decide whether another retrieval, reranking, or multimodal step should run. We ask whether these signals add routing value once the query itself is already known. Across document, audio, and video RAG, we compare matched query-only and query+retrieval routers while holding the optional actions, router family, training procedure, and evaluation fixed. On the held-out final evaluation, adding the tested retrieval signals does not produce a reliable routing improvement over the query-only baseline. Some retrieval signals are associated with whether a later step will help, but that predictability does not consistently lead to bet- ter RUN/SKIP decisions. The main lesson is therefore methodological: retrieval-state features should not be credited with routing value unless they improve over a matched query-only control. Our results do not show that routing or retrieval state is generally useless; they show that the incremental value of retrieval signals must be demonstrated rather than assumed.

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

arXiv:2508.16129v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities under reinforcement learning (RL) paradigm. However, most existing multimodal medical reasoning models focus on basic reasoning, which refers to shallow inference based on visual feature matching. In contrast, real-world clinical diagnosis extends beyond basic reasoning, demanding complex reasoning that integrates heterogeneous clinical information (such as chief complaints and medical history) with multimodal medical imaging data. To bridge this gap, we introduce MM-Retinal-Reason, an ophthalmic multimodal dataset covering the full spectrum of perception and reasoning. Specifically, it is the first dataset in ophthalmology to encompass both basic and complex reasoning tasks with Chain-of-Thought (CoT) trajectories, aiming to enhance visual-centric reasoning and emulate realistic clinical decision-making. Building upon MM-Retinal-Reason, we propose OphthaReason, the first RL-enhanced ophthalmic multimodal reasoning model with step-by-step reasoning traces. To enable flexible adaptation to both basic and complex reasoning tasks, we further introduce Uncertainty-Aware Dynamic Thinking (UADT), which estimates sample-level uncertainty via entropy and dynamically modulates exploration depth through a shaped advantage mechanism. Comprehensive experiments demonstrate the effectiveness of our model on both basic and complex reasoning tasks, outperforming general-purpose MLLMs, medical MLLMs, RL-based medical MLLMs, and ophthalmic MLLMs by at least 15.47\%. Project Page: \href{https://github.com/lxirich/OphthaReason}{link}.

Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval

arXiv:2510.20486v2 Announce Type: replace-cross Abstract: Imbalanced labels can cause frequent samples to dominate AI-based quantitative remote sensing, degrading rare-event retrieval. In rain-rate retrieval based on satellite infrared brightness temperatures, this imbalance leads to systematic underestimation of rare high-intensity rainfall. In this study, Hurdle-Retrieval Model Imbalanced Learning (RMIL) is proposed. Following a divide-and-conquer strategy, Hurdle-RMIL separates zero inflation from the long-tailed distribution of positive rain. A hurdle model handles zero inflation, whereas RMIL exploits invariance under fixed observation conditions of the rainfall-to-satellite forward process to derive a Bayes-based transformation linking conditional distributions under naturally long-tailed and hypothetical balanced rainfall. This transformation enables the balanced-distribution model to be learned from natural samples without constructing a balanced dataset. Comparisons with conventional learning, classification-regression modeling, cost-sensitive learning, and generative learning using test data from multiple regions in China show that Hurdle-RMIL mitigates systematic underestimation and improves detection of rare high-intensity and extreme rainfall without markedly degrading lower-threshold accuracy. At 0.1-10 mm per hour, its root mean square error remains close to those of the best baselines, and it yields the highest equitable threat score (ETS) at most evaluated thresholds, with its advantage becoming more pronounced at high thresholds. At 30 mm per hour, its ETS is 0.051 versus 0.015 for the best baseline, and its mean error is -25.41 mm per hour versus -28.98 mm per hour. Case studies further show improved representations of rainfall intensity and spatial extent, demonstrating that Hurdle-RMIL effectively addresses rainfall-distribution imbalance and improves the retrieval of rare high-intensity rainfall.

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

arXiv:2608.26105v2 Announce Type: replace-cross Abstract: Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

An engineered nanopore identifies saccharides, amino acids, peptides and ribonucleotides

Nature Biotechnology, Published online: 14 September 2026; doi:10.1038/s41587-026-03308-9

Modified nanopore simultaneously identifies diverse biomolecules and their modifications.

Key Experimental Therapeutics and Knowledge Gaps in Metabolic Dysfunction-Associated Steatohepatitis (MASH)

10 September 2026 at 18:00

Drug Des Devel Ther. 2026 Sep 5;20:543657. doi: 10.2147/DDDT.S543657. eCollection 2026.

ABSTRACT

Metabolic dysfunction-associated steatohepatitis (MASH) is not solely a disorder of hepatocellular lipid accumulation, but a multicellular disease driven by coordinated metabolic stress, sterile inflammation, fibrogenesis, and niche remodeling. Recent therapeutic progress with the provisional approval of resmetirom and semaglutide has validated MASH as a tractable clinical target. However, many experimental agents have shown limited or inconsistent efficacy, particularly for regression of hepatic fibrosis or cirrhosis, reflecting the biological heterogeneity and dynamic cellular architecture of the disease. Distinct from conventional pathway- or drug class-based reviews, we summarize emerging therapeutics through a liver cell-centered framework, integrating hepatocyte-directed metabolic therapies, immune-cell modulation, hepatic stellate cell-targeted antifibrotic strategies, niche-directed approaches involving liver sinusoidal endothelial cells and cholangiocytes, systemic multi-cell modulators, and precision-delivery technologies. We further compare how these interventions reshape pathogenic communication among hepatic and extrahepatic compartments, while emphasizing unresolved challenges in drug target selection, cellular specificity, disease-stage dependency, safety, and patient stratification. This perspective emphasizes the need to move from isolated pathway targeting toward cell- and network-informed therapeutic strategies supported by spatial multi-omics, human-relevant models, and precision delivery.

PMID:42719321 | PMC:PMC13557022 | DOI:10.2147/DDDT.S543657

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review

npj Digital Medicine, Published online: 10 September 2026; doi:10.1038/s41746-026-03228-7

Impact of LLM-supported patient education on patient perspectives and patient-reported outcomes: a mixed-methods systematic review

Transduction Efficiency in Clinical CAR T-Cell Products: A Retrospective Study at a Single Center

Transduction efficiency is a critical determinant of CAR T-cell manufacturing quality. Analysis of 204 clinical CAR T-cell products revealed that transduction efficiency is shaped primarily by manufacturing workflows and protocol-dependent starting material composition. Higher transduction efficiency was associated with early memory-like cellular states, providing insights into optimizing CAR T-cell.

Ammonium tetrathiomolybdate improves auditory and vestibular function after gentamicin exposure via the NRF2–GPX4 axis

Zhang and colleagues reveal that GPX4 serves as a critical regulator of NRF2-mediated otoprotection against aminoglycoside-induced hair cell injury. Their findings identify a GPX4-dependent antioxidant mechanism that enables therapeutic activation of NRF2 and provides new insights into strategies for preventing drug-induced hearing loss.

MITF-SCD1 Lipid Metabolic Axis Prevents Ouabain-Induced Spiral Ganglion Neuron Ferroptosis and Hearing Loss

Ouabain triggers cochlear spiral ganglion neuron (SGN) ferroptosis and hearing loss via SCD1 downregulation. MITF directly activates Scd1 transcription, and the MITF–SCD1 axis mitigates SGN ferroptosis and hearing impairment in ototoxic ouabain and cisplatin models, revealing a lipid metabolic vulnerability and therapeutic target for sensorineural hearing loss.

Antisense oligonucleotides against Il6ra ameliorate cancer cachexia in mice

Cancer cachexia is a devastating metabolic syndrome for which there are no approved treatments. Li and colleagues developed an RNA-targeted therapy, which ameliorates cachectic symptoms, reduces inflammation, and extends survival in mouse cancer models. The study provides an approach for treating cancer cachexia and paves the road for clinical studies.

A complement C5-targeted GalNAc-conjugated siRNA with sustained efficacy in a non-human primate model of IgA nephropathy

This study characterizes a GalNAc-C5 small interfering RNA with potent in vitro and in vivo activity. Single subcutaneous dosing sustains long-term C5 suppression in cynomolgus monkeys with IgA nephropathy, outperforming Nefecon in blocking glomerular complement deposition, supporting its standalone or combinational clinical application.

The DreAM-plus integrative RNA switch enhances transient AAV expression and reduces side effects of gene editing

This study developed a multi-layer inducible RNA switch that achieves transient expression of gene-delivery vectors in hepatic and non-hepatic tissues. As an exemplary application, this RNA switch triggers pulsive expression of gene editors that reduces the off-target effects and immunotoxicity of gene editing.

A helicase-fused Cas9 improves large-size fragment knockin

By fusing MCM5, a subunit of the eukaryotic MCM2–7 helicase complex, to the N terminus of spCas9 (MCCas), the MCCas fusion protein enhances large-size fragment knockin via homologous recombination, reduces insertions and deletions (indels), and enables efficient large-size fragment insertions in human cells and rabbit embryos.

Talking to Itself While Coding: What Makes Comments Help Code Generation?

arXiv:2609.09242v1 Announce Type: cross Abstract: Large Language Models (LLMs) often generate natural-language comments while writing code, and these comments become part of the context used to generate the code that follows. However, it remains unclear which properties of comments affect code-generation performance. We study this question through observational analyses and controlled interventions. On LiveCodeBench, neither comment frequency nor broad comment intent reliably predicts pass@1. We then prefill weaker recipient models with comment blocks written by stronger source models, allowing us to separate comment surface form from the solution content they convey. Comments from source solutions that pass the tests raise recipient pass@1 by 17.2% on average. In contrast, comments describing failed solutions provide no reliable gain, while comments written for a different problem reduce pass@1 by 20.8%. Finally, across a wide range of models and prompt variants, most recipient models show no significant recovery of the external-comment gain, and the best case recovers only 24%. These results show that comments help code generation not merely because they are comments, but because they can provide correct solution content that prompting cannot reliably elicit.

HiRAD: A Flexible Large-Scale AGV Routing System

arXiv:2609.09752v1 Announce Type: cross Abstract: Automatic Guided Vehicles (AGVs) substantially boost warehouse throughput, but routing large-scale AGV fleets remains challenging. Classical Multi-Agent Pathfinding solvers suffer from exploding combinatorial complexity and super-quadratic runtime, while relying on idealized grid or piecewise-linear motion models that mismatch real-world kinematics. Recent Reinforcement Learning (RL) solutions improve flexibility via decentralized agent policies but depend on discretized spatiotemporal representations, require millions of episodes to converge, and incur full-map observation at every step, which leads to large models, slow convergence, and high inference latency that violates real-time industrial control constraints. To address these bottlenecks, we propose HiRAD, a hierarchical RL framework for continuous-space AGV routing with real-time guarantees: (1) a step-level spatiotemporal representation that translates continuous motion into a differentiable RL problem, (2) a hierarchical strategy that splits heading choice from velocity control to reduce the action space, and (3) an asynchronous event-driven decision pipeline that lowers inference complexity from O(n^2) to O(n) and cuts per-step latency by as much as 71 percent. Across random graphs and two warehouse maps, HiRAD reduces makespan by 45 percent to 63 percent and shortens end-to-end runtime.
❌