❌

Normal view

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.

MCAT-mediated mitochondrial fatty acid metabolism regulates Lauren subtype divergence and suppresses gastric cancer progression through ROS/P53-dependent mitophagy and ferroptosis

Cell Death Differ. 2026 Sep 8. doi: 10.1038/s41418-026-01867-7. Online ahead of print.

ABSTRACT

Gastric cancer (GC) displays marked heterogeneity under the Lauren classification, yet the metabolic determinants of subtype divergence remain unclear. Here, we identify Malonyl-CoA:ACP transacylase (MCAT), a Lauren subtype-associated gene encoding a key mitochondrial fatty acid synthesis (mtFAS) enzyme, as a subtype-specific tumor suppressor in GC. Integrative multi-omics profiling revealed that MCAT expression is enriched in intestinal-type GC and correlates with favorable prognosis. Mechanistically, MCAT overexpression drives metabolic reprogramming through mitochondrial free fatty acid overload, suppressing β-oxidation while elevating mitochondrial reactive oxygen species (ROS), which triggers P53 phosphorylation at Ser15. This event concurrently activates PINK1/Parkin-mediated mitophagy and suppresses the SLC7A11/GPX4 axis to induce ferroptosis. Genetic rescue experiments confirmed that P53-Ser15 phosphorylation is essential for both mitophagy and ferroptosis induction. Endogenous MCAT levels are sufficient to determine basal ROS/P53/mitophagy/ferroptosis axis activity, and knockdown in high-expressing cells reverses these phenotypes, supporting a physiological, threshold-dependent role. In vivo, MCAT overexpression suppresses tumor growth and enhances mitophagy and ferroptosis markers. Collectively, these findings establish MCAT as a metabolic switch that links mtFAS to ROS/P53-dependent cell death, providing a potential biomarker and therapeutic target for GC.

PMID:42711380 | DOI:10.1038/s41418-026-01867-7

MCAT-mediated mitochondrial fatty acid metabolism regulates Lauren subtype divergence and suppresses gastric cancer progression through ROS/P53-dependent mitophagy and ferroptosis

Cell Death Differ. 2026 Sep 8. doi: 10.1038/s41418-026-01867-7. Online ahead of print.

ABSTRACT

Gastric cancer (GC) displays marked heterogeneity under the Lauren classification, yet the metabolic determinants of subtype divergence remain unclear. Here, we identify Malonyl-CoA:ACP transacylase (MCAT), a Lauren subtype-associated gene encoding a key mitochondrial fatty acid synthesis (mtFAS) enzyme, as a subtype-specific tumor suppressor in GC. Integrative multi-omics profiling revealed that MCAT expression is enriched in intestinal-type GC and correlates with favorable prognosis. Mechanistically, MCAT overexpression drives metabolic reprogramming through mitochondrial free fatty acid overload, suppressing β-oxidation while elevating mitochondrial reactive oxygen species (ROS), which triggers P53 phosphorylation at Ser15. This event concurrently activates PINK1/Parkin-mediated mitophagy and suppresses the SLC7A11/GPX4 axis to induce ferroptosis. Genetic rescue experiments confirmed that P53-Ser15 phosphorylation is essential for both mitophagy and ferroptosis induction. Endogenous MCAT levels are sufficient to determine basal ROS/P53/mitophagy/ferroptosis axis activity, and knockdown in high-expressing cells reverses these phenotypes, supporting a physiological, threshold-dependent role. In vivo, MCAT overexpression suppresses tumor growth and enhances mitophagy and ferroptosis markers. Collectively, these findings establish MCAT as a metabolic switch that links mtFAS to ROS/P53-dependent cell death, providing a potential biomarker and therapeutic target for GC.

PMID:42711380 | DOI:10.1038/s41418-026-01867-7

Knowledge Graph Modulated Deep Learning for Limited-Sample Clinical Data Analysis

arXiv:2605.24162v1 Announce Type: cross Abstract: Biological systems are governed by structured molecular interactions, where pathways, regulatory circuits, and functional gene relationships shape cellular behavior and disease progression. Much of this knowledge is naturally represented as graphs. However, most biomedical AI models cannot directly use graph-encoded biological knowledge and instead require compressed low-dimensional representations, which can lose important structure and reduce performance, especially in limited-sample clinical studies. Here, we introduce Graph-in-Graph (GiG), a knowledge graph-modulated deep learning framework for data-efficient clinical prediction. GiG represents each patient as a standalone modular graph, in which curated biological knowledge graphs define edges and patient-specific measurements, such as gene expression, define node features. This design allows multiple biological knowledge graphs to be integrated while preserving gene-gene interactions and pathway topology during patient-level representation learning. Across cohorts comprising nearly 9,700 patients and five clinical tasks, including liquid biopsy cancer detection, prostate cancer diagnosis, and 32-class pan-cancer classification, GiG consistently outperforms traditional and state-of-the-art methods, with the largest gains in limited-sample settings. On the challenging prostate cancer diagnosis task, GiG improves macro-F1 by up to 49 percentage points relative to competing methods. Control experiments replacing real pathway graphs with random topologies confirm that these gains arise from biologically grounded knowledge graph structure rather than graph modeling alone. These findings show that knowledge graph-modulated deep learning can improve robustness, interpretability, and sample efficiency in clinical data analysis, and provide a principled framework for integrating biological knowledge graphs into predictive modeling.

MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems

arXiv:2605.22794v2 Announce Type: replace Abstract: Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a fix. Self-evolving agents have emerged in response, but all confine evolution to text-mutable artifacts -- skill files, prompt configurations, memory schemas, workflow graphs -- and leave the agent harness untouched. Since routing, hook ordering, state invariants, and dispatch live in code rather than in any text artifact, an entire class of structural failure is physically unreachable from the text layer. We argue that source-level adaptation is a fundamentally more general medium: it is Turing-complete, a strict superset of every text-mutable scope, takes effect deterministically rather than through base-model compliance, and does not erode under long-context drift. We present MOSS, a system that performs self-rewriting at the source level on production agentic substrates. Each evolution is anchored to an automatically curated batch of production-failure evidence and proceeds through a deterministic multi-stage pipeline; code modification is delegated to a pluggable external coding-agent CLI while MOSS retains stage ordering and verdicts. Candidates are verified by replaying the batch against the candidate image in ephemeral trial workers, then promoted via user-consent-gated, in-place container swap with health-probe-gated rollback. On OpenClaw, MOSS lifts a four-task mean grader score from 0.25 to 0.61 in a single cycle without human intervention.

Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks

arXiv:2509.23067v2 Announce Type: replace-cross Abstract: The rising cost of acquiring supervised data has driven significant interest in self-improvement for large language models (LLMs). Straightforward unsupervised signals like majority voting have proven effective in generating pseudo-labels for verifiable tasks, while their applicability to unverifiable tasks (e.g., translation) is limited by the open-ended character of responses. As a result, self-evaluation mechanisms (e.g., self-judging and entropy minimization) are predominantly used to derive pseudo-labels. However, self-evaluation relying on LLMs typically incurs high computational overhead and introduces overconfidence issues due to intrinsic biases. To address these challenges, we propose a novel self-evaluation-free approach for unverifiable tasks, designed for lightweight yet effective self-improvement. Inspired by majority voting commonly employed in verifiable tasks, we propose semantic voting as a novel mechanism that relaxes the principle of hard matching (i.e., exact matching) toward soft matching (i.e., semantic similarity). Soft matching is achieved by leveraging a lightweight sentence embedding model to quantify semantic similarity, thereby mitigating excessive computational burden and intrinsic bias-associated limitations of self-evaluation. Comprehensive experiments demonstrate that our method achieves substantial gains in computational efficiency and overall better performance than self-evaluation methods across diverse model architectures and tasks.

Correction: The multifunctional RNA helicase DDX39A drives glioblastoma progression by modulating WISP1 alternative splicing that induces an immunosuppressive macrophage polarization

Oncogene, Published online: 01 April 2026; doi:10.1038/s41388-026-03756-2

Correction: The multifunctional RNA helicase DDX39A drives glioblastoma progression by modulating WISP1 alternative splicing that induces an immunosuppressive macrophage polarization

DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning

arXiv:2603.08095v1 Announce Type: cross Abstract: In scientific reasoning tasks, the veracity of the reasoning process is as critical as the final outcome. While Process Reward Models (PRMs) offer a solution to the coarse-grained supervision problems inherent in Outcome Reward Models (ORMs), their deployment is hindered by the prohibitive cost of obtaining expert-verified step-wise labels. This paper addresses the challenge of training reliable PRMs using abundant but noisy "weak" supervision. We argue that existing Weak-to-Strong Generalization (W2SG) theories lack prescriptive guidelines for selecting high-quality training signals from noisy data. To bridge this gap, we introduce the Dual-Consensus Weak-to-Strong (DC-W2S) framework. By intersecting Self-Consensus (SC) metrics among weak supervisors with Neighborhood-Consensus (NC) metrics in the embedding space, we stratify supervision signals into distinct reliability regimes. We then employ a curriculum of instance-level balanced sampling and label-level reliability-aware masking to guide the training process. We demonstrate that DC-W2S enables the training of robust PRMs for complex reasoning without exhaustive expert annotation, proving that strategic data curation is more effective than indiscriminate training on large-scale noisy datasets.
❌