❌

Normal view

HOX code-based stratification reveals RUNX1T1-HDAC reprogramming as a targetable driver of lineage plasticity across cancers

Cancer Lett. 2026 Mar 28;648:218465. doi: 10.1016/j.canlet.2026.218465. Online ahead of print.

ABSTRACT

Cancer remains a leading cause of death worldwide, with lineage plasticity emerging as a hallmark that drives therapy resistance and tumor progression by enabling cancer cells to alter identity and evade targeted therapies. Although genomic and transcriptomic aberrations correlate with lineage plasticity, the absence of scalable cross-cancer markers to rapidly identify plastic subtypes has limited predictive utility. Homeobox (HOX) genes encode transcription factors that define tissue identity through distinct expression patterns, or HOX codes, within specific lineages. By analyzing multi-omics data encompassing 39 HOX genes across more than 80,000 RNA-seq samples across 23 cancer types spanning 114 cancer subtypes, we found that HOX code expression robustly stratifies lineage-constrained and lineage-plastic states at a cross-cancer level. This framework revealed previously unrecognized lineage-plastic subtypes in prostate cancer, lung cancer, and acute myeloid leukemia (AML), each displaying distinct HOX code divergence compared to non-plastic counterparts. Differential expression analysis across these representative malignancies identified RUNX1T1 as a consistent regulator associated with HOX-defined plastic states. We validated RUNX1T1 upregulation in bulk and single-cell RNA-seq from extensive preclinical and clinical cohorts and demonstrated that RUNX1T1 is functionally required for lineage-plastic programs in prostate cancer models. AI-based structural modeling and co-immunoprecipitation established the NCOR/HDAC3 complex as a critical binding partner of RUNX1T1. CUT&RUN profiling revealed that RUNX1T1 remodels chromatin by globally reducing active enhancer marks, thereby repressing lineage-defining differentiation programs and reshaping HOX positional identity. Selective pharmacologic inhibition of HDAC3 or targeted gene silencing via lipid nanoparticles suppressed the growth of lineage-plastic cancer cells, uncovering a therapeutically actionable vulnerability. Together, these findings establish RUNX1T1 as a cross-lineage regulator of HOX code-defined plasticity and identify the RUNX1T1-HDAC axis as a targetable mechanism underlying cancer lineage plasticity.

PMID:41912135 | DOI:10.1016/j.canlet.2026.218465

HOX Code-Based Stratification Reveals RUNX1T1-HDAC Reprogramming as a Targetable Driver of Lineage Plasticity Across Cancers

Cancer Lett. 2026 Mar 28:218465. doi: 10.1016/j.canlet.2026.218465. Online ahead of print.

ABSTRACT

Cancer remains a leading cause of death worldwide, with lineage plasticity emerging as a hallmark that drives therapy resistance and tumor progression by enabling cancer cells to alter identity and evade targeted therapies. Although genomic and transcriptomic aberrations correlate with lineage plasticity, the absence of scalable cross-cancer markers to rapidly identify plastic subtypes has limited predictive utility. Homeobox (HOX) genes encode transcription factors that define tissue identity through distinct expression patterns, or HOX codes, within specific lineages. By analyzing multi-omics data encompassing 39 HOX genes across more than 80,000 RNA-seq samples across 23 cancer types spanning 114 cancer subtypes, we found that HOX code expression robustly stratifies lineage-constrained and lineage-plastic states at a cross-cancer level. This framework revealed previously unrecognized lineage-plastic subtypes in prostate cancer, lung cancer, and acute myeloid leukemia (AML), each displaying distinct HOX code divergence compared to non-plastic counterparts. Differential expression analysis across these representative malignancies identified RUNX1T1 as a consistent regulator associated with HOX-defined plastic states. We validated RUNX1T1 upregulation in bulk and single-cell RNA-seq from extensive preclinical and clinical cohorts and demonstrated that RUNX1T1 is functionally required for lineage-plastic programs in prostate cancer models. AI-based structural modeling and co-immunoprecipitation established the NCOR/HDAC3 complex as a critical binding partner of RUNX1T1. CUT&RUN profiling revealed that RUNX1T1 remodels chromatin by globally reducing active enhancer marks, thereby repressing lineage-defining differentiation programs and reshaping HOX positional identity. Selective pharmacologic inhibition of HDAC3 or targeted gene silencing via lipid nanoparticles suppressed the growth of lineage-plastic cancer cells, uncovering a therapeutically actionable vulnerability. Together, these findings establish RUNX1T1 as a cross-lineage regulator of HOX code-defined plasticity and identify the RUNX1T1-HDAC axis as a targetable mechanism underlying cancer lineage plasticity.

PMID:41912135 | DOI:10.1016/j.canlet.2026.218465

CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News

arXiv:2603.22305v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly transitioning from static Natural Language Processing (NLP) tasks including sentiment analysis and event extraction to acting as dynamic decision-making agents in complex financial environments. However, the evolution of LLMs into autonomous financial agents faces a significant dilemma in evaluation paradigms. Direct live trading is irreproducible and prone to outcome bias by confounding luck with skill, whereas existing static benchmarks are often confined to entity-level stock picking and ignore broader market attention. To facilitate the rigorous analysis of these challenges, we introduce CN-Buzz2Portfolio, a reproducible benchmark grounded in the Chinese market that maps daily trending news to macro and sector asset allocation. Spanning a rolling horizon from 2024 to mid-2025, our dataset simulates a realistic public attention stream, requiring agents to distill investment logic from high-exposure narratives instead of pre-filtered entity news. We propose a Tri-Stage CPA Agent Workflow involving Compression, Perception, and Allocation to evaluate LLMs on broad asset classes such as Exchange Traded Funds (ETFs) rather than individual stocks, thereby reducing idiosyncratic volatility. Extensive experiments on nine LLMs reveal significant disparities in how models translate macro-level narratives into portfolio weights. This work provides new insights into the alignment between general reasoning and financial decision-making, and all data, codes, and experiments are released to promote sustainable financial agent research.

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

arXiv:2603.07980v1 Announce Type: cross Abstract: As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench \$OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, \$OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.

DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

arXiv:2603.08090v1 Announce Type: cross Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit critical limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across different subject difficulty levels and prompt scenarios, and 3) a profound lack of actionable insights and diagnostic guidance for subsequent model refinement. To address these limitations, we propose DSH-Bench, a comprehensive benchmark that enables systematic multi-perspective analysis of subject-driven T2I models through four principal innovations: 1) a hierarchical taxonomy sampling mechanism ensuring comprehensive subject representation across 58 fine-grained categories, 2) an innovative classification scheme categorizing both subject difficulty level and prompt scenario for granular capability assessment, 3) a novel Subject Identity Consistency Score (SICS) metric demonstrating a 9.4\% higher correlation with human evaluation compared to existing measures in quantifying subject preservation, and 4) a comprehensive set of diagnostic insights derived from the benchmark, offering critical guidance for optimizing future model training paradigms and data construction strategies. Through an extensive empirical evaluation of 19 leading models, DSH-Bench uncovers previously obscured limitations in current approaches, establishing concrete directions for future research and development.

Unveiling Downstream Performance Scaling of LLMs: A Clustering-Based Perspective

arXiv:2502.17262v4 Announce Type: replace-cross Abstract: The escalating scale and cost of Large Language Models (LLMs) training necessitate accurate pre-training prediction of downstream task performance for comprehensive understanding of scaling properties. This is challenged by: 1) the emergence phenomenon, where unpredictable capabilities appearing suddenly at critical model scales; and 2) uneven task difficulty and inconsistent performance scaling patterns, leading to high metric variability. Current prediction methods lack accuracy and reliability. We propose a Clustering-On-Difficulty (COD) framework for downstream performance prediction. The COD framework clusters tasks by their difficulty scaling features, thereby constructing a more stable and predictable task subset that exhibits well-behaved scaling characteristics with the increase of compute budget. We adopt a performance scaling law to predict cluster-wise performance with theoretical support. Predictable subset performance acts as an intermediate predictor for the full evaluation set. We further derive a mapping function to accurately extrapolate the performance of the subset to the full set. Applied to an LLM with 70B parameters, COD achieved a 1.55\% average prediction error across eight key LLM benchmarks, thus providing actionable insights for scaling properties and training monitoring during LLM pre-training.

Not All Candidates are Created Equal: A Heterogeneity-Aware Approach to Pre-ranking in Recommender Systems

arXiv:2603.03770v1 Announce Type: cross Abstract: Most large-scale recommender systems follow a multi-stage cascade of retrieval, pre-ranking, ranking, and re-ranking. A key challenge at the pre-ranking stage arises from the heterogeneity of training instances sampled from coarse-grained retrieval results, fine-grained ranking signals, and exposure feedback. Our analysis reveals that prevailing pre-ranking methods, which indiscriminately mix heterogeneous samples, suffer from gradient conflicts: hard samples dominate training while easy ones remain underutilized, leading to suboptimal performance. We further show that the common practice of uniformly scaling model complexity across all samples is inefficient, as it overspends computation on easy cases and slows training without proportional gains. To address these limitations, this paper presents Heterogeneity-Aware Adaptive Pre-ranking (HAP), a unified framework that mitigates gradient conflicts through conflict-sensitive sampling coupled with tailored loss design, while adaptively allocating computational budgets across candidates. Specifically, HAP disentangles easy and hard samples, directing each subset along dedicated optimization paths. Building on this separation, it first applies lightweight models to all candidates for efficient coverage, and further engages stronger models on the hard ones, maintaining accuracy while reducing cost. This approach not only improves pre-ranking effectiveness but also provides a practical perspective on scaling strategies in industrial recommender systems. HAP has been deployed in the Toutiao production system for 9 months, yielding up to 0.4% improvement in user app usage duration and 0.05% in active days, without additional computational cost. We also release a large-scale industrial hybrid-sample dataset to enable the systematic study of source-driven candidate heterogeneity in pre-ranking.

From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity

arXiv:2510.25232v2 Announce Type: replace Abstract: Psychiatric comorbidity is clinically significant yet challenging due to the complexity of multiple co-occurring disorders. To address this, we develop a novel approach integrating synthetic patient electronic medical record (EMR) construction and multi-agent diagnostic dialogue generation. We create 502 synthetic EMRs for common comorbid conditions using a pipeline that ensures clinical relevance and diversity. Our multi-agent framework transfers the clinical interview protocol into a hierarchical state machine and context tree, supporting over 130 diagnostic states while maintaining clinical standards. Through this rigorous process, we construct PsyCoTalk, the first large-scale dialogue dataset supporting comorbidity, containing 3,000 multi-turn diagnostic dialogues validated by psychiatrists. This dataset enhances diagnostic accuracy and treatment planning, offering a valuable resource for psychiatric comorbidity research. Compared to real-world clinical transcripts, PsyCoTalk exhibits high structural and linguistic fidelity in terms of dialogue length, token distribution, and diagnostic reasoning strategies. Licensed psychiatrists confirm the realism and diagnostic validity of the dialogues. This dataset enables the development and evaluation of models capable of multi-disorder psychiatric screening in a single conversational pass.

ROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems

arXiv:2602.01848v2 Announce Type: replace Abstract: Current agentic frameworks underperform on long-horizon tasks. As reasoning depth increases, sequential orchestration becomes brittle, context windows impose hard limits that degrade performance, and opaque execution traces make failures difficult to localize or debug. We introduce ROMA (Recursive Open Meta-Agents), a domain-agnostic framework that addresses these limitations through recursive task decomposition and structured aggregation. ROMA decomposes goals into dependency-aware subtask trees that can be executed in parallel, while aggregation compresses and validates intermediate results to control context growth. Our framework standardizes agent construction around four modular roles --Atomizer (which decides whether a task should be decomposed), Planner, Executor, and Aggregator -- which cleanly separate orchestration from model selection and enable transparent, hierarchical execution traces. This design supports heterogeneous multi-agent systems that mix models and tools according to cost, latency, and capability. To adapt ROMA to specific tasks without fine-tuning, we further introduce GEPA$+$, an improved Genetic-Pareto prompt proposer that searches over prompts within ROMA's component hierarchy while preserving interface contracts. We show that ROMA, combined with GEPA+, delivers leading system-level performance on reasoning and long-form generation benchmarks. On SEAL-0, which evaluates reasoning over conflicting web evidence, ROMA instantiated with GLM-4.6 improves accuracy by 9.9\% over Kimi-Researcher. On EQ-Bench, a long-form writing benchmark, ROMA enables DeepSeek-V3 to match the performance of leading closed-source models such as Claude Sonnet 4.5. Our results demonstrate that recursive, modular agent architectures can scale reasoning depth while remaining interpretable, flexible, and model-agnostic.
❌