❌

Normal view

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

arXiv:2609.11977v1 Announce Type: new Abstract: Co-work agents execute complex workflows that combine information gathering, tool use, coding, and file manipulation across many model invocations. Because cost and latency accumulate over the full episode, their practical value depends not only on peak capability but also on how efficiently that capability is delivered. Yet many steps in everyday work emphasize state tracking, coordination, recovery, and follow-through rather than frontier-scale reasoning. We present Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint. We construct execution-grounded data and environments, capture replayable long-horizon trajectories across multiple harnesses, and use staged post-training to develop and consolidate complementary execution capabilities. Across a broad suite of co-work benchmarks, Occamy-1.0 is consistently among the strongest comparably sized models and remains competitive with substantially larger frontier systems on several tasks. Under our stated evaluation and pricing protocol, its aggregate performance across four representative benchmarks places it at the low-cost knee of the observed cost--performance Pareto frontier. Supporting evaluations in tool calling, coding, and instruction following further show that this specialization preserves broad agentic capability. We release the model weights and a subset of the training data to support research on practical co-work agents and agentic post-training.

Unified Text-Image Generation with Weakness-Targeted Post-Training

arXiv:2601.04339v3 Announce Type: replace-cross Abstract: Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many existing systems rely on explicit modality switching, generating reasoning text before switching manually to image generation. This separate, sequential inference process limits cross-modal coupling and prohibits automatic multimodal generation. This work explores post-training to achieve fully unified text-image generation, where a model autonomously transitions from textual reasoning to visual synthesis within a single inference process. We study this on BAGEL, a 14B mixture-of-transformers model that pairs autoregressive text generation with flow-matching image synthesis. We examine the impact of joint text-image generation on T2I performance and the relative importance of each modality during post-training. We additionally explore different post-training data strategies, showing that a targeted dataset addressing specific limitations achieves superior results compared to broad image-caption corpora or benchmark-aligned data. Using offline, reward-weighted post-training with fully self-generated synthetic data, our approach enables improvements in multimodal image generation across four diverse, independent T2I benchmarks, demonstrating the effectiveness of reward-weighting both modalities and strategically designed post-training data.

A programmed cell death learning signature predicts immunotherapy response and identifies AP1S1 as a regulator of immune exclusion in breast cancer

Chin J Cancer Res. 2026 Aug 30;38(4):480-500. doi: 10.21147/j.issn.1000-9604.2026.04.08.

ABSTRACT

OBJECTIVE: Breast cancer remains a leading cause of global cancer mortality, characterized by profound heterogeneity. While immune checkpoint blockade (ICB) has transformed oncology, its efficacy in breast cancer is often hindered by "immune-cold" microenvironments and immune exclusion. Programmed cell death (PCD) is a critical regulator of tumor immune microenvironment (TIME). However, its role in the breast cancer immune microenvironment remains poorly understood.

METHODS: We integrated multi-omics data from six breast cancer cohorts (N=3,764) to develop a programmed cell death learning signature (PCDsig) using over 100 machine learning combinations. The model was benchmarked against 29 published signatures. Single-cell transcriptomic analysis decoded the immune landscape and cellular crosstalk. The role of adaptor-related protein complex 1 subunit sigma 1 (AP1S1) was validated through a clinical cohort, in vitro functional assays, and in vivo syngeneic mouse models.

RESULTS: PCDsig significantly stratified patient prognosis across all cohorts, consistently outperforming 29 existing models. High PCDsig scores correlated with immune-excluded phenotypes, reduced CD8+ T cell infiltration, and lower immunophenoscores. Single-cell analysis revealed that high-PCDsig tumors utilize vascular endothelial growth factor A (VEGFA) signaling to foster an immunosuppressive microenvironment. AP1S1 was identified as the core driver of immune exclusion. And our clinical cohort supported the immune exclusion effect of AP1S1. AP1S1 knockdown impaired tumor progression in vitro and fundamentally remodeled the tumor immune ecosystem in vivo. Combining AP1S1 inhibition with anti-programmed cell death ligand 1 (anti-PD-L1) therapy exerted profound synergistic effects, driven by massive infiltration and functional activation of cytotoxic Granzyme B (GZMB)+CD8+ T cells.

CONCLUSIONS: Our study establishes the PCDsig we developed is a potential prognostic and predictive biomarker for breast cancer. We provide the first evidence of AP1S1 as a core immunomodulatory oncogene that mediates immune exclusion. Targeting AP1S1 represents a highly promising strategy to sensitize cold breast tumors to ICB, offering a new perspective for precision immunotherapy.

PMID:42712842 | PMC:PMC13551362 | DOI:10.21147/j.issn.1000-9604.2026.04.08

A programmed cell death learning signature predicts immunotherapy response and identifies AP1S1 as a regulator of immune exclusion in breast cancer

Chin J Cancer Res. 2026 Aug 30;38(4):480-500. doi: 10.21147/j.issn.1000-9604.2026.04.08.

ABSTRACT

OBJECTIVE: Breast cancer remains a leading cause of global cancer mortality, characterized by profound heterogeneity. While immune checkpoint blockade (ICB) has transformed oncology, its efficacy in breast cancer is often hindered by "immune-cold" microenvironments and immune exclusion. Programmed cell death (PCD) is a critical regulator of tumor immune microenvironment (TIME). However, its role in the breast cancer immune microenvironment remains poorly understood.

METHODS: We integrated multi-omics data from six breast cancer cohorts (N=3,764) to develop a programmed cell death learning signature (PCDsig) using over 100 machine learning combinations. The model was benchmarked against 29 published signatures. Single-cell transcriptomic analysis decoded the immune landscape and cellular crosstalk. The role of adaptor-related protein complex 1 subunit sigma 1 (AP1S1) was validated through a clinical cohort, in vitro functional assays, and in vivo syngeneic mouse models.

RESULTS: PCDsig significantly stratified patient prognosis across all cohorts, consistently outperforming 29 existing models. High PCDsig scores correlated with immune-excluded phenotypes, reduced CD8+ T cell infiltration, and lower immunophenoscores. Single-cell analysis revealed that high-PCDsig tumors utilize vascular endothelial growth factor A (VEGFA) signaling to foster an immunosuppressive microenvironment. AP1S1 was identified as the core driver of immune exclusion. And our clinical cohort supported the immune exclusion effect of AP1S1. AP1S1 knockdown impaired tumor progression in vitro and fundamentally remodeled the tumor immune ecosystem in vivo. Combining AP1S1 inhibition with anti-programmed cell death ligand 1 (anti-PD-L1) therapy exerted profound synergistic effects, driven by massive infiltration and functional activation of cytotoxic Granzyme B (GZMB)+CD8+ T cells.

CONCLUSIONS: Our study establishes the PCDsig we developed is a potential prognostic and predictive biomarker for breast cancer. We provide the first evidence of AP1S1 as a core immunomodulatory oncogene that mediates immune exclusion. Targeting AP1S1 represents a highly promising strategy to sensitize cold breast tumors to ICB, offering a new perspective for precision immunotherapy.

PMID:42712842 | PMC:PMC13551362 | DOI:10.21147/j.issn.1000-9604.2026.04.08

Flow-OPD: On-Policy Distillation for Flow Matching Models

arXiv:2605.08063v5 Announce Type: replace-cross Abstract: Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignment: the reward sparsity induced by scalar-valued rewards, and the gradient interference arising from jointly optimizing heterogeneous objectives, which together give rise to a 'seesaw effect' of competing metrics and pervasive reward hacking. Inspired by the success of On-Policy Distillation (OPD) in the large language model community, we propose Flow-OPD, the first unified post-training framework that integrates on-policy distillation into Flow Matching models. Flow-OPD adopts a two-stage alignment strategy: it first cultivates domain-specialized teacher models via single-reward GRPO fine-tuning, allowing each expert to reach its performance ceiling in isolation; it then establishes a robust initial policy through a Flow-based Cold-Start scheme and seamlessly consolidates heterogeneous expertise into a single student via a three-step orchestration of on-policy sampling, task-routing labeling, and dense trajectory-level supervision. We further introduce Manifold Anchor Regularization (MAR), which leverages a task-agnostic teacher to provide full-data supervision that anchors generation to a high-quality manifold, effectively mitigating the aesthetic degradation commonly observed in purely RL-driven alignment. Built upon Stable Diffusion 3.5 Medium, Flow-OPD raises the GenEval score from 63 to 92 and the OCR accuracy from 59 to 94, yielding an overall improvement of roughly 10 points over vanilla GRPO, while preserving image fidelity and human-preference alignment and exhibiting an emergent 'teacher-surpassing' effect. These results establish Flow-OPD as a scalable alignment paradigm for building generalist text-to-image models. The codes and weights will be released in: https://github.com/CostaliyA/Flow-OPD .

FDX1 as a predictive biomarker and therapeutic target for lymph node metastasis in gastric cancer

Clin Exp Med. 2026 May 10. doi: 10.1007/s10238-026-02160-0. Online ahead of print.

ABSTRACT

The prognostic values of cuproptosis-related genes (CRGs) in gastric cancer with lymph node metastasis (GCLM), especially in the tumor immune microenvironment (TIME), remain unclear. We analyzed the expression, mutation, immunity, drug sensitivity, and prognostic value of CRGs in GCLM using TCGA and GEO cohorts. Consensus clustering was performed to identify CRG subtypes, with differences characterized by multi-omics analysis. A CRG-based prognostic risk score and immune score were constructed for individualized assessment, and the role of CRGs was validated through in vitro and in vivo experiments. Consensus clustering revealed that CRGs were significantly enriched in biological processes related to mitosis and energy metabolism, as well as in immune-related and cancer-associated pathways. Four distinct CRG subtypes were identified, showing marked differences in expression profiles, prognosis, genetic alterations, TIME, and chemotherapeutic drug sensitivity. We developed an exploratory CRG-based prognostic risk score for preliminary individualized assessment, and the functional relevance of CRGs in GCLM was further validated through in vitro experiments. Among these, FDX1, LIAS, DLAT, MTF1, and GLS were identified as key determinants of overall survival in patients with GCLM, with FDX1 emerging as a potential independent prognostic factor. Notably, upregulation of FDX1 significantly suppressed lymph node metastasis of gastric cancer cells in a mouse popliteal lymph node metastasis model. Our data uncovers FDX1 might be a potential favorable prognostic factors in GCLM patients. These findings may improve our understanding of CRGs in GCLM and provide new in-sights for assessing prognosis and developing more effective treatment strategies.

PMID:42107026 | DOI:10.1007/s10238-026-02160-0

PIAS4 inhibition induces cell cycle arrest and exhibits a synergistic effect in combination with CDK4/6 inhibitor in breast cancer treatment

Oncogene, Published online: 07 April 2026; doi:10.1038/s41388-026-03753-5

PIAS4 inhibition induces cell cycle arrest and exhibits a synergistic effect in combination with CDK4/6 inhibitor in breast cancer treatment

Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies

arXiv:2604.00830v2 Announce Type: replace-cross Abstract: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is an adaptation policy that updates the actor policy based on experience from previous episodes, thereby improving future behavior. Existing methods rely on fixed, hand-crafted adaptation policies rather than optimizing them for downstream improvement. We argue that optimal adaptation policies should be learned from task environments, not hand-engineered based on human intuition. To achieve this, we introduce Meta-TTL, a framework that formulates the discovery of effective adaptation policies as a bi-level optimization problem. Within this framework, the inner loop executes the standard TTL process, measuring how effectively a candidate adaptation policy helps an agent correct errors across sequential episodes. Guided by the agent's performance, the outer loop employs evolutionary search over a diverse distribution of training tasks to iteratively refine the adaptation policy. We evaluate Meta-TTL on Jericho and WebArena-Lite across both in-distribution (ID) and out-of-distribution (OOD) settings, using multiple meta-agent backbones. Results on both benchmarks show that Meta-TTL consistently outperforms hand-crafted baselines, suggesting that the optimized adaptation policy encodes transferable strategies that generalize beyond the training task distribution.

Accuracy of Radiomics-Based Machine Learning for Predicting Risk of Recurrence in Non–Small Cell Lung Cancer: Systematic Review and Meta-Analysis

Background: During the diagnosis and treatment of non–small cell lung cancer (NSCLC), detecting the risk of its recurrence in an early phase is still challenging. Recent studies have investigated the radiomics-based machine learning (ML) models for detecting the risk of recurrence in NSCLC. However, there is still insufficient systematic evidence to prove its efficiency. Objective: This study is designed to systematically evaluate the effectiveness of radiomics-based ML in predicting the risk of recurrence in NSCLC, aiming to provide evidence-based support for the subsequent development of scoring tools to forecast recurrence risk. Methods: For acquiring research on radiomics-based models for forecasting the risk of recurrence in NSCLC, Cochrane Library, Web of Science, PubMed, and Embase were systematically retrieved, up to October 24, 2025. Studies on analyzing the recurrence of NSCLC using radiomics-based ML were included, while those in which only texture analysis was conducted or radiomics-based ML was not constructed were excluded. The Radiomics Quality Score (RQS) was used to appraise the eligible studies. Subgroup analyses were conducted according to the variables of the model, the background of treatment, the stage of lung cancer, and the pathological type. Results: Ultimately, 30 eligible studies in total were included, covering 7964 patients with NSCLC. According to the meta-analysis, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.850 (95% CI 0.834‐0.866, 95% prediction interval [PI] 0.623‐1.004) in the training set. Specifically, the pooled c-index was 0.876 (95% CI 0.853‐0.900) among the patients receiving the stereotactic body radiation therapy and 0.825 (95% CI 0.804‐0.848) among those who received surgeries combined with other adjuvant treatment regimens. The c-index of the radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.833 (95% CI 0.822‐0.854, 95% PI 0.717‐0.945) in the training set. In contrast, the c-index of radiomics-based ML models for forecasting the risk of recurrence in NSCLC was 0.878 (95% CI 0.854‐0.902, 95% PI 0.681‐1.000) in the validation set. The c-index of radiomics-based ML models combined with clinical features for forecasting the risk of recurrence in NSCLC was 0.854 (95% CI 0.830‐0.878, 95% PI 0.655‐0.992) in the validation set. The average RQS across the included studies was 27.4%, revealing methodological limitations and an absence of standardization. Conclusions: This study is the first to confirm that radiomics-based ML models effectively predict the risk of recurrence in NSCLC. This study provides evidence-based support for the subsequent development or updating of radiomics-based ML models. However, the current methodological application of radiomics remains concerning. Therefore, in the future, research should standardize the workflow for implementing radiomics-based ML and incorporate multicenter imaging data to enhance its generalizability. Trial Registration: PROSPERO CRD42025631191; https://www.crd.york.ac.uk/PROSPERO/view/CRD42025631191

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

arXiv:2601.22060v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting MLLMs by ``reasoning-then-tool-call'' for visual and textual search engines to obtain substantial gains on tasks requiring extensive factual information. However, these approaches typically define multimodal search in a naive setting, assuming that a single full-level or entity-level image query and few text query suffices to retrieve the key evidence needed to answer the question, which is unrealistic in real-world scenarios with substantial visual noise. Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision-DeepResearch, which proposes one new multimodal deep-research paradigm, i.e., performs multi-turn, multi-entity and multi-scale visual and textual search to robustly hit real-world search engines under heavy noise. Our Vision-DeepResearch supports dozens of reasoning steps and hundreds of engine interactions, while internalizing deep-research capabilities into the MLLM via cold-start supervision and RL training, resulting in a strong end-to-end multimodal deep-research MLLM. It substantially outperforming existing multimodal deep-research MLLMs, and workflows built on strong closed-source foundation model such as GPT-5, Gemini-2.5-pro and Claude-4-Sonnet. The code will be released in https://github.com/Osilly/Vision-DeepResearch.

Fluid-Derived Organoids from Pleural Effusion and Ascites: Emerging Models for Drug Resistance and Personalized Oncology

J Cancer. 2026 Mar 4;17(3):614-625. doi: 10.7150/jca.127511. eCollection 2026.

ABSTRACT

Malignant pleural effusion (MPE) and malignant ascites (MA) are common complications in advanced-stage cancers, often signifying disease progression and resistance to treatment. Compared to tissue biopsies or surgical specimens, materials derived from effusions offer advantages such as minimal invasiveness, ease of accessibility, and the feasibility of repeated collection during therapeutic interventions. Organoids generated from tumor cells in effusions, termed fluid-derived organoids (FDOs), have demonstrated the ability to maintain genetic heterogeneity and accurately replicate patient-specific tumor phenotypes. These characteristics position FDOs as promising models for investigating drug resistance mechanisms and informing personalized oncology strategies. In the context of lung cancer, organoids derived from pleural effusions have been employed to study acquired resistance to epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors and immunotherapy. Similarly, in ovarian and gastrointestinal cancers, organoids derived from ascites have proven to be valuable platforms for examining chemotherapy resistance and conducting drug sensitivity testing. FDOs have shown significant potential for translational applications by effectively correlating ex vivo drug responses with clinical outcomes, thus facilitating real-time monitoring of resistance evolution. However, several challenges remain, such as achieving culture standardization, maintaining the integrity of tumor microenvironment components, and integrating with multi-omics approaches. This review provides a comprehensive overview of recent advancements in the use of pleural effusion- and ascites-derived organoids for drug resistance research, underscores their applications in personalized oncology, and explores future research directions.

PMID:41869438 | PMC:PMC13003542 | DOI:10.7150/jca.127511

Fluid-Derived Organoids from Pleural Effusion and Ascites: Emerging Models for Drug Resistance and Personalized Oncology

J Cancer. 2026 Mar 4;17(3):614-625. doi: 10.7150/jca.127511. eCollection 2026.

ABSTRACT

Malignant pleural effusion (MPE) and malignant ascites (MA) are common complications in advanced-stage cancers, often signifying disease progression and resistance to treatment. Compared to tissue biopsies or surgical specimens, materials derived from effusions offer advantages such as minimal invasiveness, ease of accessibility, and the feasibility of repeated collection during therapeutic interventions. Organoids generated from tumor cells in effusions, termed fluid-derived organoids (FDOs), have demonstrated the ability to maintain genetic heterogeneity and accurately replicate patient-specific tumor phenotypes. These characteristics position FDOs as promising models for investigating drug resistance mechanisms and informing personalized oncology strategies. In the context of lung cancer, organoids derived from pleural effusions have been employed to study acquired resistance to epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors and immunotherapy. Similarly, in ovarian and gastrointestinal cancers, organoids derived from ascites have proven to be valuable platforms for examining chemotherapy resistance and conducting drug sensitivity testing. FDOs have shown significant potential for translational applications by effectively correlating ex vivo drug responses with clinical outcomes, thus facilitating real-time monitoring of resistance evolution. However, several challenges remain, such as achieving culture standardization, maintaining the integrity of tumor microenvironment components, and integrating with multi-omics approaches. This review provides a comprehensive overview of recent advancements in the use of pleural effusion- and ascites-derived organoids for drug resistance research, underscores their applications in personalized oncology, and explores future research directions.

PMID:41869438 | PMC:PMC13003542 | DOI:10.7150/jca.127511

Aryl hydrocarbon receptor is critical for both AR-dependent and AR-indifferent enzalutamide resistance in castration-resistant prostate cancer

Oncogene, Published online: 23 March 2026; doi:10.1038/s41388-026-03723-x

Aryl hydrocarbon receptor is critical for both AR-dependent and AR-indifferent enzalutamide resistance in castration-resistant prostate cancer

Extracellular Vesicles in Osteosarcoma: Mechanisms, Diagnostics and Therapeutic Applications

Drug Des Devel Ther. 2026 Jan 6;20:565059. doi: 10.2147/DDDT.S565059. eCollection 2026.

ABSTRACT

Osteosarcoma is a primary bone malignancy of adolescents and young adults with marked heterogeneity and a high metastatic propensity. Five-year survival exceeds 70% in localized disease but falls to about 20% with pulmonary metastasis or chemoresistance, and overall outcomes have plateaued for decades. Extracellular vesicles (EVs) have emerged as critical mediators of osteosarcoma progression and metastasis. EVs remodel the tumor microenvironment (TME) by promoting immune evasion, extracellular matrix reprogramming, and angiogenesis, while also facilitating invasion, epithelial-mesenchymal transition (EMT)-like plasticity, and formation of lung pre-metastatic niches through organotropic integrins and glycoproteins. Their cargo, including proteins, lipids, and nucleic acids, drives intercellular communication that sustains proliferation, migration, and therapy resistance under metabolic or hypoxic stress. Clinically, the stability of EVs in body fluids and their tumor-specific molecular signatures highlight their promise as liquid-biopsy biomarkers for early diagnosis, prognosis, and treatment monitoring. Therapeutically, EVs are being engineered as delivery vehicles for drugs or RNA therapeutics, and interventions targeting their biogenesis, cargo sorting, or uptake are under exploration. Future research should integrate single-EV multi-omics, longitudinal cohort validation, and causal perturbation models to delineate functional mechanisms. Rational strategies that modulate EV dynamics and incorporate standardized analytic pipelines may transform EVs into actionable biomarkers and therapeutic targets, offering new avenues to overcome resistance and improve clinical outcomes in osteosarcoma.

PMID:41858917 | PMC:PMC12998350 | DOI:10.2147/DDDT.S565059

LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma

Oncogene, Published online: 14 March 2026; doi:10.1038/s41388-026-03714-y

LINC-AC092535.5 regulates MICAL2 mRNA level to inhibit p53-mediated ferroptosis in nasopharyngeal carcinoma

A Novel Multi-Agent Architecture to Reduce Hallucinations of Large Language Models in Multi-Step Structural Modeling

arXiv:2603.07728v1 Announce Type: new Abstract: Large language models (LLMs) such as GPT and Gemini have demonstrated remarkable capabilities in contextual understanding and reasoning. The strong performance of LLMs has sparked growing interest in leveraging them to automate tasks traditionally dependent on human expertise. Recently, LLMs have been integrated into intelligent agents capable of operating structural analysis software (e.g., OpenSees) to construct structural models and perform analyses. However, existing LLMs are limited in handling multi-step structural modeling due to frequent hallucinations and error accumulation during long-sequence operations. To this end, this study presents a novel multi-agent architecture to automate the structural modeling and analysis using OpenSeesPy. First, problem analysis and construction planning agents extract key parameters from user descriptions and formulate a stepwise modeling plan. Node and element agents then operate in parallel to assemble the frame geometry, followed by a load assignment agent. The resulting geometric and load information is translated into executable OpenSeesPy scripts by code translation agents. The proposed architecture is evaluated on a benchmark of 20 frame problems over ten repeated trials, achieving 100% accuracy in 18 cases and 90% in the remaining two. The architecture also significantly improves computational efficiency and demonstrates scalability to larger structural systems.

xLLM Technical Report

arXiv:2510.14686v2 Announce Type: replace-cross Abstract: We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locates online and offline tasks through unified elastic scheduling to maximize cluster utilization. This module also relies on a workload-adaptive dynamic Prefill-Decode (PD) disaggregation policy and a novel Encode-Prefill-Decode (EPD) disaggregation policy designed for multimodal inputs. Furthermore, it incorporates a distributed architecture to provide global KV Cache management and robust fault-tolerant capabilities for high availability. At the engine layer, xLLM-Engine co-optimizes system and algorithm designs to fully saturate computing resources. This is achieved through comprehensive multi-layer execution pipeline optimizations, an adaptive graph mode and an xTensor memory management. xLLM-Engine also further integrates algorithmic enhancements such as optimized speculative decoding and dynamic EPLB, collectively serving to substantially boost throughput and inference efficiency. Extensive evaluations demonstrate that xLLM delivers significantly superior performance and resource efficiency. Under identical TPOT constraints, xLLM achieves throughput up to 1.7x that of MindIE and 2.2x that of vLLM-Ascend with Qwen-series models, while maintaining an average throughput of 1.7x that of MindIE with Deepseek-series models. xLLM framework is publicly available at https://github.com/jd-opensource/xllm and https://github.com/jd-opensource/xllm-service.

WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality

arXiv:2510.18560v3 Announce Type: replace-cross Abstract: The paradigm of LLM-as-a-judge is emerging as a scalable and efficient alternative to human evaluation, demonstrating strong performance on well-defined tasks. However, its reliability in open-ended tasks with dynamic environments and complex interactions remains unexplored. To bridge the gap, we introduce WebDevJudge, a systematic benchmark for assessing LLM-as-a-judge performance in web development, with support for both non-interactive evaluation based on static observations and continuous interactive evaluation with a dynamic web environment. WebDevJudge comprises human preference labels over paired web implementations, annotated with structured and query-grounded rubrics to ensure high-quality ground truth. Using this benchmark, we comprehensively evaluate various evaluators, including LLMs, MLLMs, and agentic workflows. We systematically investigate the impact of different paradigms and guidance mechanisms. Our experiments reveal a significant gap between LLM judges and human experts. In-depth analysis indicates this gap stems from fundamental model limitations, including failures in recognizing functional equivalence, verifying task feasibility, and mitigating bias. Overall, WebDevJudge presents a challenge to LLM-as-a-judge, offering insights to guide future research toward developing more reliable and capable automated evaluators for complicated scenarios. Code and data are available at https://github.com/lcy2723/WebDevJudge.
❌