❌

Normal view

ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC

arXiv:2609.13083v1 Announce Type: cross Abstract: In urban search and rescue, articulated tracked robots (ATRs) must traverse structured but contact-rich environments such as stairwells and cluttered building interiors. Reliable autonomy remains challenging because robot-terrain interaction (RTI) is hybrid and discontinuous, and effective flipper-track coordination is difficult to model analytically. We present ASTRIL-MPC, a language-guided neural kinematics model predictive control (MPC) framework for autonomous traversal. A learned kinematics model predicts short-horizon task-state increments from a height sequence and recent trajectories; NMPC plans with multi-objective costs and strict feasibility constraints; and a large language model (LLM) proposes bounded updates to selected weights and bounds through a safety-checked interface with range clipping, rate limiting, and consistency checks. The compiled predictor enables a full control cycle within 100 ms. Across three traversal tasks and a multi-height generalization setting, ASTRIL-MPC improves an aggregate traversal-quality score by up to 71% over a non-adaptive NMPC and by 67% over a PPO baseline, while eliminating measurable collision impacts during descent. These results indicate that combining learned kinematics, optimization-based planning, and language-guided retuning yields data-efficient and robust autonomy for articulated tracked robots.

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

arXiv:2609.04298v2 Announce Type: replace Abstract: Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

Terminalia chebula Retz. aqueous extract exerts anti-adhesive and anti-inflammatory effects against Helicobacter pylori: Insights from lysine metabolism remodeling and fecal metabolomics

J Ethnopharmacol. 2026 Aug 29;373:122318. doi: 10.1016/j.jep.2026.122318. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: Helicobacter pylori (H. pylori) infection is a leading risk factor for chronic gastritis, peptic ulcers, and gastric cancer. The escalating antibiotic resistance of H. pylori and adverse effects arising from standard antibiotic regimens have underscored an urgent need for natural, food-complementary therapeutic alternatives. Terminalia chebula Retz., commonly named "Hezi" in traditional Chinese medicine and "Haritaki" in Ayurveda, is a well-recognized edible fruit with a long history of use in alleviating gastrointestinal disorders.

AIM OF THE STUDY: While it has been traditionally recognized for its anti-inflammatory and antimicrobial properties, its anti-adhesive efficacy against H. pylori and the metabolic mechanisms underlying its in vivo effects remain largely unexplored.

MATERIALS AND METHODS: This study adopted multiple analytical and experimental approaches: ultra-high-performance liquid chromatography-tandem mass spectrometry (UPLC-MS/MS), RNA-seq, cell viability and adhesion assays, western blotting, hematoxylin and eosin (H&E) staining, enzyme-linked immunosorbent assay (ELISA), metabolomics, and proteomics.

RESULTS: In this study, 15 primary compounds in T. chebula aqueous extract were identified, predominantly tannins. Multi-omics analyses (transcriptomics, metabolomics, and proteomics) revealed that T. chebula aqueous extract significantly inhibited H. pylori adhesion and enriched the lysine degradation pathway. Notably, N-alpha-acetyl-L-lysine was characterized as a key metabolite in fecal metabolomic profiling, which was associated with glycosphingolipid biosynthesis, ferroptosis, HIF-1 signaling, lysosomal function, and arginine metabolism. In vitro assays confirmed that N-alpha-acetyl-L-lysine reduced H. pylori adhesion to GES-1 cells, reversed H. pylori-induced cellular damage, and suppressed the secretion of pro-inflammatory cytokines (IL-6 and TNF-α). In vivo, T. chebula aqueous extract administration markedly alleviated H. pylori-induced gastric inflammation by downregulating IL-6, IL-1β, TNF-α, TGF-β, and IFN-γ. Collectively, these findings demonstrate that T. chebula aqueous extract exerts anti-H. pylori effects by regulating lysine metabolism, with N-alpha-acetyl-L-lysine serving as a key metabolite via fecal metabolomics.

CONCLUSION: These findings highlight T. chebula's promising potential as a functional food ingredient or adjunctive therapy for H. pylori-related diseases, providing a natural, mechanism-based option for clinical intervention.

PMID:42665167 | DOI:10.1016/j.jep.2026.122318

DBPnet: Damper Characteristics-Based Bayesian Physics-Informed Neural Network for Wheel Load Estimation

arXiv:2605.24860v1 Announce Type: cross Abstract: Advanced driver assistance systems (ADAS) play an important role in modern automotive intelligence, significantly enhancing vehicle safety and stability. The performance of ADAS critically relies on accurate and reliable vehicle state estimation, particularly from vehicle dynamic sensors. Among these signals, wheel load is a key variable for chassis control and safety-critical functions, yet it remains difficult to estimate robustly due to complex suspension geometry, nonlinear dynamics, and measurement noise. To address this issue, we propose DBPnet, a Bayesian physics-informed neural network (PINN) with a physics-aware embedding module inspired by damper characteristics. First, this paper presents a suspension linkage-level modeling (SLLM) approach that constructs a nonlinear instantaneous dynamic model by explicitly considering the complex geometric structure of the suspension. Building upon SLLM, Bayesian inference is integrated into the PINN to effectively cope with noise and uncertainty in the vehicle chassis system, thereby improving the model's robustness. Then, a physics-informed loss function is employed to ensure consistency with fundamental physical principles, while the damper characteristics-inspired embedding module extracts temporal variation features of input signals and incorporates them into each layer of the PINN, ensuring that physical observations guide the neural network without being constrained by fixed physical models. Extensive evaluations on high-fidelity simulations and real-world experiments demonstrate that our DBPnet consistently achieves lower RMSE and MaxError than baseline methods. These results highlight the potential of our DBPnet to advance wheel load estimation and contribute to the development of more reliable ADAS actuator functions.

EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis

arXiv:2604.08213v2 Announce Type: replace-cross Abstract: High-quality source-target image pairs with precise editing instructions are essential for instruction-guided image editing, yet constructing such training triplets at scale remains costly. Recent pipelines often rely on vision-language models to synthesize editing instructions automatically, but we find that strong VLMs still struggle to describe visual transformations between image pairs. In particular, they exhibit three recurring failure modes: orientation inconsistency, viewpoint ambiguity, and missing fine-grained attributes. In a human evaluation on 400 image pairs, several open-source VLM baselines produce critical-error rates above 47\%, making many synthesized instructions unsuitable for downstream training. To address this, we propose EditCaption, a two-stage post-training pipeline for image editing instruction synthesis. First, we construct a 100K supervised fine-tuning dataset through GLM-based auto-captioning, EditScore filtering, and human refinement. Second, we collect 10K human-annotated preference pairs, where each rejected instruction is labeled with its primary error type and severity. Based on this dataset, we propose Hardness-Adaptive Error-Aware DPO (HAE-DPO), a task-adapted DPO objective that introduces an adaptive margin based on human-labeled severity, failure-mode type, and reference-model hardness. Experiments across three benchmarks demonstrate that our 235B model with SFT+HAE-DPO achieves state-of-the-art performance among open-source and closed models, scoring 4.720 on Eval-400, 4.672 on HQ-Edit, and 4.651 on ByteMorph-Bench -- surpassing Gemini-3-Pro on all three. Human evaluation confirms critical error rates drop from 47.75\% to 17.50\%, with correct rates improving from 41.75\% to 70.25\%, surpassing Gemini-3-Pro (66.00\%).

A framework for building a synthetic cell from the SynCell Asia Initiative

Nature Biotechnology, Published online: 26 May 2026; doi:10.1038/s41587-026-03153-w

Building a living cell from scratch requires overcoming a bottleneck that has remained unresolved despite decades of progress: orchestrating the spatiotemporal integration of core functional modules. To tackle this barrier, the SynCell Asia Initiative outlines a strategy for developing core functional modules followed by their systems-level integration through the establishment of a centralized, artificial intelligence (AI)-driven biofoundry.

Multi-omics biomarkers for predicting resistance, hyperprogression, and immune-related toxicity during PD-1/PD-L1 therapy in lung cancer: a literature review

Front Immunol. 2026 May 8;17:1780459. doi: 10.3389/fimmu.2026.1780459. eCollection 2026.

ABSTRACT

Immune checkpoint inhibitors targeting programmed cell death protein 1 (PD-1) and its ligand programmed death-ligand 1 (PD-L1) have transformed the management of advanced lung cancer, yet most patients experience primary resistance, hyperprogressive disease (HPD), or clinically significant immune-related adverse events (irAEs). Multi-omics technologies now enable integrated interrogation of tumor, microenvironmental, host, and clinical determinants of these divergent outcomes. In this review, we first discuss the biological and clinical foundations of PD-1/PD-L1 blockade in non-small cell and small cell lung cancer, and summarize the spectrum of resistance, HPD, and irAEs observed in trials and real-world practice. We then describe multi-omics study frameworks that connect genomics, transcriptomics, epigenomics, proteomics, metabolomics, radiomics, and microbiome profiling with these outcome phenotypes. Building on this foundation, we synthesize evidence for composite biomarkers of primary and acquired resistance, delineate emerging multi-omics signatures of HPD, and examine host- and tumor-derived multi-omics correlates of organ-specific and systemic irAEs. We further propose an efficacy-risk quadrant framework to guide clinical decision-making when favorable efficacy predictors coexist with elevated risk of severe adverse outcomes, and outline a three-step approach for high-efficacy/high-risk patients: joint probability reporting, multi-omics guided mitigation, and dynamic reassessment. Finally, we evaluate translational strategies that integrate multi-omics scores into baseline risk stratification, dynamic monitoring with attention to technical challenges such as distinguishing true progression from ctDNA pseudoprogression, and biomarker-driven trial design, while assessing the evidence level and translational readiness of candidate assays from retrospective discovery to clinical implementation. A clinical case illustrates how multi-omics can link baseline risk stratification, regimen selection, and longitudinal monitoring into a coherent action plan, while acknowledging that artificial intelligence-driven models remain investigational and real-world application still relies on clinician judgment. Collectively, this review defines how integrated multi-omics biomarkers can be leveraged to predict resistance, HPD, and immune-related toxicity, and to refine patient selection and management during PD-1/PD-L1 therapy in lung cancer.

PMID:42183274 | PMC:PMC13194140 | DOI:10.3389/fimmu.2026.1780459

Multi-omics biomarkers for predicting resistance, hyperprogression, and immune-related toxicity during PD-1/PD-L1 therapy in lung cancer: a literature review

Front Immunol. 2026 May 8;17:1780459. doi: 10.3389/fimmu.2026.1780459. eCollection 2026.

ABSTRACT

Immune checkpoint inhibitors targeting programmed cell death protein 1 (PD-1) and its ligand programmed death-ligand 1 (PD-L1) have transformed the management of advanced lung cancer, yet most patients experience primary resistance, hyperprogressive disease (HPD), or clinically significant immune-related adverse events (irAEs). Multi-omics technologies now enable integrated interrogation of tumor, microenvironmental, host, and clinical determinants of these divergent outcomes. In this review, we first discuss the biological and clinical foundations of PD-1/PD-L1 blockade in non-small cell and small cell lung cancer, and summarize the spectrum of resistance, HPD, and irAEs observed in trials and real-world practice. We then describe multi-omics study frameworks that connect genomics, transcriptomics, epigenomics, proteomics, metabolomics, radiomics, and microbiome profiling with these outcome phenotypes. Building on this foundation, we synthesize evidence for composite biomarkers of primary and acquired resistance, delineate emerging multi-omics signatures of HPD, and examine host- and tumor-derived multi-omics correlates of organ-specific and systemic irAEs. We further propose an efficacy-risk quadrant framework to guide clinical decision-making when favorable efficacy predictors coexist with elevated risk of severe adverse outcomes, and outline a three-step approach for high-efficacy/high-risk patients: joint probability reporting, multi-omics guided mitigation, and dynamic reassessment. Finally, we evaluate translational strategies that integrate multi-omics scores into baseline risk stratification, dynamic monitoring with attention to technical challenges such as distinguishing true progression from ctDNA pseudoprogression, and biomarker-driven trial design, while assessing the evidence level and translational readiness of candidate assays from retrospective discovery to clinical implementation. A clinical case illustrates how multi-omics can link baseline risk stratification, regimen selection, and longitudinal monitoring into a coherent action plan, while acknowledging that artificial intelligence-driven models remain investigational and real-world application still relies on clinician judgment. Collectively, this review defines how integrated multi-omics biomarkers can be leveraged to predict resistance, HPD, and immune-related toxicity, and to refine patient selection and management during PD-1/PD-L1 therapy in lung cancer.

PMID:42183274 | PMC:PMC13194140 | DOI:10.3389/fimmu.2026.1780459

circPARPBP promotes cancer stemness and chemoresistance in triple-negative breast cancer through recruiting SRCAP complex to activate CCL20 transcription

Oncogene, Published online: 21 May 2026; doi:10.1038/s41388-026-03819-4

circPARPBP promotes cancer stemness and chemoresistance in triple-negative breast cancer through recruiting SRCAP complex to activate CCL20 transcription

The 2025 lung cancer landscape: advances in screening, molecular taxonomy and therapeutic strategy: a narrative review

Transl Lung Cancer Res. 2026 Mar 23;15(3):62. doi: 10.21037/tlcr-2025-1-1477. Epub 2026 Mar 18.

ABSTRACT

BACKGROUND AND OBJECTIVE: In 2025, lung cancer research advanced rapidly across the disease continuum, from population-level risk assessment and screening to mechanistic studies of early carcinogenesis and therapeutic innovation in perioperative and metastatic settings. A key shift moved beyond a smoking-centred paradigm toward a multidimensional risk framework reflecting the growing burden among never-smokers and the roles of air pollution, occupational exposures, and systemic metabolic-inflammatory states. This narrative review aims to synthesize influential 2025 evidence across prevention, diagnosis, treatment, and survivorship, and to identify convergent themes and translational gaps relevant to clinical practice and policy.

METHODS: We performed a narrative synthesis of influential lung cancer studies published in major international journals in 2025. Evidence was organized along a clinically oriented pathway spanning carcinogenesis and screening, precision diagnosis, treatment optimization in resectable and advanced disease, and survivorship, emphasizing practice-informing trials, high-impact translational research, and implementation-relevant technologies.

KEY CONTENT AND FINDINGS: Lineage tracing, single-cell and spatial omics, and evolutionary inference refined concepts of field cancerization, clonal selection, and copy-number-driven fitness. In small-cell lung cancer, evidence further supported neuronal coupling and synapse-like programs as potentially tractable vulnerabilities. Clinically, low-dose computed tomography (CT) strategies and data-informed nodule thresholds aimed to balance under-detection against over-surveillance harms. In diagnostics, artificial intelligence (AI) models increasingly inferred molecular features from routine histopathology ("virtual molecular testing") and should be regarded as decision support requiring prospective validation, population calibration, and explicit failure-mode reporting. Multimodal approaches integrating imaging with circulating tumor DNA (ctDNA) improved feasibility in tissue-limited settings, but clinical utility remains contingent on assay standardization and pathway-level implementation. In resectable disease, longer follow-up consolidated neoadjuvant chemo-immunotherapy for selected patients, while ctDNA kinetics emerged as a candidate biomarker for response-adaptive escalation and de-escalation. In advanced non-small cell lung cancer (NSCLC), phase III evidence for antibody-drug conjugates and bispecific antibodies began reshaping sequencing, while highlighting challenges in toxicity, access, affordability, and immature overall survival in several programs.

CONCLUSIONS: The 2025 landscape reflects coordinated progress in risk conceptualization, biology, diagnostics, and therapeutics, yet gaps in validation, standardization, and real-world deliverability persist. Priorities include prospective evaluation of AI- and ctDNA-enabled pathways, toxicity-informed sequencing, and equitable implementation aligned with health-system capacity.

PMID:41982682 | PMC:PMC13071762 | DOI:10.21037/tlcr-2025-1-1477

Scaling Attention via Feature Sparsity

arXiv:2603.22300v1 Announce Type: cross Abstract: Scaling Transformers to ultra-long contexts is bottlenecked by the $O(n^2 d)$ cost of self-attention. Existing methods reduce this cost along the sequence axis through local windows, kernel approximations, or token-level sparsity, but these approaches consistently degrade accuracy. In this paper, we instead explore an orthogonal axis: feature sparsity. We propose Sparse Feature Attention (SFA), where queries and keys are represented as $k$-sparse codes that preserve high-dimensional expressivity while reducing the cost of attention from $\Theta(n^2 d)$ to $\Theta(n^2 k^2/d)$. To make this efficient at scale, we introduce FlashSFA, an IO-aware kernel that extends FlashAttention to operate directly on sparse overlaps without materializing dense score matrices. Across GPT-2 and Qwen3 pretraining, SFA matches dense baselines while improving speed by up to $2.5\times$ and reducing FLOPs and KV-cache by nearly 50\%. On synthetic and downstream benchmarks, SFA preserves retrieval accuracy and robustness at long contexts, outperforming short-embedding baselines that collapse feature diversity. These results establish feature-level sparsity as a complementary and underexplored axis for efficient attention, enabling Transformers to scale to orders-of-magnitude longer contexts with minimal quality loss. Code is available at https://github.com/YannX1e/Sparse-Feature-Attention.

PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding

arXiv:2603.22796v1 Announce Type: cross Abstract: Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multimodal Models (LMMs) reasoning with a novel control paradigm. PhotoAgent first translates subjective aesthetic goals into solvable geometric constraints via LMM-driven, chain-of-thought (CoT) reasoning, allowing an analytical solver to compute a high-quality initial viewpoint. This initial pose is then iteratively refined through visual reflection within a photorealistic internal world model built with 3D Gaussian Splatting (3DGS). This ``mental simulation'' replaces costly and slow physical trial-and-error, enabling rapid convergence to aesthetically superior results. Evaluations confirm that PhotoAgent excels in spatial reasoning and achieves superior final image quality.

Low-dose intestinal irradiation enhances the efficacy and prognosis of PD-1 blockade in metastatic non-small cell lung cancer

Clin Cancer Res. 2026 Mar 18. doi: 10.1158/1078-0432.CCR-25-4153. Online ahead of print.

ABSTRACT

PURPOSE: Intestinal low-dose irradiation (ILDR) may enhance immunotherapy efficacy by modulating the gut microbiota and metabolism; however, its role in metastatic non-small cell lung cancer (mNSCLC), particularly in the first-line setting, remains unclear.

EXPERIMENTAL DESIGN: This multicenter retrospective and prospective study included mNSCLC patients receiving first- and second-line programmed cell death protein 1 (PD-1) inhibitors along with abdominopelvic radiotherapy between 2018 and 2025. Patients were stratified by the mean intestinal radiation dose into <1 Gy, 1-3 Gy, and >3 Gy groups and treatment outcomes were compared. The blood and fecal samples were subjected to multi-omics profiling.

RESULTS: g>309 patients were included in the retrospective analysis. Optimal efficacy was observed with a small intestinal mean radiation dose (SIMRD) of 1-3 Gy, showing longer progression-free survival (PFS, 10.2 months) and overall survival (OS, 22.8 months) (P < 0.01), which was consistent across subgroups. Compared with 1-3 Gy, SIMRD >3 Gy (Hazard ratio [HR] = 4.87, P < 0.001) and <1 Gy (HR = 1.85, P < 0.001) independently predicted worse OS. Prospective results confirmed the best disease control rate (P = 0.041) and PFS (P = 0.046) with SIMRD of 1-3 Gy. Responders were enriched in Bacillota, Clostridia, and indole derivatives, particularly indole-3-carboxylic acid. Moreover, the 1-3 Gy group exhibited increased circulating macrophage inflammatory protein-3α and reduced circulating α4β7+ regulatory T cells.

CONCLUSIONS: ILDR influences the efficacy of PD-1 blockade in patients with mNSCLC, particularly when SIMRD is maintained within the 1-3 Gy range, likely through modulation of the gut microbiota-metabolite-immune axis.

PMID:41849236 | DOI:10.1158/1078-0432.CCR-25-4153

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

arXiv:2602.12670v3 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurations over 7,308 trajectories. Curated Skills raise average pass rate by 16.2 percentage points(pp), but effects vary widely by domain (+4.5pp for Software Engineering to +51.9pp for Healthcare) and 16 of 84 tasks show negative deltas. Self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming. Focused Skills with 2--3 modules outperform comprehensive documentation, and smaller models with Skills can match larger models without them.

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

arXiv:2602.12670v2 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurations over 7,308 trajectories. Curated Skills raise average pass rate by 16.2 percentage points(pp), but effects vary widely by domain (+4.5pp for Software Engineering to +51.9pp for Healthcare) and 16 of 84 tasks show negative deltas. Self-generated Skills provide no benefit on average, showing that models cannot reliably author the procedural knowledge they benefit from consuming. Focused Skills with 2--3 modules outperform comprehensive documentation, and smaller models with Skills can match larger models without them.

Lysophosphatidylcholine acyltransferase 1 promotes head and neck squamous cell carcinoma progression by enhancing COX17-dependent oxidative phosphorylation

Cell Death Discovery, Published online: 06 March 2026; doi:10.1038/s41420-026-02994-3

Lysophosphatidylcholine acyltransferase 1 promotes head and neck squamous cell carcinoma progression by enhancing COX17-dependent oxidative phosphorylation

Kaiwu-PyTorch-Plugin: Bridging Deep Learning and Photonic Quantum Computing for Energy-Based Models and Active Sample Selection

arXiv:2602.19114v1 Announce Type: cross Abstract: This paper introduces the Kaiwu-PyTorch-Plugin (KPP) to bridge Deep Learning and Photonic Quantum Computing across multiple dimensions. KPP integrates the Coherent Ising Machine into the PyTorch ecosystem, addressing classical inefficiencies in Energy-Based Models. The framework facilitates quantum integration in three key aspects: accelerating Boltzmann sampling, optimizing training data via Active Sampling, and constructing hybrid architectures like QBM-VAE and Q-Diffusion. Empirical results on single-cell and OpenWebText datasets demonstrate KPPs ability to achieve SOTA performance, validating a comprehensive quantum-classical paradigm.

KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider

arXiv:2506.02634v5 Announce Type: replace-cross Abstract: Serving large language models (LLMs) is important for cloud providers, and caching intermediate results (KV\$) after processing each request substantially improves serving throughput and latency. However, there is limited understanding of how LLM serving benefits from KV\$ caching, where system design decisions like cache eviction policies are highly workload-dependent. In this paper, we present the first systematic characterization of the KV\$ workload patterns from one of the leading LLM service providers. We draw observations that were not covered by previous studies focusing on synthetic workloads, including: KV\$ reuses are skewed across requests, where reuses between single-turn requests are equally important as multi-turn requests; the reuse time and probability are diverse considering all requests, but for a specific request category, the pattern tends to be predictable; and the overall cache size required for an ideal cache hit ratio is moderate. Based on the characterization, we further propose a workload-aware cache eviction policy that improves the serving performance under real-world traces, especially with limited cache capacity.
❌