❌

Normal view

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605.25246v2 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines. Existing benchmarks are limited to small or simplified examples far below real-world scale and complexity. We introduce FrontierOR, among the first benchmarks to systematically evaluate LLM-based efficient algorithm design for realistic large-scale optimization problems. FrontierOR includes 180 tasks derived from methodologically diverse papers published in top-tier operations research venues, each with standardized instances and a hidden, expert-verified evaluation suite. We evaluate seven LLMs spanning frontier, cost-effective, and open-source models both in one-shot and test-time evolution settings. The results reveal that frontier models still struggle to move from executable formulations to efficient optimization algorithms: the strongest one-shot model outperforms Gurobi in only 31% of cases in both solution quality and computational efficiency, and even strong coding agents with test-time evolution achieve only 50% on selected hard tasks. FrontierOR establishes a practical evaluation platform for LLM-based optimization algorithm design, which enables future LLMs and agents to be systematically tested on whether they can move beyond correct formulation toward a feasible, high-quality, and efficient algorithm.

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

arXiv:2605.24041v2 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure that struggles to resolve high-frequency details, a limitation known as spectral bias. We introduce the Iterative Refinement Neural Operator (IRNO), which augments pre-trained operators with a learned refinement module iteratively applied via fixed-point iteration. IRNO decomposes the prediction into a coarse initialization followed by successive residual corrections, paralleling classical numerical solvers. Under local assumptions, we establish contraction of the induced operator, ensuring convergence to a unique fixed point. To explicitly target high-frequency errors, we propose a progressive spectral loss that adaptively increases penalty on high-frequency components over refinement steps during training. Across physical systems, IRNO consistently lowers error, with up to 56.05% improvement on turbulent flow. On Active Matter, spectral analysis reveals that, relative to base operator, the normalized error ratios decrease to 27.72-36.10% in low-, 5.07-6.68% in mid-, and 1.48-2.04% in high-frequencies, remaining stable beyond the trained iteration count. Code is available at https://github.com/xiaotianliu-dartmouth/Iterative_Refinement_Neural_Operator

NSR-Boost: A Neuro-Symbolic Residual Boosting Framework for Industrial Legacy Models

arXiv:2601.10457v3 Announce Type: replace Abstract: Although the Gradient Boosted Decision Trees (GBDTs) dominate industrial tabular applications, upgrading legacy models in high-concurrency production environments still faces prohibitive retraining costs and systemic risks. To address this problem, we present NSR-Boost, a neuro-symbolic residual boosting framework designed specifically for industrial scenarios. Its core advantage lies in being ``non-intrusive''. It treats the legacy model as a frozen model and performs targeted repairs on "hard regions" where predictions fail. The framework comprises three key stages: First, finding hard regions through residuals, then generating interpretable experts by generating symbolic code structures using Large Language Model (LLM) and fine-tuning parameters using Bayesian optimization, and finally dynamically integrating experts with legacy model output through a lightweight aggregator. Experimental results demonstrate that the framework significantly outperforms state-of-the-art (SOTA) baselines across six public datasets and one private dataset. More importantly, we report the successful deployment of NSR-Boost within the core financial risk control system of Qfin Holdings, where empirical results on real-world online traffic exhibit superior performance improvements and a significant reduction in the bad rate. In conclusion, it effectively captures long-tail risks missed by traditional models and offers a safe, low-cost evolutionary paradigm for industry.

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

arXiv:2605.23904v2 Announce Type: replace Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization. Code: https://aka.ms/skillopt

Forest carbon protocols underestimate climate-driven carbon loss risks

Nature, Published online: 20 May 2026; doi:10.1038/s41586-026-10571-y

The buffer pool designed to compensate for unintended carbon losses from the largest forest climate mitigation programme in the United States is too small when considering the impact of future climate change scenarios.

Stromal ACTA2 Counteracts TCDD-Induced Hepatocarcinogenesis via Suppression of the PI3K-AKT-mTOR Pathway

J Hepatocell Carcinoma. 2026 May 10;13:586916. doi: 10.2147/JHC.S586916. eCollection 2026.

ABSTRACT

PURPOSE: 2,3,7,8-Tetrachlorodibenzo-p-dioxin (TCDD) is a persistent environmental pollutant that promotes hepatocellular carcinoma (HCC) through non-genotoxic mechanisms. However, stromal regulatory factors that counteract its tumor-promoting effects remain poorly defined. This study aimed to elucidate the role of actin alpha-2 (ACTA2) in TCDD-associated hepatocarcinogenesis.

METHODS: An integrative strategy combining network toxicology, Mendelian randomization, multi-omics and single-cell analyses, molecular docking and molecular dynamics simulations, along with in vitro experiments, was employed to investigate the functional role of ACTA2.

RESULTS: ACTA2 was identified as a stromal-associated factor linked to reduced HCC risk and improved patient survival. Single-cell and multi-omics analyses revealed that ACTA2 is predominantly expressed in hepatic stellate cells and fibroblast-like populations, reflecting tumor microenvironment composition rather than tumor cell-intrinsic expression. Functional enrichment analyses indicated that ACTA2 is associated with extracellular matrix remodeling and PI3K-AKT signaling. Molecular simulations demonstrated stable binding of TCDD to ACTA2 (ΔG_bind ≈ -7.05 kcal/mol), suggesting potential structural perturbation. In vitro experiments showed that TCDD downregulated ACTA2 expression, promoted proliferation of LX-2 and cancer-associated fibroblasts (CAFs), and activated PI3K-AKT-mTOR signaling, whereas ACTA2 overexpression attenuated these effects.

CONCLUSION: ACTA2 acts as a context-dependent stromal regulator that modulates PI3K-AKT-mTOR signaling in TCDD-induced hepatocarcinogenesis. These findings highlight the importance of stromal remodeling in environmental carcinogenesis and suggest ACTA2 as a potential biomarker and therapeutic target in dioxin-associated HCC.

PMID:42148320 | PMC:PMC13175077 | DOI:10.2147/JHC.S586916

Refined immune-based molecular subtypes of gastric cancer: Integrating mismatch repair status and tumor microenvironment for enhanced immunotherapy prediction

Chin J Cancer Res. 2026 Apr 30;38(2):234-251. doi: 10.21147/j.issn.1000-9604.2026.02.09.

ABSTRACT

OBJECTIVE: Gastric cancer (GC) is heterogeneous, and current mismatch repair (MMR)-based classifications incompletely predict response to immune checkpoint inhibitors (ICIs).

METHODS: RNA sequencing (RNA-seq) and immune infiltration profiles from 189 resected GC were used to derive four refined immune-MMR subtypes (R1-R4) by integrating MMR status, survival, and tumor microenvironment (TME) features. Multi-omics profiling and pathway analysis defined subtype biology. External transcriptomic cohorts and an ICI-treated cohort were classified with Nearest Template Prediction (NTP). Immune response-associated genes were identified from responder vs. non-responder comparisons within the ICI-sensitive subtype and validated by multiplex immunohistochemistry (mIHC).

RESULTS: R1 showed the best prognosis and highest immunotherapy response with objective response rate (ORR) 54.5%, while R4 had the worst prognosis. R2 represented an immune-unresponsive deficient mismatch repair (dMMR) subset, and R3 captured an immune-active proficient mismatch repair (pMMR) subgroup with moderate therapy sensitivity. Multi-omics integration revealed subtype-specific pathways (e.g., ECM remodeling in R1, metabolic reprogramming in R2). Reclassification of pMMR tumors based on transcriptional similarity to R1 identified a New R3 subset with enhanced immune features and higher ICI response. Eight immune response-associated genes (e.g., CXCL10, CXCL11, ELN, GAD1, IL32, MT1E, OR2I1P, SLC3A1) were identified and validated by mIHC for predictive relevance.

CONCLUSIONS: This immune-based molecular framework refines risk stratification beyond conventional MMR categories, identifies ICI-sensitive subsets among both dMMR and pMMR tumors, and proposes candidate biomarkers for patient selection.

PMID:42147371 | PMC:PMC13171420 | DOI:10.21147/j.issn.1000-9604.2026.02.09

❌