❌

Reading view

FrontierChallenge: Evaluating Scientific Workflow Completion

arXiv:2608.24979v2 Announce Type: replace Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. Complementary HDS6 process scores correlate strongly with task outcomes, supporting FrontierChallenge as a benchmark of Heavy Duty Solver capabilities. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
  •  

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

arXiv:2605.24834v1 Announce Type: cross Abstract: Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversarial jailbreak attacks that disguise malicious intent through role-play scenarios, fictional framing, and indirect requests. We present Reflect-Guard, a method that augments LLM-based safety classifiers with chain-of-thought self-reflection capabilities through parameter-efficient fine-tuning. Our approach distills analytical reasoning from GPT-4o-mini into structured reflection annotations, then trains Llama-Guard-3-8B via QLoRA to generate logical self-reflections before issuing safety verdicts. Using only 1000 training examples and updating just 0.5% of model parameters (~42M), Reflect-Guard achieves substantial improvements on two challenging benchmarks. On WildGuardTest, F1 score improves from 0.770 to 0.842 (+7.2 pp), with recall on adversarial prompts increasing from 0.513 to 0.921 (+40.8 pp). On JailbreakBench, the attack success rate drops from 10.3% to 1.8%, representing an 82.5% relative reduction. These gains are especially pronounced on adversarial inputs, where the explicit reasoning step enables the model to see through obfuscation techniques that defeat standard pattern-matching approaches. Our results demonstrate that teaching safety classifiers to reason about adversarial intent, rather than simply classify surface patterns, is a promising direction for robust LLM safety.
  •  

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

arXiv:2602.02474v2 Announce Type: replace-cross Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixed procedures hard-code human priors about what to store and how to revise memory, making them rigid under diverse interaction patterns and inefficient on long histories. To this end, we present \textbf{MemSkill}, which reframes these operations as learnable and evolvable memory skills, structured and reusable routines for extracting, consolidating, and pruning information from interaction traces. Inspired by the design philosophy of agent skills, MemSkill employs a \emph{controller} that learns to select a small set of relevant skills, paired with an LLM-based \emph{executor} that produces skill-guided memories. Beyond learning skill selection, MemSkill introduces a \emph{designer} that periodically reviews hard cases where selected skills yield incorrect or incomplete memories, and evolves the skill set by proposing refinements and new skills. Together, MemSkill forms a closed-loop procedure that improves both the skill-selection policy and the skill set itself. Experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate that MemSkill improves task performance over strong baselines and generalizes well across settings. Further analyses shed light on how skills evolve, offering insights toward more adaptive, self-evolving memory management for LLM agents.
  •  

Stromal ACTA2 Counteracts TCDD-Induced Hepatocarcinogenesis via Suppression of the PI3K-AKT-mTOR Pathway

J Hepatocell Carcinoma. 2026 May 10;13:586916. doi: 10.2147/JHC.S586916. eCollection 2026.

ABSTRACT

PURPOSE: 2,3,7,8-Tetrachlorodibenzo-p-dioxin (TCDD) is a persistent environmental pollutant that promotes hepatocellular carcinoma (HCC) through non-genotoxic mechanisms. However, stromal regulatory factors that counteract its tumor-promoting effects remain poorly defined. This study aimed to elucidate the role of actin alpha-2 (ACTA2) in TCDD-associated hepatocarcinogenesis.

METHODS: An integrative strategy combining network toxicology, Mendelian randomization, multi-omics and single-cell analyses, molecular docking and molecular dynamics simulations, along with in vitro experiments, was employed to investigate the functional role of ACTA2.

RESULTS: ACTA2 was identified as a stromal-associated factor linked to reduced HCC risk and improved patient survival. Single-cell and multi-omics analyses revealed that ACTA2 is predominantly expressed in hepatic stellate cells and fibroblast-like populations, reflecting tumor microenvironment composition rather than tumor cell-intrinsic expression. Functional enrichment analyses indicated that ACTA2 is associated with extracellular matrix remodeling and PI3K-AKT signaling. Molecular simulations demonstrated stable binding of TCDD to ACTA2 (ΔG_bind ≈ -7.05 kcal/mol), suggesting potential structural perturbation. In vitro experiments showed that TCDD downregulated ACTA2 expression, promoted proliferation of LX-2 and cancer-associated fibroblasts (CAFs), and activated PI3K-AKT-mTOR signaling, whereas ACTA2 overexpression attenuated these effects.

CONCLUSION: ACTA2 acts as a context-dependent stromal regulator that modulates PI3K-AKT-mTOR signaling in TCDD-induced hepatocarcinogenesis. These findings highlight the importance of stromal remodeling in environmental carcinogenesis and suggest ACTA2 as a potential biomarker and therapeutic target in dioxin-associated HCC.

PMID:42148320 | PMC:PMC13175077 | DOI:10.2147/JHC.S586916

  •  

M<sup>6</sup>A-dependent upregulation of ZNF460 promotes epithelial-mesenchymal transition and metastasis of gastric cancer through a histone modification-mediated positive feedback loop

Oncogene, Published online: 02 April 2026; doi:10.1038/s41388-026-03755-3

M6A-dependent upregulation of ZNF460 promotes epithelial-mesenchymal transition and metastasis of gastric cancer through a histone modification-mediated positive feedback loop
  •  

Deciphering lung adenocarcinoma heterogeneity: a multi-omics approach reveals nuclear division fibroblasts as prognosticators and therapeutic targets

J Transl Med. 2026 Mar 20. doi: 10.1186/s12967-026-08022-3. Online ahead of print.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) is a predominant contributor to cancer‑related mortality globally. Lung‑associated fibroblasts (LAFs) are intricately linked to tumorigenesis and the tumor microenvironment (TME), but their heterogeneity and prognostic relevance in LUAD remain incompletely understood. This study aimed to systematically characterize LAF subsets across the spectrum of pulmonary disease, identify LAF subpopulations associated with LUAD prognosis, and construct a robust LAF‑based prognostic signature.

METHODS: We employed a multi-omics approach, leveraging bulk RNA data of 2719 patients from 19 LUAD cohorts, single-cell RNA (scRNA) sequencing data of 368,904 cells from 93 samples, and spatial transcriptomics data of 15,673 spots from 6 samples to characterize the landscape of LAFs across various stages of pulmonary disease. We employed multiple advanced machine learning algorithms to construct and validate a robust nuclear division LAFs (nLAFs) risk score (nLRS) prediction model.

RESULTS: We observed a dynamic and gradual increase in the proportion of LAFs during the progression of LUAD. Throughout this process, we identified nine LAFs subtypes and found nLAFs are significantly associated with the prognosis of LUAD. Utilizing 100 machine learning algorithm combinations and integrating nLAFs marker genes, we developed a five gene based nLRS model, which demonstrated superior performance than other 49 published models in predicting clinical outcomes for LUAD. Additionally, we observed distinct biological functions and immune cell infiltration in the TME between high and low nLRS groups. Exploratory analysis of pan-cancer immunotherapy cohorts suggested that patients with high nLRS scores may exhibit resistance to immunotherapy in some cancer types, but prospective validation in LUAD-specific cohorts is required. Conversely, high nLRS patients displayed increased sensitivity to chemotherapeutic and targeted therapies in preclinical models.

CONCLUSION: Our study introduces a candidate five-gene signature derived from nLAFs that may serve as a robust prognostic biomarker pending prospective validation, offering insights into personalized therapeutic strategies for LUAD patients.

PMID:41862916 | DOI:10.1186/s12967-026-08022-3

  •  

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

  •  

Multi-Omics Characterization of Lactate-Associated Molecular Subtypes in Lung Cancer Suggests a Role for DKK1 in Lactate-Linked Migration, Invasion, and Lactylation Programs

Cancers (Basel). 2026 Feb 25;18(5):735. doi: 10.3390/cancers18050735.

ABSTRACT

BACKGROUND: Lactate accumulation is increasingly recognized as a feature of tumor metabolic reprogramming that can coincide with immune dysregulation and aggressive phenotypes. The prognostic and immunologic relevance of lactate-associated heterogeneity in lung cancer remains to be clarified.

METHODS: We curated lactate-related genes and identified prognostic candidates in lung cancer cohorts. Consensus clustering was applied to define lactate-associated molecular subtypes, followed by characterization of survival and tumor microenvironment features. A LASSO-based gene signature was developed to generate an individual-level risk score and an integrated nomogram. Multi-omics analyses were used to evaluate concordance between transcriptomic and proteomic alterations. Single-cell transcriptomic data were analyzed to explore cellular heterogeneity in lactate-related programs. In vitro assays evaluated the response of candidate genes to lactate exposure and assessed cell migration and invasion under proliferation-inhibited conditions after genetic perturbation.

RESULTS: Two lactate-associated molecular subtypes were identified with distinct overall survival and divergent immune microenvironment features. Subtype 1 was associated with better outcomes and a more immune-inflamed profile, whereas Subtype 2 was associated with poorer outcomes and a myeloid-enriched, immunosuppressive contexture. Pathway analyses indicated subtype-associated differences in extracellular matrix-related processes and apoptosis-associated signaling. We developed an 11-gene prognostic signature and nomogram that stratified patients by risk across TCGA and GEO cohorts. Multi-omics integration highlighted ANLN, FGA, and DKK1 as consistently dysregulated at both transcript and protein levels. Among these candidates, DKK1 showed lactate-responsive induction in vitro. DKK1 perturbation altered lactate-enhanced migratory and invasive phenotypes and was accompanied by changes in intracellular lactate levels and global protein lactylation, supporting a potential feedforward relationship between lactate exposure, DKK1 expression, and lactylation.

CONCLUSIONS: This study characterizes lactate-associated molecular heterogeneity in lung cancer and provides a lactate-related subtype framework and prognostic risk model for patient stratification. The findings nominate DKK1 as a lactate-responsive candidate linked to migration/invasion phenotypes and lactate/lactylation changes in vitro.

PMID:41827671 | PMC:PMC12985219 | DOI:10.3390/cancers18050735

  •  

Understand Then Memory: A Cognitive Gist-Driven RAG Framework with Global Semantic Diffusion

arXiv:2602.15895v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in LLMs by incorporating external knowledge. However, the inherent discrete representation of text in existing frameworks often results in a loss of semantic integrity, leading to retrieval deviations. Inspired by the human episodic memory mechanism, we propose CogitoRAG, a RAG framework that simulates human cognitive memory processes. The core of this framework lies in the extraction and evolution of the Semantic Gist. During the offline indexing stage, CogitoRAG first deduces unstructured corpora into gist memory corpora, which are then transformed into a multi-dimensional knowledge graph integrating entities, relational facts, and memory nodes. In the online retrieval stage, the framework handles complex queries via Query Decomposition Module that breaks them into comprehensive sub-queries, mimicking the cognitive decomposition humans employ for complex information. Subsequently, Entity Diffusion Module performs associative retrieval across the graph, guided by structural relevance and an entity-frequency reward mechanism. Furthermore, we propose the CogniRank algorithm, which precisely reranks candidate passages by fusing diffusion-derived scores with semantic similarity. The final evidence is delivered to the generator in a passage-memory pairing format, providing high-density information support. Experimental results across five mainstream QA benchmarks and multi-task generation on GraphBench demonstrate that CogitoRAG significantly outperforms state-of-the-art RAG methods, showcasing superior capabilities in complex knowledge integration and reasoning.
  •  

From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning

arXiv:2603.03825v1 Announce Type: cross Abstract: The cold-start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attention Score (VAS), an attention-based metric that quantifies how much a model attends to visual tokens. We find that reasoning performance is strongly correlated with VAS (r=0.9616): models with higher VAS achieve substantially stronger multimodal reasoning. Surprisingly, multimodal cold-start fails to elevate VAS, resulting in attention distributions close to the base model, whereas text-only cold-start leads to a clear increase. We term this counter-intuitive phenomenon Lazy Attention Localization. To validate its causal role, we design training-free interventions that directly modulate attention allocation during inference, performance gains of 1$-$2% without any retraining. Building on these insights, we further propose Attention-Guided Visual Anchoring and Reflection (AVAR), a comprehensive cold-start framework that integrates visual-anchored data synthesis, attention-guided objectives, and visual-anchored reward shaping. Applied to Qwen2.5-VL-7B, AVAR achieves an average gain of 7.0% across 7 multimodal reasoning benchmarks. Ablation studies further confirm that each component of AVAR contributes step-wise to the overall gains. The code, data, and models are available at https://github.com/lrlbbzl/Qwen-AVAR.
  •  

CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

arXiv:2602.12268v2 Announce Type: replace Abstract: AI agents are increasingly used to solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. However, applying reinforcement learning to such settings remains difficult: realistic objectives often lack verifiable rewards and instead emphasize open-ended behaviors; moreover, RL for multi-turn, multi-step agentic tool use is still underexplored; and building and maintaining executable tool environments is costly, limiting scale and coverage. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, our method adopts a strategy of sparse reward assignment but dense evaluation criteria. Training is performed in a scalable LLM-simulated tool environment, avoiding heavy engineering for large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on tau^-Bench, by 10 points on BFCL-V4, and by 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus provides a scalable recipe for optimizing multi-turn, multi-step tool-using agents without relying on verifiable rewards. Code provided by the open-source community: https://github.com/namezhenzhang/CM2-RLCR-Tool-Agent.
  •  
❌