❌

Normal view

APT-Agent: Automated Penetration Testing using Large Language Models

arXiv:2605.24949v1 Announce Type: cross Abstract: Penetration testing is essential to securing modern web infrastructures, yet traditional manual methods struggle to keep pace with their scale and complexity. Large Language Models (LLMs) offer new opportunities for automating these tasks, but existing approaches face two persistent challenges: hallucination of technical entities and insufficient long-term contextual memory. To address these issues, we present APT-Agent, a fully automated LLM-driven penetration testing framework that systematically orchestrates reconnaissance, exploitation, and exfiltration. APT-Agent introduces a hybrid rectification module to recover hallucinated commands and a command-specific memory architecture to preserve operational context across multi-step attack sequences. We evaluate our APT-Agent on Metasploitable 2 against seven vulnerable services spanning web, database, and network protocols. APT-Agent achieves an 84.29% end-to-end exploitation success rate, compared to 48.57% (Script Kiddie) and 18.57% (PentestGPT) under matched conditions. By reducing cognitive burden and minimizing reliance on human intervention, APT-Agent represents a step toward scalable, reliable, and cognitively efficient automation for penetration testing.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents Ting Liu
    arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, a skill often needs to express more than task guidance: goals, input boundaries, permissions, human approval points, evidence requirements, output contracts, quality criteria, verification steps, and handoff rules. This paper proposes contractual skills, a GovernSpec-inspired design fram
     

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings, however, a skill often needs to express more than task guidance: goals, input boundaries, permissions, human approval points, evidence requirements, output contracts, quality criteria, verification steps, and handoff rules. This paper proposes contractual skills, a GovernSpec-inspired design framework for organizing SKILL.md files as readable task contracts while preserving lightweight skill discovery and progressive loading. The framework clarifies the boundary between contractual skills, GovernSpec YAML contracts, Model Context Protocol (MCP) surfaces, tool adapters, runtime guardrails, tracing, and evaluation systems. We evaluate the framework with three offline empirical studies. The first text-generation experiment covers three enterprise skills, fifteen synthetic tasks, four instruction conditions, and eight generation models, producing 960 outputs and 1680 cross-judge score records. The second study is a public-skill A/B expansion: eight public skills are compared with contractual rewrites across forty-eight synthetic tasks, six generation models, two repeats, 1152 outputs, and two complete judge files. In this setting, contractual skills raise mean quality from 4.692 to 4.914 and reduce critical-error rate from 0.083 to 0.013. The third study is an offline tool-calling challenge with eight models and 192 simulated tool-call records. The results suggest that contractual skills are best understood as a governance layer that makes task intent, boundaries, and acceptance criteria explicit, not as a standalone safety mechanism.

Pan-neurodegeneration proteomics reveals disease subtypes and molecular signatures

A pan-neurodegeneration atlas built from multilayer, deep proteomics of 2,279 brain samples across 6 major diseases integrates whole proteome, detergent-insoluble proteome, and posttranslational modifications to enable intra- and inter-disease comparisons to reveal disease-specific subtypes and dysregulated pathways, while identifying shared changes such as GPNMB upregulation and NPTX2 downregulation.

WNT7A correlates with immunosuppression and predicts adverse prognosis in lung adenocarcinoma: Potential implication of the NF-kappaB/CCL2 Axis

Cytokine. 2026 Apr 6;202:157144. doi: 10.1016/j.cyto.2026.157144. Online ahead of print.

ABSTRACT

BACKGROUND: The remodeling of the tumor immune microenvironment (TME) is a pivotal determinant of therapeutic efficacy and clinical outcome in Lung Adenocarcinoma (LUAD). While WNT signaling is a known oncogenic driver, the specific immunomodulatory role of WNT7A and its potential crosstalk with inflammatory pathways in LUAD remain to be fully elucidated. We sought to define the prognostic value of WNT7A and explore the molecular mechanisms by which it may foster an immunosuppressive TME.

METHODS: We performed a multi-omics analysis utilizing the TCGA-LUAD cohort (N = 508) and validated findings in an independent external cohort (GSE30219, N = 293). The prognostic significance of WNT7A was evaluated using Kaplan-Meier and multivariate Cox regression analyses. TME composition was dissected via ssGSEA, focusing on myeloid-derived suppressor cell (MDSC) infiltration. Mechanistic pathways were identified using Gene Set Enrichment Analysis (GSEA) and gene co-expression networks.

RESULTS: High WNT7A expression was identified as a significant predictor of poor Overall Survival (OS) in the TCGA cohort (P < 0.05) and validated in the external cohort (P < 0.05). Multivariate analysis confirmed WNT7A as an independent prognostic risk factor (HR = 1.085, P = 0.036). Immunologically, WNT7A expression was positively correlated with MDSC infiltration (R = 0.43, P < 0.001), suggesting a shift towards an immune-tolerant phenotype. Mechanistically, GSEA revealed a robust activation of inflammatory signaling in the high-WNT7A group. Specifically, the TNFA Signaling via NF-κB pathway was significantly enriched(NES = 2.52, P < 0.001). Consistent with this pathway activation, WNT7A showed a statistically significant positive correlation with CCL2 (P < 0.001), a critical chemokine for MDSC recruitment, implicating the NF-κB/CCL2 axis in this process.

CONCLUSION: WNT7A serves as a prognostic biomarker linked to immune evasion in LUAD, potentially by modulating the NF-κB/CCL2/MDSC axis. This study identifies WNT7A as a potential therapeutic target to remodel the immune microenvironment, providing a rationale for future investigations into WNT-targeted strategies to improve immunotherapy efficacy.

PMID:41946008 | DOI:10.1016/j.cyto.2026.157144

WNT7A correlates with immunosuppression and predicts adverse prognosis in lung adenocarcinoma: Potential implication of the NF-kappaB/CCL2 Axis

Cytokine. 2026 Apr 6;202:157144. doi: 10.1016/j.cyto.2026.157144. Online ahead of print.

ABSTRACT

BACKGROUND: The remodeling of the tumor immune microenvironment (TME) is a pivotal determinant of therapeutic efficacy and clinical outcome in Lung Adenocarcinoma (LUAD). While WNT signaling is a known oncogenic driver, the specific immunomodulatory role of WNT7A and its potential crosstalk with inflammatory pathways in LUAD remain to be fully elucidated. We sought to define the prognostic value of WNT7A and explore the molecular mechanisms by which it may foster an immunosuppressive TME.

METHODS: We performed a multi-omics analysis utilizing the TCGA-LUAD cohort (N = 508) and validated findings in an independent external cohort (GSE30219, N = 293). The prognostic significance of WNT7A was evaluated using Kaplan-Meier and multivariate Cox regression analyses. TME composition was dissected via ssGSEA, focusing on myeloid-derived suppressor cell (MDSC) infiltration. Mechanistic pathways were identified using Gene Set Enrichment Analysis (GSEA) and gene co-expression networks.

RESULTS: High WNT7A expression was identified as a significant predictor of poor Overall Survival (OS) in the TCGA cohort (P < 0.05) and validated in the external cohort (P < 0.05). Multivariate analysis confirmed WNT7A as an independent prognostic risk factor (HR = 1.085, P = 0.036). Immunologically, WNT7A expression was positively correlated with MDSC infiltration (R = 0.43, P < 0.001), suggesting a shift towards an immune-tolerant phenotype. Mechanistically, GSEA revealed a robust activation of inflammatory signaling in the high-WNT7A group. Specifically, the TNFA Signaling via NF-κB pathway was significantly enriched(NES = 2.52, P < 0.001). Consistent with this pathway activation, WNT7A showed a statistically significant positive correlation with CCL2 (P < 0.001), a critical chemokine for MDSC recruitment, implicating the NF-κB/CCL2 axis in this process.

CONCLUSION: WNT7A serves as a prognostic biomarker linked to immune evasion in LUAD, potentially by modulating the NF-κB/CCL2/MDSC axis. This study identifies WNT7A as a potential therapeutic target to remodel the immune microenvironment, providing a rationale for future investigations into WNT-targeted strategies to improve immunotherapy efficacy.

PMID:41946008 | DOI:10.1016/j.cyto.2026.157144

Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents

arXiv:2604.04157v1 Announce Type: new Abstract: Theory of Mind (ToM) -- the ability to model others' mental states -- is fundamental to human social cognition. Whether large language models (LLMs) can develop ToM has been tested exclusively through static vignettes, leaving open whether ToM-like reasoning can emerge through dynamic interaction. Here we report that autonomous LLM agents playing extended sessions of Texas Hold'em poker progressively develop sophisticated opponent models, but only when equipped with persistent memory. In a 2x2 factorial design crossing memory (present/absent) with domain knowledge (present/absent), each with five replications (N = 20 experiments, ~6,000 agent-hand observations), we find that memory is both necessary and sufficient for ToM-like behavior emergence (Cliff's delta = 1.0, p = 0.008). Agents with memory reach ToM Level 3-5 (predictive to recursive modeling), while agents without memory remain at Level 0 across all replications. Strategic deception grounded in opponent models occurs exclusively in memory-equipped conditions (Fisher's exact p

Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging

arXiv:2604.01538v1 Announce Type: cross Abstract: Large language models have been adopted in the medical domain for clinical documentation to reduce clinician burden. However, studies have reported that LLMs often "forget" a significant amount of instruction-following ability when fine-tuned using a task-specific medical dataset, a critical challenge in adopting general-purpose LLMs for clinical applications. This study presents a model merging framework to efficiently adapt general-purpose LLMs to the medical domain by countering this forgetting issue. By merging a clinical foundation model (GatorTronLlama) with a general instruct model (Llama-3.1-8B-Instruct) via interpolation-based merge methods, we seek to derive a domain-adapted model with strong performance on clinical tasks while retaining instruction-following ability. Comprehensive evaluation across medical benchmarks and five clinical generation tasks (e.g., radiology and discharge summarization) shows that merged models can effectively mitigate catastrophic forgetting, preserve clinical domain expertise, and retain instruction-following ability. In addition, our model merging strategies demonstrate training efficiency, achieving performance on par with fully fine-tuned baselines under severely constrained supervision (e.g., 64-shot vs. 256-shot). Consequently, weight-space merging constitutes a highly scalable solution for adapting open-source LLMs to clinical applications, facilitating broader deployment in resource-constrained healthcare environments.

Dual Antitumor Effects of Solanum incanum L. in Lung and Colonic Adenocarcinomas via TRAIL-Mediated Extrinsic Apoptotic Pathway: A Network Pharmacology and Multi-Omics Investigation

Curr Med Chem. 2026 Mar 17. doi: 10.2174/0109298673415187251212085809. Online ahead of print.

ABSTRACT

BACKGROUND: Solanum incanum L. (S. incanum) is a Traditional Chinese Medicine (TCM) known for its heat-clearing and detoxifying properties. This study evaluates the anti-cancer potential of S. incanum in lung and colorectal adenocarcinoma. Also, we aim to confirm the connection between the lung and the large intestine, as proposed in TCM theory.

METHODS: This study employed network pharmacology analysis, MTT and morphological analysis, Western blotting, and molecular docking to investigate the anti-cancer effects of S. incanum, particularly through its activation of the TRAIL-mediated extrinsic apoptotic pathway.

RESULTS: This study identified 15 active compounds in S. incanum, with glycosylated alkaloids such as α-solanine and solasonine highlighted as key bioactive compounds. KEGG and GO analyses revealed that S. incanum influences apoptosis-related pathways, particularly the extrinsic apoptotic pathway, regulating mediators such as CASP3, CASP8, and various TNF- related receptors. The anticancer potential was assessed through cellular studies using A549 and HT29 cell lines, with MTT cell viability assays showing that α-solanine induces apoptosis, with IC50 values between 20 and 30 μM. Western blot analysis demonstrated that S. incanum activates apoptosis via the TRAIL-mediated extrinsic apoptotic pathway, promoting CASP8, CASP3, and PARP1 cleavage in a dose-dependent manner, affecting DNA synthesis. Molecular docking further identified α-solanine and solasonine as compounds that bind to critical proteins in the TRAIL-related apoptotic pathway, confirming their roles in promoting apoptosis.

DISCUSSION: The shared activation of the TRAIL-mediated apoptotic pathway in both lung and colorectal adenocarcinoma models suggests a common molecular mechanism. These findings provide experimental evidence linking traditional ethnopharmacological concepts with apoptosis-related anti-cancer activity.

CONCLUSIONS: This study emphasizes that the main compound of S. incanum, α-solanine, supports the ethnopharmacological concept of "Exterior-interior Pairing of the Lung and Large Intestine", thereby demonstrating that both lung and colorectal adenocarcinomas share a common anticancer pathway through a TRAIL-related apoptotic mechanism.

PMID:41863159 | DOI:10.2174/0109298673415187251212085809

Human-specific features of the cerebellum and ZP2-regulated synapse development

Human-specific transcriptomic and regulatory features are present in the cerebellum, with ZP2 playing a key role in synapse regulation. ZP2 expression is induced by pontine mossy fibers, leading to decreased synaptic proteins and neuronal activity, which provides insights into the evolutionary development of the human cerebellum.

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Jun 1;50(6):695-704. doi: 10.1097/PAS.0000000000002533. Epub 2026 Mar 13.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2 :: NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases ( NFATC2::NUTM2A , n=2; NFATC2::NUTM2E , n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2 -associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

NFATC2::NUTM2 Fusion Defines a Novel Primary Pulmonary Epithelial Tumor With a Distinctive Immunophenotype

Am J Surg Pathol. 2026 Mar 13. doi: 10.1097/PAS.0000000000002533. Online ahead of print.

ABSTRACT

With the application of molecular techniques in pathologic diagnosis, several novel primary pulmonary epithelial tumors have been continuously discovered and classified under the WHO classification of thoracic tumors. Recently, a pulmonary tumor with NFATC2::NUTM2B fusion was first documented, but the spectrum of NFATC2::NUTM2 fusion variants and their associated pathologic features remains incompletely characterized. Coincidentally, we also found and described 6 primary pulmonary tumors harboring recurrent NFATC2::NUTM2A/E fusions through integrated genomic analysis. These patients, including 4 females and 2 males, with a median age of 53 years, presented with incidentally detected peripheral lung nodules composed of monotonous epithelioid cells arranged in cords, nests, and trabeculae within a prominent desmoplastic stroma. All tumors exhibited a consistent immunophenotype: CK5/6+/GATA3+/calponin+/EMA+/DOG1 (perinuclear dot-like staining)/p63-. High-throughput chromosome conformation capture (Hi-C) analysis showed the structural variation of NFATC2::NUTM2E in all 6 cases, whereas RNA sequencing detected the fusion transcripts in 5 cases (NFATC2::NUTM2A, n=2; NFATC2::NUTM2E, n=3). Ultrastructural examination of 1 case suggested epithelial differentiation. All patients remained disease-free after complete resection (median follow-up: 24 mo; range: 9 to 41 mo). These findings define a novel primary pulmonary tumor entity driven by NFATC2::NUTM2 fusions, and characterized by a distinctive immunophenotype, expanding the spectrum of NUTM2-associated neoplasms. Our study underscores the utility of multiomics approaches for characterizing rare neoplasms and provides a diagnostic framework for this entity.

PMID:41821426 | DOI:10.1097/PAS.0000000000002533

T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning

arXiv:2603.03790v1 Announce Type: cross Abstract: Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to enhance text-processing performance? To explore it, in this work, we first introduce Structure of Thought (SoT), a prompting technique that explicitly guides models to construct intermediate text structures, consistently boosting performance across eight tasks and three model families. Building upon this insight, we present T2S-Bench, the first benchmark designed to evaluate and improve text-to-structure capabilities of models. T2S-Bench includes 1.8K samples across 6 scientific domains and 32 structural types, rigorously constructed to ensure accuracy, fairness, and quality. Evaluation on 45 mainstream models reveals substantial improvement potential: the average accuracy on the multi-hop reasoning task is only 52.1%, and even the most advanced model achieves 58.1% node accuracy in end-to-end extraction. Furthermore, on Qwen2.5-7B-Instruct, SoT alone yields an average +5.7% improvement across eight diverse text-processing tasks, and fine-tuning on T2S-Bench further increases this gain to +8.6%. These results highlight the value of explicit text structuring and the complementary contributions of SoT and T2S-Bench. Dataset and eval code have been released at https://t2s-bench.github.io/T2S-Bench-Page/.

IDER: IDempotent Experience Replay for Reliable Continual Learning

arXiv:2603.00624v2 Announce Type: replace-cross Abstract: Catastrophic forgetting, the tendency of neural networks to forget previously learned knowledge when learning new tasks, has been a major challenge in continual learning (CL). To tackle this challenge, CL methods have been proposed and shown to reduce forgetting. Furthermore, CL models deployed in mission-critical settings can benefit from uncertainty awareness by calibrating their predictions to reliably assess their confidences. However, existing uncertainty-aware continual learning methods suffer from high computational overhead and incompatibility with mainstream replay methods. To address this, we propose idempotent experience replay (IDER), a novel approach based on the idempotent property where repeated function applications yield the same output. Specifically, we first adapt the training loss to make model idempotent on current data streams. In addition, we introduce an idempotence distillation loss. We feed the output of the current model back into the old checkpoint and then minimize the distance between this reprocessed output and the original output of the current model. This yields a simple and effective new baseline for building reliable continual learners, which can be seamlessly integrated with other CL approaches. Extensive experiments on different CL benchmarks demonstrate that IDER consistently improves prediction reliability while simultaneously boosting accuracy and reducing forgetting. Our results suggest the potential of idempotence as a promising principle for deploying efficient and trustworthy continual learning systems in real-world applications.Our code is available at https://github.com/YutingLi0606/Idempotent-Continual-Learning.

Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos

arXiv:2602.18466v1 Announce Type: cross Abstract: K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these interactions has made automated analysis elusive. Existing benchmarks for classroom discourse focus primarily on mathematics and rely solely on transcripts, overlooking the visual artifacts and model-based reasoning emphasized by the Next Generation Science Standards (NGSS). We address this gap with SciIBI, the first video benchmark for analyzing science classroom discourse, featuring 113 NGSS-aligned clips annotated with Core Instructional Practices (CIP) and sophistication levels. By evaluating eight state-of-the-art LLMs and Multimodal LLMs, we reveal fundamental limitations: current models struggle to distinguish pedagogically similar practices, suggesting that CIP coding requires instructional reasoning beyond surface pattern matching. Furthermore, adding video input yields inconsistent gains across architectures. Crucially, our evidence-based evaluation reveals that models often succeed through surface shortcuts rather than genuine pedagogical understanding. These findings establish science classroom discourse as a challenging frontier for multimodal AI and point toward human-AI collaboration, where models retrieve evidence to accelerate expert review rather than replace it.

Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering

arXiv:2602.19569v1 Announce Type: cross Abstract: Question Answering over Temporal Knowledge Graphs (TKGQA) has attracted growing interest for handling time-sensitive queries. However, existing methods still struggle with: 1) weak incorporation of temporal constraints in question representation, causing biased reasoning; 2) limited ability to perform explicit multi-hop reasoning; and 3) suboptimal fusion of language and graph representations. We propose a novel framework with temporal-aware question encoding, multi-hop graph reasoning, and multi-view heterogeneous information fusion. Specifically, our approach introduces: 1) a constraint-aware question representation that combines semantic cues from language models with temporal entity dynamics; 2) a temporal-aware graph neural network for explicit multi-hop reasoning via time-aware message passing; and 3) a multi-view attention mechanism for more effective fusion of question context and temporal graph knowledge. Experiments on multiple TKGQA benchmarks demonstrate consistent improvements over multiple baselines.
❌