❌

Reading view

Pan-cancer landscape of protein kinase D3: An integrative TCGA multi-omics analysis of clinical, molecular, and immunological roles

PLoS One. 2026 Apr 3;21(4):e0346173. doi: 10.1371/journal.pone.0346173. eCollection 2026.

ABSTRACT

Cancer remains a leading cause of mortality worldwide and a significant barrier to improving quality of life across all populations. The protein kinase D family, including PRKD3, has been demonstrated to play a crucial role in cancer development through its involvement in regulating key cellular processes. Although growing evidence highlights the role of PRKD3 in the tumorigenesis of certain cancers, a comprehensive pan-cancer analysis of PRKD3 remains unavailable. To address this, we performed an integrative pan-cancer analysis of PRKD3 using multi-omics datasets from The Cancer Genome Atlas, the Genotype-Tissue Expression project, and cBioPortal. We examined PRKD3 expression, copy number variation, mutation, and DNA methylation, and evaluated their associations with clinicopathological features, patient survival, and diagnostic potential across 33 cancer types. Immune relevance was further assessed through correlations with immune infiltration, checkpoint gene expression, and immunotherapy response-related genomic biomarkers. Our results revealed that PRKD3 expression was highly heterogeneous, showing significant upregulation in liver cancer, gastric cancer, and adrenocortical carcinoma, and downregulation in others. Elevated expression was consistently associated with poor prognosis and increased stromal, neutrophil, and cancer-associated fibroblast infiltration in adrenocortical carcinoma, liver cancer, and stomach cancer, whereas paradoxical associations with favorable outcomes were observed in kidney clear cell carcinoma. PRKD3 expression also correlated with immune checkpoint molecules including PD-1, PD-L1, and CTLA-4, supporting an immunosuppressive role, while context-dependent associations with TMB and MSI highlighted its potential influence on tumor immunogenicity and responsiveness to immune checkpoint blockade. Collectively, these findings identify PRKD3 as a potential context-dependent modulator of tumor biology, prognosis, and immune interactions, underscoring its potential as a biomarker of diagnostic, prognostic, and therapeutic relevance in precision oncology.

PMID:41931575 | PMC:PMC13048501 | DOI:10.1371/journal.pone.0346173

  •  

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

  •  

Single-Cell and Multi-Omics-Based Characterization of Gastric Cancer Identifies TPP1 as a Potential Target for Gastric Cancer Progression and Treatment

Oncol Res. 2026 Mar 23;34(4):27. doi: 10.32604/or.2026.070208. eCollection 2026.

ABSTRACT

BACKGROUND: Cancer-associated fibroblasts (CAFs) play critical roles in tumor progression and immunosuppression; however, their contribution to the functional classification and personalized treatment of gastric cancer remains poorly defined. This study aimed to identify effective therapeutic targets to facilitate individualized treatment strategies for patients with gastric cancer.

METHODS: Single-cell and bulk transcriptomic analyses were integrated to characterize gastric cancer fibroblasts. "Seurat", "Slingshot", and "CellChat" were used for dimensionality reduction, trajectory inference, and cell-cell communication analyses, respectively. Key metastasis-associated fibroblast modules were identified using High-dimensional weighted gene co-expression network analysis (hdWGCNA) to construct a prognostic model, which was further evaluated for immune infiltration, therapeutic response, and mutational features. The expression and function of the core gene tripeptidyl peptidase 1 (TPP1) were validated through immunoblotting, PCR, and functional assays.

RESULTS: Eight fibroblast subpopulations associated with gastric cancer metastasis exhibited distinct differentiation trajectories and transcriptional heterogeneity. Prognostic analysis indicated that metastasis-associated fibroblasts correlated with poor clinical outcomes. The high-risk subgroup showed marked immunosuppression, resistance to immunotherapy, and reduced mutational burden, with tumor progression-related pathways significantly enriched in this group. In vitro experiments further confirmed that TPP1 knockdown suppressed gastric cancer cell metastasis, invasion, and clonogenic capacity while inducing apoptosis.

CONCLUSION: This study characterized the heterogeneity of gastric cancer-associated fibroblasts using single-cell transcriptomic analysis and established a prognostic model based on metastasis-related fibroblast markers. The model demonstrated strong predictive performance for patient prognosis, immune landscape, and immunotherapy response. Furthermore, the findings highlighted the pivotal role of TPP1 in gastric cancer progression and its potential as a therapeutic target.

PMID:41930144 | PMC:PMC13040347 | DOI:10.32604/or.2026.070208

  •  

Interpretable Machine Learning to Understand Wildfire Toxicity: Bridging Chemicals, Omics, and Toxicological Outcomes via Symbolic Regression with Novel Feature Scoring

Chem Res Toxicol. 2026 Apr 3. doi: 10.1021/acs.chemrestox.5c00440. Online ahead of print.

ABSTRACT

Wildfire smoke exposures are increasingly common, consisting of complex mixtures of gases and particulates known to cause diverse pulmonary health effects. While health outcomes are regularly studied, quantitative links between smoke chemical composition and toxicological outcomes remain poorly defined, limiting interpretation of wildfire smoke health risks. This study explores symbolic regression (SR) as an interpretable artificial intelligence/machine learning method to generate closed-form mathematical models linking chemical exposure to biological responses relevant to wildfire smoke. Prior to application on wildfire-relevant data sets, we benchmarked three Python-based SR packages on simulated data, assessing performance across varying noise levels and operator complexities. Insights from these simulation tests, such as the importance of including necessary operators, were incorporated when applying SR to lab-generated wildland fire exposure-toxicity data. This data set included chemical characterizations of biomass smoke exposures and corresponding pulmonary responses in female CD-1 mice (n = 60). Specifically, we evaluated the ability to predict a lung injury marker using (1) targeted measures of over 80 chemicals measured in smoke (RMSE = 17.57 mg/mL) and (2) lung tissue measures of hundreds of transcripts (RMSE = 15.12 mg/mL). Resulting error metrics were comparable to Random Forest and XGBoost models. To aid model interpretation, we developed directional ensemble contribution scores (DECS), a novel feature importance scoring method that quantifies the direction and magnitude of predictor contributions across top-performing models. Expert toxicologists also contributed to model prioritization, integrating a "biologists-in-the-loop" approach. Results highlighted polycyclic aromatic hydrocarbons as drivers of lung injury and methoxyphenols as suppressors. Transcriptomic analyses highlighted a small set of genes, which have roles in metabolism, cell proliferation, immune regulation, and oncogenic processes, with MYC proto-oncogene (Myc) showing the strongest association. Overall, this study demonstrates SR and associated DECS as practical, interpretable tools for modeling environmental mixtures, such as wildfire smoke, and their toxicological effects.

PMID:41928614 | DOI:10.1021/acs.chemrestox.5c00440

  •  

Pan-cancer landscape of protein kinase D3: An integrative TCGA multi-omics analysis of clinical, molecular, and immunological roles

PLoS One. 2026 Apr 3;21(4):e0346173. doi: 10.1371/journal.pone.0346173. eCollection 2026.

ABSTRACT

Cancer remains a leading cause of mortality worldwide and a significant barrier to improving quality of life across all populations. The protein kinase D family, including PRKD3, has been demonstrated to play a crucial role in cancer development through its involvement in regulating key cellular processes. Although growing evidence highlights the role of PRKD3 in the tumorigenesis of certain cancers, a comprehensive pan-cancer analysis of PRKD3 remains unavailable. To address this, we performed an integrative pan-cancer analysis of PRKD3 using multi-omics datasets from The Cancer Genome Atlas, the Genotype-Tissue Expression project, and cBioPortal. We examined PRKD3 expression, copy number variation, mutation, and DNA methylation, and evaluated their associations with clinicopathological features, patient survival, and diagnostic potential across 33 cancer types. Immune relevance was further assessed through correlations with immune infiltration, checkpoint gene expression, and immunotherapy response-related genomic biomarkers. Our results revealed that PRKD3 expression was highly heterogeneous, showing significant upregulation in liver cancer, gastric cancer, and adrenocortical carcinoma, and downregulation in others. Elevated expression was consistently associated with poor prognosis and increased stromal, neutrophil, and cancer-associated fibroblast infiltration in adrenocortical carcinoma, liver cancer, and stomach cancer, whereas paradoxical associations with favorable outcomes were observed in kidney clear cell carcinoma. PRKD3 expression also correlated with immune checkpoint molecules including PD-1, PD-L1, and CTLA-4, supporting an immunosuppressive role, while context-dependent associations with TMB and MSI highlighted its potential influence on tumor immunogenicity and responsiveness to immune checkpoint blockade. Collectively, these findings identify PRKD3 as a potential context-dependent modulator of tumor biology, prognosis, and immune interactions, underscoring its potential as a biomarker of diagnostic, prognostic, and therapeutic relevance in precision oncology.

PMID:41931575 | PMC:PMC13048501 | DOI:10.1371/journal.pone.0346173

  •  

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

  •  

Presentation: Panel: Taking Architecture Out of the Echo Chamber

Andrew Harmel-Law and a panel of expert architects discuss the shifting practice of architecture in 2025. They explain strategies for communicating technical debt to stakeholders, the benefits of decentralized decision-making through ADRs, and the career paths of modern leaders. The panel shares insights on bridging the gap between mobile and backend teams to ensure a holistic system.

By Andrew Harmel-Law, Cat Morris, Diana Montalion, Shana Dacres-Lawrence, Vanessa Formicola, Elena Stojmilova, Peter Hunter
  •  

Runtime Burden Allocation for Structured LLM Routing in Agentic Expert Systems: A Full-Factorial Cross-Backend Methodology

arXiv:2604.01235v1 Announce Type: new Abstract: Structured LLM routing is often treated as a prompt-engineering problem. We argue that it is, more fundamentally, a systems-level burden-allocation problem. As large language models (LLMs) become core control components in agentic AI systems, reliable structured routing must balance correctness, latency, and implementation cost under real deployment constraints. We show that this balance is shaped not only by prompts or schemas, but also by how structural work is allocated across the generation stack: whether output structure is emitted directly by the model, compressed during transport, or reconstructed locally after generation. We evaluate this formulation through a comprehensive full-factorial benchmark covering 48 deployment configurations and 15,552 requests across OpenAI, Gemini, and Llama backends. Our central finding is consequential: there is no universal best routing mode. Instead, backend-specific interaction effects dominate performance. Modes that remain highly reliable on Gemini and OpenAI can suffer substantial correctness degradation on Llama, while efficiency gains from compressed realization are strongly backend-dependent. Rather than presenting another isolated model comparison, this work contributes a deployable framework for reasoning about structured routing under heterogeneous backend conditions. We provide a cross-backend evaluation methodology and practical deployment guidance for navigating the correctness-cost-latency frontier in production-grade agentic expert systems.
  •  

Parallelized Hierarchical Connectome: A Spatiotemporal Recurrent Framework for Spiking State-Space Models

arXiv:2604.01295v1 Announce Type: new Abstract: This work presents the Parallelized Hierarchical Connectome (PHC), a general framework that upgrades temporal-only State-Space Models (SSMs) into spatiotemporal recurrent networks. Conventional SSMs achieve high-speed sequence processing through parallel scans, yet are limited to temporal recurrence without lateral or feedback interactions within a single timestep. PHC maps the diagonal SSM core to a shared Neuron Layer and inter-neuronal communication to a shared Synapse Layer, where neurons are partitioned into hierarchical regions governed by the connectome topology. A Multi-Transmission Loop enables intra-slice spatial recurrence, allowing signals to propagate across the hierarchical connectome within each temporal window while preserving O(logT) parallelism. This framework enables integration of neuro-physical priors typically intractable for standard SSMs, including adaptive leaky integrate-and-fire dynamics, Dale's Law, short-term plasticity, and reward-modulated spike-timing-dependent plasticity. The framework is instantiated as PHCSSM, the first model to unify recurrent spiking neural network dynamics with diagonal SSM parallelism while enforcing all five biological constraints and learnable lateral connections within a fully parallelizable training pipeline. Empirical results on physiological benchmarks from the UEA multivariate time-series archive demonstrate that PHCSSM achieves performance competitive with state-of-the-art SSMs while reducing parameter complexity from Theta(D^2 L) for L-layer stacked architectures to Theta(D^2). These findings suggest that biologically grounded inductive biases offer a principled route to parameter-efficient sequence modeling, opening diagonal SSMs to spatiotemporal recurrence and enabling fully parallelizable recurrent spiking neural network training.
  •  

IDEA2: Expert-in-the-loop competency question elicitation for collaborative ontology engineering

arXiv:2604.01344v1 Announce Type: new Abstract: Competency question (CQ) elicitation represents a critical but resource-intensive bottleneck in ontology engineering. This foundational phase is often hampered by the communication gap between domain experts, who possess the necessary knowledge, and ontology engineers, who formalise it. This paper introduces IDEA2, a novel, semi-automated workflow that integrates Large Language Models (LLMs) within a collaborative, expert-in-the-loop process to address this challenge. The methodology is characterised by a core iterative loop: an initial LLM-based extraction of CQs from requirement documents, a co-creational review and feedback phase by domain experts on an accessible collaborative platform, and an iterative, feedback-driven reformulation of rejected CQs by an LLM until consensus is achieved. To ensure transparency and reproducibility, the entire lifecycle of each CQ is tracked using a provenance model that captures the full lineage of edits, anonymised feedback, and generation parameters. The workflow was validated in 2 real-world scenarios (scientific data, cultural heritage), demonstrating that IDEA2 can accelerate the requirements engineering process, improve the acceptance and relevance of the resulting CQs, and exhibit high usability and effectiveness among domain experts. We release all code and experiments at https://github.com/KE-UniLiv/IDEA2
  •  

Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

arXiv:2604.01363v1 Announce Type: new Abstract: We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and (ii) rising tides where the increase in AI capabilities is more continuous and broad-based. We test for these effects in preliminary evidence from an ongoing evaluation of AI capabilities across over 3,000 broad-based tasks derived from the U.S. Department of Labor O*NET categorization that are text-based and thus LLM-addressable. Based on more than 17,000 evaluations by workers from these jobs, we find little evidence of crashing waves (in contrast to recent work by METR), but substantial evidence that rising tides are the primary form of AI automation. AI performance is high and improving rapidly across a wide range of tasks. We estimate that, in 2024-Q2, AI models successfully complete tasks that take humans approximately 3-4 hours with about a 50% success rate, increasing to about 65% by 2025-Q3. If recent trends in AI capability growth persist, this pace of AI improvement implies that LLMs will be able to complete most text-related tasks with success rates of, on average, 80%-95% by 2029 at a minimally sufficient quality level. Achieving near-perfect success rates at this quality level or comparable success rates at superior quality would require several additional years. These AI capability improvements would impact the economy and labor market as organizations adopt AI, which could have a substantially longer timeline.
  •  

CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models

arXiv:2604.01366v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making contexts. While prior work has shown that LLMs exhibit cognitive biases behaviorally, whether these biases correspond to identifiable internal representations and can be mitigated through targeted intervention remains an open question. We define LLM cognitive bias as systematic, reproducible deviations from correct answers in tasks with computable ground-truth baselines, and introduce LLM CogBias, a benchmark organized around four families of cognitive biases: Judgment, Information Processing, Social, and Response. We evaluate three LLMs and find that cognitive biases emerge systematically across all four families, with magnitudes and debiasing responses that are strongly family-dependent: prompt-level debiasing substantially reduces Response biases but backfires for Judgment biases. Using linear probes under a contrastive design, we show that these biases are encoded as linearly separable directions in model activation space. Finally, we apply activation steering to modulate biased behavior, achieving 26--32\% reduction in bias score (fraction of biased responses) while preserving downstream capability on 25 benchmarks (Llama: negligible degradation; Qwen: up to $-$19.0pp for Judgment biases). Despite near-orthogonal bias representations across models (mean cosine similarity 0.01), steering reduces bias at similar rates across architectures ($r(246)$=.621, $p$
  •  

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics

arXiv:2604.01375v1 Announce Type: new Abstract: Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubrics using downstream signals such as reinforcement learning outcomes, there remains no principled way to diagnose rubric quality issues from such aggregated or downstream signals alone. To address this gap, we introduce RIFT: RubrIc Failure mode Taxonomy, a taxonomy for systematically characterizing failure modes in rubric composition and design. RIFT consists of eight failure modes organized into three high-level categories: Reliability Failures, Content Validity Failures, and Consequential Validity Failures. RIFT is developed using grounded theory by iteratively annotating rubrics drawn from five diverse benchmarks spanning general instruction following, code generation, creative writing, and expert-level deep research, until no new failure modes are identified. We evaluate the consistency of the taxonomy by measuring agreement among independent human annotators, observing fair agreement overall (87% pairwise agreement and 0.64 average Cohen's kappa). Finally, to support scalable diagnosis, we propose automated rubric quality metrics and show that they align with human failure-mode annotations, achieving up to 0.86 F1.
  •  

Leveraging the Value of Information in POMDP Planning

arXiv:2604.01434v1 Announce Type: new Abstract: Partially observable Markov decision processes (POMDPs) offer a principled formalism for planning under state and transition uncertainty. Despite advances made towards solving large POMDPs, obtaining performant policies under limited planning time remains a major challenge due to the curse of dimensionality and the curse of history. For many POMDP problems, the value of information (VOI) - the expected performance gain from reasoning about observations - varies over the belief space. We introduce a dynamic programming framework that exploits this structure by conditionally processing observations based on the value of information at each belief. Building on this framework, we propose Value of Information Monte Carlo planning (VOIMCP), a Monte Carlo Tree Search algorithm that allocates computational effort more efficiently by selectively disregarding observation information when the VOI is low, avoiding unnecessary branching of observations. We provide theoretical guarantees on the near-optimality of our VOI reasoning framework and derive non-asymptotic convergence bounds for VOIMCP. Simulation evaluations demonstrate that VOIMCP outperforms baselines on several POMDP benchmarks.
  •  

ClawSafety: "Safe" LLMs, Unsafe Agents

arXiv:2604.01438v1 Announce Type: new Abstract: Personal AI agents like OpenClaw run with elevated privileges on users' local machines, where a single successful prompt injection can leak credentials, redirect financial transactions, or destroy files. This threat goes well beyond conventional text-level jailbreaks, yet existing safety evaluations fall short: most test models in isolated chat settings, rely on synthetic environments, and do not account for how the agent framework itself shapes safety outcomes. We introduce CLAWSAFETY, a benchmark of 120 adversarial test scenarios organized along three dimensions (harm domain, attack vector, and harmful action type) and grounded in realistic, high-privilege professional workspaces spanning software engineering, finance, healthcare, law, and DevOps. Each test case embeds adversarial content in one of three channels the agent encounters during normal work: workspace skill files, emails from trusted senders, and web pages. We evaluate five frontier LLMs as agent backbones, running 2,520 sandboxed trials across all configurations. Attack success rates (ASR) range from 40\% to 75\% across models and vary sharply by injection vector, with skill instructions (highest trust) consistently more dangerous than email or web content. Action-trace analysis reveals that the strongest model maintains hard boundaries against credential forwarding and destructive actions, while weaker models permit both. Cross-scaffold experiments on three agent frameworks further demonstrate that safety is not determined by the backbone model alone but depends on the full deployment stack, calling for safety evaluation that treats model and framework as joint variables.
  •  

When AI Gets it Wong: Reliability and Risk in AI-Assisted Medication Decision Systems

arXiv:2604.01449v1 Announce Type: new Abstract: Artificial intelligence (AI) systems are increasingly integrated into healthcare and pharmacy workflows, supporting tasks such as medication recommendations, dosage determination, and drug interaction detection. While these systems often demonstrate strong performance under standard evaluation metrics, their reliability in real-world decision-making remains insufficiently understood. In high-risk domains such as medication management, even a single incorrect recommendation can result in severe patient harm. This paper examines the reliability of AI-assisted medication systems by focusing on system failures and their potential clinical consequences. Rather than evaluating performance solely through aggregate metrics, this work shifts attention towards how errors occur and what happens when AI systems produce incorrect outputs. Through a series of controlled, simulated scenarios involving drug interactions and dosage decisions, we analyse different types of system failures, including missed interactions, incorrect risk flagging, and inappropriate dosage recommendations. The findings highlight that AI errors in medication-related contexts can lead to adverse drug reactions, ineffective treatment, or delayed care, particularly when systems are used without sufficient human oversight. Furthermore, the paper discusses the risks of over-reliance on AI recommendations and the challenges posed by limited transparency in decision-making processes. This work contributes a reliability-focused perspective on AI evaluation in healthcare, emphasising the importance of understanding failure behavior and real-world impact. It highlights the need to complement traditional performance metrics with risk-aware evaluation approaches, particularly in safety-critical domains such as pharmacy practice.
  •  

A Multi-Agent Human-LLM Collaborative Framework for Closed-Loop Scientific Literature Summarization

arXiv:2604.01452v1 Announce Type: new Abstract: Scientific discovery is slowed by fragmented literature that requires excessive human effort to gather, analyze, and understand. AI tools, including autonomous summarization and question answering, have been developed to aid in understanding scientific literature. However, these tools lack the structured, multi-step approach necessary for extracting deep insights from scientific literature. Large Language Models (LLMs) offer new possibilities for literature analysis, but remain unreliable due to hallucinations and incomplete extraction. We introduce Elhuyar, a multi-agent, human-in-the-loop system that integrates LLMs, structured AI, and human scientists to extract, analyze, and iteratively refine insights from scientific literature. The framework distributes tasks among specialized agents for filtering papers, extracting data, fitting models, and summarizing findings, with human oversight ensuring reliability. The system generates structured reports with extracted data, visualizations, model equations, and text summaries, enabling deeper inquiry through iterative refinement. Deployed in materials science, it analyzed literature on tungsten under helium-ion irradiation, showing experimentally correlated exponential helium bubble growth with irradiation dose and temperature, offering insight for plasma-facing materials (PFMs) in fusion reactors. This demonstrates how AI-assisted literature review can uncover scientific patterns and accelerate discovery.
  •  

Infeasibility Aware Large Language Models for Combinatorial Optimization

arXiv:2604.01455v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored for NP-hard combinatorial optimization problems, but most existing methods emphasize feasible-instance solution generation and do not explicitly address infeasibility detection. We propose an infeasibility-aware framework that combines certifiable dataset construction, supervised fine-tuning, and LLM-assisted downstream search. For the minor-embedding problem, we introduce a new mathematical programming formulation together with provable zero-phase infeasibility screening, which enables scalable construction of training instances labeled either as feasible with structured certificates or as certifiably infeasible. Using training data generated through this exact optimization pipeline, we show that an 8B-parameter LLM can be fine-tuned to jointly perform solution generation and infeasibility detection. We further utilize LLM outputs as warm starts for downstream local search, providing a practical way to accelerate optimization even when the LLM outputs are imperfect. Experiments show that our fine-tuned model improves overall accuracy by up to 30\% over GPT-5.2; meanwhile LLM-guided warm starts provide up to $2\times$ speedup compared with starting from scratch in downstream local search.
  •  
❌