❌

Normal view

Liquid biopsies using circulating tumor DNA for surveillance of gastrointestinal cancers in Hispanics: first real-world data report

ESMO Real World Data Digit Oncol. 2026 Jan 14;11:100652. doi: 10.1016/j.esmorw.2025.100652. eCollection 2026 Mar.

ABSTRACT

BACKGROUND: Malignant tumors release circulating tumor DNA (ctDNA) into the bloodstream, providing insights into tumor-specific mutations and pathways driving cancer progression. ctDNA testing is currently approved as a type of liquid biopsy to monitor disease burden and detect minimal residual disease (MRD). This study aimed to evaluate the adoption of ctDNA testing in a community oncology practice and assess the overall diagnostic performance of ctDNA and its association with disease progression in stage IV colorectal cancer (CRC), as determined by imaging studies.

PATIENTS AND METHODS: This retrospective study analyzed the medical records of 88 patients with gastrointestinal cancers (80 CRC, 5 gastric, 3 esophageal) who underwent ctDNA molecular testing between January 2020 and April 2022. Electronic medical records from patients aged ≥21 years who had two or more ctDNA tests with concurrent imaging studies or a pathology-confirmed CRC, gastric cancer, or esophageal cancer diagnosis were evaluated.

RESULTS: At baseline, 47 (53.4%) patients had negative and 41 (46.6%) had positive results. Most patients had CRC (90.1%). In stage IV CRC, ctDNA was increasing before radiologic progression in all documented cases (100%), with a median lead time of 2.5 months (range 0.5-15 months). In early-stage CRC (I-III), ctDNA preceded radiologic progression in 40% of cases, with a median lead time of 6 months (range 6-10 months).

CONCLUSIONS: Using real-world data, we report the first-time results of the ctDNA testing adoption in a community oncology setting among patients with gastrointestinal cancers, predominantly CRC. Our findings suggest that integration of ctDNA testing may support disease monitoring in routine clinical practice.

PMID:41930304 | PMC:PMC13040887 | DOI:10.1016/j.esmorw.2025.100652

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

Advances in Metabolic Reprogramming and Immune Regulatory Mechanisms in Lung Cancer

Oncol Res. 2026 Mar 23;34(4):11. doi: 10.32604/or.2026.076176. eCollection 2026.

ABSTRACT

Lung cancer remains the leading cause of cancer-related mortality worldwide, primarily driven by metabolic reprogramming and immune evasion mechanisms within tumor cells. To adapt to the nutrient-deprived tumor microenvironment (TME), lung cancer cells undergo profound metabolic reprogramming, characterized by enhanced glycolysis (the Warburg effect), increased glutamine dependency (mediated by GLS1), and accelerated lipid synthesis (involving enzymes such as FASN). These metabolic alterations not only remodel the TME but also dampen antitumor immune responses by promoting immunosuppressive cell populations (e.g., Tregs and M2 macrophages) and inhibiting effector functions of CD8+ T cells and natural killer (NK) cells. Critically, a bidirectional crosstalk operates between tumor cell metabolism and the immunosuppressive TME: metabolic reprogramming drives immune suppression through metabolite accumulation, whereas the immunosuppressive TME, in turn, promotes tumor cell adaptability-thus forming a positive feedback loop that reinforces immune evasion and therapy resistance. This review elucidates key molecular pathways governing metabolic reprogramming in lung cancer-spanning glucose, amino acid, and lipid metabolism-and their dynamic crosstalk with immune regulation, including epigenetic modifications and non-coding RNA-mediated mechanisms. Additionally, it evaluates emerging therapeutic strategies targeting the metabolic-immune axis, such as inhibitors of HK2 or GLS1 combined with anti-PD-1/PD-L1 agents, which aim to reverse immunosuppression and improve clinical outcomes. By synthesizing recent advances, this work provides a theoretical framework for precision oncology interventions, highlighting the potential of metabolic immunotherapies and future directions integrating AI and multi-omics data to overcome resistance in lung cancer.

PMID:41930159 | PMC:PMC13040304 | DOI:10.32604/or.2026.076176

Interpretable Machine Learning to Understand Wildfire Toxicity: Bridging Chemicals, Omics, and Toxicological Outcomes via Symbolic Regression with Novel Feature Scoring

Chem Res Toxicol. 2026 Apr 3. doi: 10.1021/acs.chemrestox.5c00440. Online ahead of print.

ABSTRACT

Wildfire smoke exposures are increasingly common, consisting of complex mixtures of gases and particulates known to cause diverse pulmonary health effects. While health outcomes are regularly studied, quantitative links between smoke chemical composition and toxicological outcomes remain poorly defined, limiting interpretation of wildfire smoke health risks. This study explores symbolic regression (SR) as an interpretable artificial intelligence/machine learning method to generate closed-form mathematical models linking chemical exposure to biological responses relevant to wildfire smoke. Prior to application on wildfire-relevant data sets, we benchmarked three Python-based SR packages on simulated data, assessing performance across varying noise levels and operator complexities. Insights from these simulation tests, such as the importance of including necessary operators, were incorporated when applying SR to lab-generated wildland fire exposure-toxicity data. This data set included chemical characterizations of biomass smoke exposures and corresponding pulmonary responses in female CD-1 mice (n = 60). Specifically, we evaluated the ability to predict a lung injury marker using (1) targeted measures of over 80 chemicals measured in smoke (RMSE = 17.57 mg/mL) and (2) lung tissue measures of hundreds of transcripts (RMSE = 15.12 mg/mL). Resulting error metrics were comparable to Random Forest and XGBoost models. To aid model interpretation, we developed directional ensemble contribution scores (DECS), a novel feature importance scoring method that quantifies the direction and magnitude of predictor contributions across top-performing models. Expert toxicologists also contributed to model prioritization, integrating a "biologists-in-the-loop" approach. Results highlighted polycyclic aromatic hydrocarbons as drivers of lung injury and methoxyphenols as suppressors. Transcriptomic analyses highlighted a small set of genes, which have roles in metabolism, cell proliferation, immune regulation, and oncogenic processes, with MYC proto-oncogene (Myc) showing the strongest association. Overall, this study demonstrates SR and associated DECS as practical, interpretable tools for modeling environmental mixtures, such as wildfire smoke, and their toxicological effects.

PMID:41928614 | DOI:10.1021/acs.chemrestox.5c00440

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Integrating liquid biopsies and artificial intelligence for early cancer detection: A systematic review and meta-analysis

Eur J Cancer. 2026 Mar 24;239:116699. doi: 10.1016/j.ejca.2026.116699. Online ahead of print.

ABSTRACT

INTRODUCTION: The latest generation of liquid biopsies incorporates multi-omic features, including genomics, methylomics, and fragmentomics. Machine learning (ML) approaches have been proposed to synthesize these complex biological data for the development of diagnostic classifiers. This study aims to evaluate the integration of ML with circulating cell-free DNA (cfDNA) analysis for early cancer detection.

METHODS: Medline, Embase, Cochrane, and Web of Science were searched in July 2025. Eligible studies combined ML and cfDNA features to distinguish cancer patients (stages I-III) from non-cancer controls. Summary diagnostic performance metrics and their 95% confidence intervals (CI) were calculated.

RESULTS: The study included 109 articles permitting analyses for lung (n = 34), liver (n = 29), colorectal (n = 28), pancreatic (n = 16), breast (n = 17), esophageal (n = 12), ovarian (n = 13), gastric (n = 9), head and neck (n = 4), and mixed (n = 27) cancer types. Specificity was consistently high across all tumor types and stages (94%-99%). Sensitivity ranged from 72% to 92% for stage I-III, 44-91% for stage I, 71-98% for stage II and 83-99% for stage III. In the pooled study population, neural networks (90%, 95% CI: 81%-95%), random forest (86%, 95% CI: 77%-92%) and heterogeneous ensemble learning (85%, 95% CI: 79%-89%) demonstrated the highest sensitivity. The stratified analysis by classifier feature revealed 86% (95% CI: 80%-90%) sensitivity for fragmentation and 81% (95% CI: 76%-85%) for methylation, with 92%-96% specificity.

CONCLUSION: ML and cfDNA profiling show potential for early cancer detection, with ensemble methods, neural networks and random forests achieving the best overall performance. Fragmentomic features provide the highest sensitivity.

PMID:41930854 | DOI:10.1016/j.ejca.2026.116699

Immune endotypes in tuberculosis: Keys to decoding disease complexity

J Intern Med. 2026 Apr 3. doi: 10.1111/joim.70092. Online ahead of print.

ABSTRACT

Tuberculosis (TB) remains a major global health challenge, with multi-drug antibiotic regimens as the current standard of care. While effective at killing Mycobacterium tuberculosis, these treatments do not resolve persistent inflammation, prevent lung damage, or reverse immune dysregulation that contribute to poor outcomes and disease recurrence. Precision medicine offers a promising alternative but requires deeper insight into disease mechanisms to enable tailored interventions. This comprehensive review introduces the concept of immune endotyping to define the underlying disease mechanisms as tools to decode clinical and immunological heterogeneity in TB. TB displays a wide spectrum of clinical phenotypes, from latent or asymptomatic infection to mild or severe disease with characteristic non-cavitary or cavitary lung pathology. Instead, distinct immune endotypes capture the diverse biological pathways that shape disease progression and treatment response. Similar clinical presentations may arise from different immune dysfunctions, underscoring the need to move beyond broad phenotypic classifications. Advances in multi-omics and computational analyses uncover immune signatures that enable stratification for host-directed therapies (HDTs) targeting hyperinflammation, immunosuppression, coagulopathy or metabolic exhaustion. Integrating clinical, radiological, and immunological data through multimodal profiling is essential for developing personalized interventions. We also explore how endotyping has transformed treatment in other diseases, offering valuable insights for TB. Additionally, we present examples of how putative immune endotypes may be targeted with appropriate HDTs. In summary, this review underscores the potential of immune endotypes to advance precision medicine in TB, moving beyond one-size-fits-all treatment to improve outcomes, especially in severe and drug-resistant cases.

PMID:41930636 | DOI:10.1111/joim.70092

Advances in Metabolic Reprogramming and Immune Regulatory Mechanisms in Lung Cancer

3 April 2026 at 18:00

Oncol Res. 2026 Mar 23;34(4):11. doi: 10.32604/or.2026.076176. eCollection 2026.

ABSTRACT

Lung cancer remains the leading cause of cancer-related mortality worldwide, primarily driven by metabolic reprogramming and immune evasion mechanisms within tumor cells. To adapt to the nutrient-deprived tumor microenvironment (TME), lung cancer cells undergo profound metabolic reprogramming, characterized by enhanced glycolysis (the Warburg effect), increased glutamine dependency (mediated by GLS1), and accelerated lipid synthesis (involving enzymes such as FASN). These metabolic alterations not only remodel the TME but also dampen antitumor immune responses by promoting immunosuppressive cell populations (e.g., Tregs and M2 macrophages) and inhibiting effector functions of CD8+ T cells and natural killer (NK) cells. Critically, a bidirectional crosstalk operates between tumor cell metabolism and the immunosuppressive TME: metabolic reprogramming drives immune suppression through metabolite accumulation, whereas the immunosuppressive TME, in turn, promotes tumor cell adaptability-thus forming a positive feedback loop that reinforces immune evasion and therapy resistance. This review elucidates key molecular pathways governing metabolic reprogramming in lung cancer-spanning glucose, amino acid, and lipid metabolism-and their dynamic crosstalk with immune regulation, including epigenetic modifications and non-coding RNA-mediated mechanisms. Additionally, it evaluates emerging therapeutic strategies targeting the metabolic-immune axis, such as inhibitors of HK2 or GLS1 combined with anti-PD-1/PD-L1 agents, which aim to reverse immunosuppression and improve clinical outcomes. By synthesizing recent advances, this work provides a theoretical framework for precision oncology interventions, highlighting the potential of metabolic immunotherapies and future directions integrating AI and multi-omics data to overcome resistance in lung cancer.

PMID:41930159 | PMC:PMC13040304 | DOI:10.32604/or.2026.076176

Presentation: Panel: Taking Architecture Out of the Echo Chamber

Andrew Harmel-Law and a panel of expert architects discuss the shifting practice of architecture in 2025. They explain strategies for communicating technical debt to stakeholders, the benefits of decentralized decision-making through ADRs, and the career paths of modern leaders. The panel shares insights on bridging the gap between mobile and backend teams to ensure a holistic system.

By Andrew Harmel-Law, Cat Morris, Diana Montalion, Shana Dacres-Lawrence, Vanessa Formicola, Elena Stojmilova, Peter Hunter

Runtime Burden Allocation for Structured LLM Routing in Agentic Expert Systems: A Full-Factorial Cross-Backend Methodology

arXiv:2604.01235v1 Announce Type: new Abstract: Structured LLM routing is often treated as a prompt-engineering problem. We argue that it is, more fundamentally, a systems-level burden-allocation problem. As large language models (LLMs) become core control components in agentic AI systems, reliable structured routing must balance correctness, latency, and implementation cost under real deployment constraints. We show that this balance is shaped not only by prompts or schemas, but also by how structural work is allocated across the generation stack: whether output structure is emitted directly by the model, compressed during transport, or reconstructed locally after generation. We evaluate this formulation through a comprehensive full-factorial benchmark covering 48 deployment configurations and 15,552 requests across OpenAI, Gemini, and Llama backends. Our central finding is consequential: there is no universal best routing mode. Instead, backend-specific interaction effects dominate performance. Modes that remain highly reliable on Gemini and OpenAI can suffer substantial correctness degradation on Llama, while efficiency gains from compressed realization are strongly backend-dependent. Rather than presenting another isolated model comparison, this work contributes a deployable framework for reasoning about structured routing under heterogeneous backend conditions. We provide a cross-backend evaluation methodology and practical deployment guidance for navigating the correctness-cost-latency frontier in production-grade agentic expert systems.

IDEA2: Expert-in-the-loop competency question elicitation for collaborative ontology engineering

arXiv:2604.01344v1 Announce Type: new Abstract: Competency question (CQ) elicitation represents a critical but resource-intensive bottleneck in ontology engineering. This foundational phase is often hampered by the communication gap between domain experts, who possess the necessary knowledge, and ontology engineers, who formalise it. This paper introduces IDEA2, a novel, semi-automated workflow that integrates Large Language Models (LLMs) within a collaborative, expert-in-the-loop process to address this challenge. The methodology is characterised by a core iterative loop: an initial LLM-based extraction of CQs from requirement documents, a co-creational review and feedback phase by domain experts on an accessible collaborative platform, and an iterative, feedback-driven reformulation of rejected CQs by an LLM until consensus is achieved. To ensure transparency and reproducibility, the entire lifecycle of each CQ is tracked using a provenance model that captures the full lineage of edits, anonymised feedback, and generation parameters. The workflow was validated in 2 real-world scenarios (scientific data, cultural heritage), demonstrating that IDEA2 can accelerate the requirements engineering process, improve the acceptance and relevance of the resulting CQs, and exhibit high usability and effectiveness among domain experts. We release all code and experiments at https://github.com/KE-UniLiv/IDEA2
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Semantic Modeling for World-Centered Architectures Andrei Mantsivoda · Darya Gavrilina
    arXiv:2604.01359v1 Announce Type: new Abstract: We introduce world-centered multi-agent systems (WMAS) as an alternative to traditional agent-centered architectures, arguing that structured domains such as enterprises and institutional systems require a shared, explicit world representation to ensure semantic consistency, explainability, and long-term stability. We classify worlds along dimensions including ontological explicitness, normativity, etc. In WMAS, learning and coordination operate o
     

Semantic Modeling for World-Centered Architectures

arXiv:2604.01359v1 Announce Type: new Abstract: We introduce world-centered multi-agent systems (WMAS) as an alternative to traditional agent-centered architectures, arguing that structured domains such as enterprises and institutional systems require a shared, explicit world representation to ensure semantic consistency, explainability, and long-term stability. We classify worlds along dimensions including ontological explicitness, normativity, etc. In WMAS, learning and coordination operate over a shared world model rather than isolated agent-local representations, enabling global consistency and verifiable system behavior. We propose semantic models as a mathematical formalism for representing such worlds. Finally, we present the Ontobox platform as a realization of WMAS.

Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

arXiv:2604.01363v1 Announce Type: new Abstract: We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and (ii) rising tides where the increase in AI capabilities is more continuous and broad-based. We test for these effects in preliminary evidence from an ongoing evaluation of AI capabilities across over 3,000 broad-based tasks derived from the U.S. Department of Labor O*NET categorization that are text-based and thus LLM-addressable. Based on more than 17,000 evaluations by workers from these jobs, we find little evidence of crashing waves (in contrast to recent work by METR), but substantial evidence that rising tides are the primary form of AI automation. AI performance is high and improving rapidly across a wide range of tasks. We estimate that, in 2024-Q2, AI models successfully complete tasks that take humans approximately 3-4 hours with about a 50% success rate, increasing to about 65% by 2025-Q3. If recent trends in AI capability growth persist, this pace of AI improvement implies that LLMs will be able to complete most text-related tasks with success rates of, on average, 80%-95% by 2029 at a minimally sufficient quality level. Achieving near-perfect success rates at this quality level or comparable success rates at superior quality would require several additional years. These AI capability improvements would impact the economy and labor market as organizations adopt AI, which could have a substantially longer timeline.

When AI Gets it Wong: Reliability and Risk in AI-Assisted Medication Decision Systems

arXiv:2604.01449v1 Announce Type: new Abstract: Artificial intelligence (AI) systems are increasingly integrated into healthcare and pharmacy workflows, supporting tasks such as medication recommendations, dosage determination, and drug interaction detection. While these systems often demonstrate strong performance under standard evaluation metrics, their reliability in real-world decision-making remains insufficiently understood. In high-risk domains such as medication management, even a single incorrect recommendation can result in severe patient harm. This paper examines the reliability of AI-assisted medication systems by focusing on system failures and their potential clinical consequences. Rather than evaluating performance solely through aggregate metrics, this work shifts attention towards how errors occur and what happens when AI systems produce incorrect outputs. Through a series of controlled, simulated scenarios involving drug interactions and dosage decisions, we analyse different types of system failures, including missed interactions, incorrect risk flagging, and inappropriate dosage recommendations. The findings highlight that AI errors in medication-related contexts can lead to adverse drug reactions, ineffective treatment, or delayed care, particularly when systems are used without sufficient human oversight. Furthermore, the paper discusses the risks of over-reliance on AI recommendations and the challenges posed by limited transparency in decision-making processes. This work contributes a reliability-focused perspective on AI evaluation in healthcare, emphasising the importance of understanding failure behavior and real-world impact. It highlights the need to complement traditional performance metrics with risk-aware evaluation approaches, particularly in safety-critical domains such as pharmacy practice.

Infeasibility Aware Large Language Models for Combinatorial Optimization

arXiv:2604.01455v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored for NP-hard combinatorial optimization problems, but most existing methods emphasize feasible-instance solution generation and do not explicitly address infeasibility detection. We propose an infeasibility-aware framework that combines certifiable dataset construction, supervised fine-tuning, and LLM-assisted downstream search. For the minor-embedding problem, we introduce a new mathematical programming formulation together with provable zero-phase infeasibility screening, which enables scalable construction of training instances labeled either as feasible with structured certificates or as certifiably infeasible. Using training data generated through this exact optimization pipeline, we show that an 8B-parameter LLM can be fine-tuned to jointly perform solution generation and infeasibility detection. We further utilize LLM outputs as warm starts for downstream local search, providing a practical way to accelerate optimization even when the LLM outputs are imperfect. Experiments show that our fine-tuned model improves overall accuracy by up to 30\% over GPT-5.2; meanwhile LLM-guided warm starts provide up to $2\times$ speedup compared with starting from scratch in downstream local search.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease

arXiv:2604.01475v1 Announce Type: new Abstract: Parkinsons disease (PD) alters cortical neural dynamics, yet reliable non-invasive electrophysiological biomarkers remain elusive. This study examined whether interpretable EEG features capturing complementary aspects of neural dynamics can discriminate Parkinsonian neural states. A comprehensive set of interpretable features was extracted and grouped into Standard descriptors (spectral power, phase synchronization, time-domain statistics) and Dynamical descriptors (aperiodic activity, cross-frequency coupling, scale-free dynamics, neuronal avalanche statistics, and instantaneous frequency measures). A multi-head attention transformer classifier was trained using strict LOSO validation. Group-level comparisons were performed to identify electrophysiological differences associated with disease and medication state. Standard feature sets achieved strongest performance in discriminating medication states (PDoff vs PDon), whereas Dynamical performed competitively in contrasts between PD patients and healthy controls. Random feature ablation analyses indicated that Dynamical descriptors provide complementary information distributed across features while correlation analysis revealed low redundancy within both feature sets. Group-level comparisons revealed medication-sensitive reductions in delta power and voltage variance, modulation of neuronal avalanche statistics, persistent increases in theta phase synchronization in PD patients, and disease-related alterations in cross-frequency interactions. Traditional spectral and synchronization features primarily reflect medication-related neural modulation, whereas dynamical descriptors reveal broader alterations in cortical network organization associated with disease but also with medication. These findings support multivariate EEG representations as a promising framework for developing non-invasive biomarkers of PD.

A Self-Evolving Agentic Framework for Metasurface Inverse Design

arXiv:2604.01480v1 Announce Type: new Abstract: Metasurface inverse design has become central to realizing complex optical functionality, yet translating target responses into executable, solver-compatible workflows still demands specialized expertise in computational electromagnetics and solver-specific software engineering. Recent large language models (LLMs) offer a complementary route to reducing this workflow-construction burden, but existing language-driven systems remain largely session-bounded and do not preserve reusable workflow knowledge across inverse-design tasks. We present an agentic framework for metasurface inverse design that addresses this limitation through context-level skill evolution. The framework couples a coding agent, evolving skill artifacts, and a deterministic evaluator grounded in physical simulation so that solver-specific strategies can be iteratively refined across tasks without modifying model weights or the underlying physics solver. We evaluate the framework on a benchmark spanning multiple metasurface inverse-design task types, with separate training-aligned and held-out task families. Evolved skills raise in-distribution task success from 38% to 74%, increase criteria pass fraction from 0.510 to 0.870, and reduce average attempts from 4.10 to 2.30. On held-out task families, binary success changes only marginally, but improvements in best margin together with shifts in error composition and agent behavior indicate partial transfer of workflow knowledge. These results suggest that the main value of skill evolution lies in accumulating reusable solver-specific expertise around reliable computational engines, thereby offering a practical path toward more autonomous and accessible metasurface inverse-design workflows.
❌