❌

Normal view

Epigenome-wide Mendelian randomization with multi-omics validation identifies epigenetic drivers of idiopathic pulmonary fibrosis

Commun Biol. 2026 Apr 11. doi: 10.1038/s42003-026-10033-1. Online ahead of print.

ABSTRACT

Idiopathic pulmonary fibrosis (IPF) is a complex disease without clear etiology or effective therapy. While DNA methylation has been implicated in IPF pathogenesis, the tissue-specific causal effects of the epigenetic factors on IPF remain undetermined. Here, we perform epigenome-wide Mendelian randomization using blood-based methylation quantitative trait loci of 420,509 CpG sites and genome-wide association study for IPF to elucidate the causal effects of the CpG sites on IPF. Totally, 452 CpG sites has shown putative causal effects on IPF risk after Bonferroni correction. Among them, 13 CpG sites have shown strong colocalization evidence with genetic factors associated with IPF. Specifically, DNA methylation at CpG sites within MAN2A2 and TRIM27 shows significant differences between IPF lungs and controls, correlating with altered mRNA expressions of these genes in lung tissues. The CpG site in MAN2A2 is a binding site of ZNF384 according to transcription factor databases. RNA sequencing in the TGFβ1-induced alveolar epithelia confirms significantly reduced expression of MAN2A2 and ZNF384 comparing to the controls. Collectively, our study suggests a putative causal link between DNA methylation within MAN2A2 and IPF risk, wherein lung-specific DNA methylation in MAN2A2 may perturb the interaction between ZNF384 and MAN2A2, revealing novel roles for these genes in IPF pathogenesis.

PMID:41965819 | DOI:10.1038/s42003-026-10033-1

Epigenome-wide Mendelian randomization with multi-omics validation identifies epigenetic drivers of idiopathic pulmonary fibrosis

Commun Biol. 2026 Apr 11. doi: 10.1038/s42003-026-10033-1. Online ahead of print.

ABSTRACT

Idiopathic pulmonary fibrosis (IPF) is a complex disease without clear etiology or effective therapy. While DNA methylation has been implicated in IPF pathogenesis, the tissue-specific causal effects of the epigenetic factors on IPF remain undetermined. Here, we perform epigenome-wide Mendelian randomization using blood-based methylation quantitative trait loci of 420,509 CpG sites and genome-wide association study for IPF to elucidate the causal effects of the CpG sites on IPF. Totally, 452 CpG sites has shown putative causal effects on IPF risk after Bonferroni correction. Among them, 13 CpG sites have shown strong colocalization evidence with genetic factors associated with IPF. Specifically, DNA methylation at CpG sites within MAN2A2 and TRIM27 shows significant differences between IPF lungs and controls, correlating with altered mRNA expressions of these genes in lung tissues. The CpG site in MAN2A2 is a binding site of ZNF384 according to transcription factor databases. RNA sequencing in the TGFβ1-induced alveolar epithelia confirms significantly reduced expression of MAN2A2 and ZNF384 comparing to the controls. Collectively, our study suggests a putative causal link between DNA methylation within MAN2A2 and IPF risk, wherein lung-specific DNA methylation in MAN2A2 may perturb the interaction between ZNF384 and MAN2A2, revealing novel roles for these genes in IPF pathogenesis.

PMID:41965819 | DOI:10.1038/s42003-026-10033-1

RAE-AR: Taming Autoregressive Models with Representation Autoencoders

arXiv:2604.01545v1 Announce Type: new Abstract: The latent space of generative modeling is long dominated by the VAE encoder. The latents from the pretrained representation encoders (e.g., DINO, SigLIP, MAE) are previously considered inappropriate for generative modeling. Recently, RAE method lights the hope and reveals that the representation autoencoder can also achieve competitive performance as the VAE encoder. However, the integration of representation autoencoder into continuous autoregressive (AR) models, remains largely unexplored. In this work, we investigate the challenges of employing high-dimensional representation autoencoders within the AR paradigm, denoted as \textit{RAE-AR}. We focus on the unique properties of AR models and identify two primary hurdles: complex token-wise distribution modeling and the high-dimensionality amplified training-inference gap (exposure bias). To address these, we introduce token simplification via distribution normalization to ease modeling difficulty and improve convergence. Furthermore, we enhance prediction robustness by incorporating Gaussian noise injection during training to mitigate exposure bias. Our empirical results demonstrate that these modifications substantially bridge the performance gap, enabling representation autoencoder to achieve results comparable to traditional VAEs on AR models. This work paves the way for a more unified architecture across visual understanding and generative modeling.

AeroTherm-GPT: A Verification-Centered LLM Framework for Thermal Protection System Engineering Workflows

arXiv:2604.01738v1 Announce Type: new Abstract: Integrating Large Language Models (LLMs) into hypersonic thermal protection system (TPS) design is bottlenecked by cascading constraint violations when generating executable simulation artifacts. General-purpose LLMs, treating generation as single-pass text completion, fail to satisfy the sequential, multi-gate constraints inherent in safety-critical engineering workflows. To address this, we propose AeroTherm-GPT, the first TPS-specialized LLM Agent, instantiated through a Constraint-Closed-Loop Generation (CCLG) framework. CCLG organizes TPS artifact generation as an iterative workflow comprising generation, validation, CDG-guided repair, execution, and audit. The Constraint Dependency Graph (CDG) encodes empirical co-resolution structure among constraint categories, directing repair toward upstream fault candidates based on lifecycle ordering priors and empirical co-resolution probabilities. This upstream-priority mechanism resolves multiple downstream violations per action, achieving a Root-Cause Fix Efficiency of 4.16 versus 1.76 for flat-checklist repair. Evaluated on HyTPS-Bench and validated against external benchmarks, AeroTherm-GPT achieves 88.7% End-to-End Success Rate (95% CI: 87.5-89.9), a gain of +12.5 pp over the matched non-CDG ablation baseline, without catastrophic forgetting on scientific reasoning and code generation tasks.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

arXiv:2604.02029v1 Announce Type: new Abstract: Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an increasing body of work shows that many critical internal processes are more naturally carried out in continuous latent space than in human-readable verbal traces. This shift is driven by the structural limitations of explicit-space computation, including linguistic redundancy, discretization bottlenecks, sequential inefficiency, and semantic loss. This survey aims to provide a unified and up-to-date landscape of latent space in language-based models. We organize the survey into five sequential perspectives: Foundation, Evolution, Mechanism, Ability, and Outlook. We begin by delineating the scope of latent space, distinguishing it from explicit or verbal space and from the latent spaces commonly studied in generative visual models. We then trace the field's evolution from early exploratory efforts to the current large-scale expansion. To organize the technical landscape, we examine existing work through the complementary lenses of mechanism and ability. From the perspective of Mechanism, we identify four major lines of development: Architecture, Representation, Computation, and Optimization. From the perspective of Ability, we show how latent space supports a broad capability spectrum spanning Reasoning, Planning, Modeling, Perception, Memory, Collaboration, and Embodiment. Beyond consolidation, we discuss the key open challenges, and outline promising directions for future research. We hope this survey serves not only as a reference for existing work, but also as a foundation for understanding latent space as a general computational and systems paradigm for next-generation intelligence.

MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models

arXiv:2509.23725v3 Announce Type: replace Abstract: Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-grained logical inconsistencies. To address this, we propose \textsc{MedLA}, a logic-driven multi-agent framework built on large language models. Each agent organizes its reasoning process into an explicit logical tree based on syllogistic triads (major premise, minor premise, and conclusion), enabling transparent inference and premise-level alignment. Agents engage in a multi-round, graph-guided discussion to compare and iteratively refine their logic trees, achieving consensus through error correction and contradiction resolution. We demonstrate that \textsc{MedLA} consistently outperforms both static role-based systems and single-agent baselines on challenging benchmarks such as MedDDx and standard medical QA tasks. Furthermore, \textsc{MedLA} scales effectively across both open-source and commercial LLM backbones, achieving state-of-the-art performance and offering a generalizable paradigm for trustworthy medical reasoning.
❌