❌

Normal view

Advancing Graph Few-Shot Learning via In-Context Learning

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning. However, existing methods often face two key limitations. First, the predominant graph few-shot learning paradigm relies on supervised tasks, failing to leverage the vast number of unlabeled nodes in the graph. Second, many approaches require complex task adaptation or fine-tuning during inference, limiting their efficiency and applicability. Inspired by the powerful in-context learning capabilities of large language models, we propose a novel model named VISION for adVancIng graph few-Shot learning via In-cOntext LearNing to address these challenges. Our model reframes graph few-shot learning as a fine-tuning-free sequence reasoning problem. At its core is a context-aware network that initializes nodes with role embeddings and employs a dual-context fusion module to synergistically integrate local topological structures and global task-level dependencies. This allows our model to dynamically generate class-aware representations for the query set conditioned on the support set context in a single forward pass. To effectively train our model, we introduce an unsupervised task generator that creates structure-adaptive features and constructs diverse pseudo-tasks from abundant unlabeled data. Our method unifies unsupervised meta-learning with graph in-context learning, achieving efficient inference. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our model. Our public code can be found

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we introduce a novel framework POLARIS that brings the rigor of specification-based software testing to AI safety. POLARIS first compiles unstructured natural-language policies into First-Order Logic (FOL) representations, establishing a traceable link between high-level rules and concrete test cases. This formalization enables the construction of a Semantic Policy Graph, where complex policy violation scenarios are encoded as traversable paths. By systematically exploring this graph, POLARIS uncovers compositional violation patterns, which are then instantiated into executable natural-language test queries, enabling coverage-driven and reproducible safety testing. Experiments demonstrate that POLARIS achieves higher policy coverage and attack success counts compared to established baselines. Crucially, by bridging formal methods and AI safety, POLARIS provides a principled, automated approach to ensuring LLMs adhere to safety-critical policies with verifiable traceability. We release our code at https://github.com/huac-lxy/POLARIS.

ADMFormer: An Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention for Traffic Forecasting

arXiv:2605.25543v1 Announce Type: new Abstract: Accurate traffic forecasting is essential for intelligent transportation systems, supporting a wide range of real-world applications. However, it remains challenging due to two key factors:~(1) Traffic series contain heterogeneous temporal patterns, where stable periodic regularities coexist with event-driven fluctuations. Existing methods often treat them within a unified representation, limiting their ability to capture fine-grained temporal dynamics.~(2)Spatial dependencies among nodes are inherently dynamic and sparse, while dense all-pairs attention often introduces redundant interactions and amplifies noise. To address these issues, we propose ADMFormer, an Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention. Specifically, ADMFormer first employs a time-node adaptive gating mechanism to decouple traffic signals into dominant regularities and residual fluctuations that vary across time and nodes. A dual-branch temporal module is then designed to separately capture global periodic dependencies and high-frequency irregular variations from these two decomposed components. Furthermore, ADMFormer introduces a time-varying masked spatial attention that sparsifies spatial interactions based on real-time traffic states, thereby effectively preserving dynamic and informative dependencies. Extensive experiments on four real-world datasets demonstrate that ADMFormer achieves state-of-the-art performance.

PHGNet: Prototype-Guided Hypergraph Construction for Heterogeneous Spatiotemporal Forecasting

arXiv:2605.25554v1 Announce Type: new Abstract: As a core task in intelligent transportation systems, traffic forecasting plays a critical role in urban traffic management. Accurate traffic forecasting relies on modeling complex spatiotemporal dependencies, which is inherently challenging due to spatial heterogeneity in traffic systems.Despite significant progress, most existing methods are still limited to pairwise spatial dependency modeling, making it difficult to capture dynamic high-order interactions among nodes with similar traffic patterns. To address this issue, we propose PHGNet, a novel spatiotemporal forecasting framework based on prototype-guided hypergraph construction. At the core of PHGNet, a prototype learning mechanism is designed to adaptively assign pattern-similar nodes to hyperedges, thereby capturing high-order interactions with time-varying structures. To improve the reliability of dynamic hypergraph construction, we further develop a global-local node representation module to extract time-consistent features. For forecasting, iterative residual refinement and Temporal Query Attention are introduced to improve forecasting accuracy while supporting efficient parallel decoding. Extensive experiments on multiple real-world datasets demonstrate that PHGNet achieves superior predictive performance compared with state-of-the-art methods.

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

arXiv:2605.24523v1 Announce Type: cross Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. We introduce a tri-modal contrastive framework for EEG-based visual decoding that aligns EEG, visual, and textual representations within a unified latent space. Our approach follows a two-stage design. First, we pre-train an EEG encoder via masked reconstruction on unlabeled trials, learning spatio-temporal regularities that transfer robustly to downstream tasks. Second, we jointly align EEG, image, and LLM-generated textual descriptions through contrastive learning, where text supervision acts as a semantic regularizer that injects linguistic structure into the shared space without overwhelming the primary EEG-image signal. The encoder integrates subject-specific adaptation, graph-attention over channels, and temporal-spatial convolutional embeddings. On the Things-EEG2 200-way zero-shot benchmark, our framework achieves 54.1% Top-1 and 83.4% Top-5 accuracy, substantially exceeding the strongest prior baseline (32.4% / 64.0%), with paired Wilcoxon tests confirming significance (p

What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval

arXiv:2605.24524v1 Announce Type: cross Abstract: In non-invasive neural language decoding, results can be inflated by sources that are not stimulus-evoked neural evidence: decoder priors, embedding-based metrics, and non-neural structural nuisances such as signal duration. The methodological challenge is therefore attribution: a reported gain is more informative when it can be traced to a specific source. We recast stimulus-locked MEG-to-audio retrieval as an auditing framework that separates apparent performance into three sources - structural shortcuts, window-level stimulus-locked evidence, and cross-window contextual aggregation - and provides a diagnostic for each. Signal-blind Gaussian noise reaches 66.3% Rank@1 (R@1) under variable-length decoding but collapses to near chance once fixed-duration windows and stimulus-identity splits are enforced, isolating structural leakage. Under these controls, fixed-window retrieval recovers measurable MEG-audio discriminability, while an oracle sentence-bucket diagnostic shows that 95.7% of Top-1 errors select the wrong sentence, localising the residual bottleneck to sentence-level competition. We audit this contextual source with Group Context Bias (GCB), an inference-time additive logit bias that pools sentence-consistent evidence across windows while leaving the base retrieval scores and candidate pool fixed. Used as a score-space intervention, GCB makes the contextual source measurable: R@1 shifts from 44% to 52% on Gwilliams and from 22% to 29% on MOUS under the same fixed setting. GCB is auditable under this design: its effect collapses under random-grouping perturbations and vanishes when local evidence is attenuated in MEG or is near chance in EEG, supporting its use as a controlled source-attribution intervention. These results suggest that brain-to-language performance should be source-attributed, not merely reported.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction

arXiv:2604.17328v2 Announce Type: replace-cross Abstract: This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially alleviate length-related phenomena, a more fundamental issue remains insufficiently characterized: the comparison units used during training lack inherent comparability. Building on this observation, we propose a new perspective: the length problem should not be viewed merely as a loss-scaling or normalization bias, but rather as a \emph{comparison unit construction} problem. We further establish a sample-construction-based training framework that, instead of applying post-hoc corrections to unequal-length responses, proactively constructs equal-length, alignable, and comparable training segments during generation. Within this framework, we propose EqLen, a concrete method applicable to group-relative comparison algorithms such as GRPO, GSPO, and RLOO. Through dual-track synchronous generation, prefix inheritance, and segment masking, EqLen efficiently collects effective equal-length training segments and enables stable

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamentally in how to transform feedback provided only at the end of a sequence into fine-grained learning signals that can guide intermediate reasoning steps. Existing approaches either rely on outcome-level rewards for sequence-level optimization, which makes precise credit assignment difficult, or depend on externally constructed process supervision, which is costly and difficult to scale sustainably. To address this, we propose a new perspective: reinforcement learning for reasoning can be understood as the problem of internalizing outcome supervision into process supervision. From this perspective, we introduce a supervision-internalization method for reinforcement learning for reasoning, enabling the model to automatically extract process-level learning signals through identifying, correcting, and reusing failed reasoning trajectories, thereby achieving finer-grained policy optimization under outcome-only supervision. We further abstract this idea into a new training paradigm, in which the model continually generates and refines its own internal process supervision during reinforcement learning, opening a new path for fine-grained credit assignment in reinforcement learning for reasoning that differs from externally provided process supervision.

CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks

arXiv:2604.04060v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed in complex applications, their vulnerability to adversarial attacks raises urgent safety concerns, especially those evolving over multi-round interactions. Existing defenses are largely reactive and struggle to adapt as adversaries refine strategies across rounds. In this work, we propose CoopGuard , a stateful multi-round LLM defense framework based on cooperative agents that maintains and updates an internal defense state to counter evolving attacks. It employs three specialized agents (Deferring Agent, Tempting Agent, and Forensic Agent) for complementary round-level strategies, coordinated by System Agent, which conditions decisions on the evolving defense state (interaction history) and orchestrates agents over time. To evaluate evolving threats, we introduce the EMRA benchmark with 5,200 adversarial samples across 8 attack types, simulating progressively LLM multi-round attacks. Experiments show that CoopGuard reduces attack success rate by 78.9% over state-of-the-art defenses, while improving deceptive rate by 186% and reducing attack efficiency by 167.9%, offering a more comprehensive assessment of multi-round defense. These results demonstrate that CoopGuard provides robust protection for LLMs in multi-round adversarial scenarios.

Embedding Enhancement via Fine-Tuned Language Models for Learner-Item Cognitive Modeling

arXiv:2604.04088v1 Announce Type: cross Abstract: Learner-item cognitive modeling plays a central role in the web-based online intelligent education system by enabling cognitive diagnosis (CD) across diverse online educational scenarios. Although ID embedding remains the mainstream approach in cognitive modeling due to its effectiveness and flexibility, recent advances in language models (LMs) have introduced new possibilities for incorporating rich semantic representations to enhance CD performance. This highlights the need for a comprehensive analysis of how LMs enhance embeddings through semantic integration across mainstream CD tasks. This paper identifies two key challenges in fully leveraging LMs in existing work: Misalignment between the training objectives of LMs and CD models creates a distribution gap in feature spaces; A unified framework is essential for integrating textual embeddings across varied CD tasks while preserving the strengths of existing cognitive modeling paradigms to ensure the robustness of embedding enhancement. To address these challenges, this paper introduces EduEmbed, a unified embedding enhancement framework that leverages fine-tuned LMs to enrich learner-item cognitive modeling across diverse CD tasks. EduEmbed operates in two stages. In the first stage, we fine-tune LMs based on role-specific representations and an interaction diagnoser to bridge the semantic gap of CD models. In the second stage, we employ a textual adapter to extract task-relevant semantics and integrate them with existing modeling paradigms to improve generalization. We evaluate the proposed framework on four CD tasks and computerized adaptive testing (CAT) task, achieving robust performance. Further analysis reveals the impact of semantic information across diverse tasks, offering key insights for future research on the application of LMs in CD for online intelligent education systems.

Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality

arXiv:2604.04418v1 Announce Type: cross Abstract: As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether these justifications help users distinguish correct answers from incorrect ones. We formalize this idea as error verifiability and propose $v_{\text{bal}}$, a balanced metric that measures whether justifications enable raters to accurately assess answer correctness, validated against human raters who show high agreement. We find that neither common approaches, such as post-training and model scaling, nor more targeted interventions recommended improve verifiability. We introduce two methods that succeed at improving verifiability: reflect-and-rephrase (RR) for mathematical reasoning and oracle-rephrase (OR) for factual QA, both of which improve verifiability by incorporating domain-appropriate external information. Together, our results establish error verifiability as a distinct dimension of response quality that does not emerge from accuracy improvements alone and requires dedicated, domain-aware methods to address.

Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

arXiv:2603.25158v3 Announce Type: replace Abstract: Equipping Large Language Model (LLM) agents with domain-specific skills is critical for tackling complex tasks. Yet, manual authoring creates a severe scalability bottleneck. Conversely, automated skill generation often yields fragile or fragmented results because it either relies on shallow parametric knowledge or sequentially overfits to non-generalizable trajectory-local lessons. To overcome this, we introduce Trace2Skill, a framework that mirrors how human experts author skills: by holistically analyzing broad execution experience before distilling it into a single, comprehensive guide. Instead of reacting sequentially to individual trajectories, Trace2Skill dispatches a parallel fleet of sub-agents to analyze a diverse pool of executions. It extracts trajectory-specific lessons and hierarchically consolidates them into a unified, conflict-free skill directory via inductive reasoning. Trace2Skill supports both deepening existing human-written skills and creating new ones from scratch. Experiments in challenging domains, such as spreadsheet, VisionQA and math reasoning, show that Trace2Skill significantly improves upon strong baselines, including Anthropic's official xlsx skills. Crucially, this trajectory-grounded evolution does not merely memorize task instances or model-specific quirks: evolved skills transfer across LLM scales and generalize to OOD settings. For example, skills evolved by Qwen3.5-35B on its own trajectories improved a Qwen3.5-122B agent by up to 57.65 absolute percentage points on WikiTableQuestions. Ultimately, our results demonstrate that complex agent experience can be packaged into highly transferable, declarative skills -- requiring no parameter updates, no external retrieval modules, and utilizing open-source models as small as 35B parameters.

Doxorubicin promotes the production of inflammatory cytokines in tumor-associated macrophages through activating lactate dehydrogenase A

Cell Death Discovery, Published online: 31 March 2026; doi:10.1038/s41420-026-03014-0

Doxorubicin promotes the production of inflammatory cytokines in tumor-associated macrophages through activating lactate dehydrogenase A

Multi-omics analysis reveals AR as a potential prognostic factor and immune-related therapeutic target in gastric cancer

Biochem Biophys Rep. 2026 Mar 16;46:102537. doi: 10.1016/j.bbrep.2026.102537. eCollection 2026 Jun.

ABSTRACT

BACKGROUND: Although studies have shown that the androgen receptor (AR) is associated with tumor progression and malignant regulation, its role in the tumor immune microenvironment and predictive value for prognosis and immunotherapy response in various cancer types have not been systematically analyzed.

METHODS: In this paper, multi-omics techniques was used to analyze AR comprehensively.

RESULTS: A comprehensive pan-cancer analysis revealed that the AR was expressed in a variety of tumors, especially as a risk factor for poor prognosis in gastric cancer. In addition, gene set enrichment analysis showed that the AR promotes cell proliferation and tumor cell invasion and regulates anti-tumor response. Immune score, immune cell infiltration, and anticancer immune cycle analysis showed that high AR levels were correlated with low infiltration of CD4+ T cells and NKT cells, high infiltration of Th2 cells and MDSCs, negatively correlated with antigen-presenting molecules, and positively correlated with various immune-negative regulatory molecules. Single-cell sequencing highlighted the heterogeneous expression of ARs in different cell types, particularly in epithelial cells, where high AR levels were associated with the enhanced activity of tumor-promoting pathways.

CONCLUSIONS: In conclusion, this study highlights the potential of the AR as a novel biomarker for gastric cancer prognosis and immunotherapy efficacy, expanding its applicability in the development of new antitumor drugs.

PMID:41890218 | PMC:PMC13014673 | DOI:10.1016/j.bbrep.2026.102537

SVNeoPP: A Workflow for Structural-Variant-Derived Neoantigen Prediction and Prioritization Using Multi-Omics Data

Biology (Basel). 2026 Mar 19;15(6):492. doi: 10.3390/biology15060492.

ABSTRACT

BACKGROUND: Tumor neoantigens are key targets for personalized vaccines and T-cell therapies, yet most pipelines focus on neoantigens derived from SNV/small indel and often yield a limited number of high-quality candidates. SVs are prevalent in tumors and can generate novel chimeric sequences and neopeptides, making them a promising additional source of neoantigens. However, SV-derived neoantigen prediction remains challenging due to breakpoint uncertainty, isoform-dependent coding inference, and limited integration of multi-dimensional evidence and reproducibility.

METHODS: We developed SVNeoPP (Structural Variant Neoantigen Prediction and Prioritization), an end-to-end workflow for SV-derived neoantigen analysis. SVNeoPP takes WGS and RNA-seq as inputs, performs SV calling and annotation, and reconstructs altered transcripts and coding sequences in a traceable, isoform-aware manner to generate candidate peptides. Candidates are prescreened by integrating antigen-processing features with HLA binding prediction, and then hierarchically filtered and prioritized based on transcript expression, LC-MS/MS proteomics evidence, immunogenicity predictions, and sequence similarity to experimentally validated neoantigen databases. SVNeoPP is implemented in Snakemake to enable modular extension, checkpoint-based restarts, and end-to-end reproducibility.

RESULTS: Using a hepatocellular carcinoma (HCC) multi-omics dataset as a proof of concept, we demonstrated the performance of SVNeoPP and obtained a high-priority shortlist of candidate peptides. Compared with other methods, SVNeoPP substantially expanded the candidate search space for SV-derived neoantigens and showed more favorable distributions of antigen-processing and HLA binding features.

CONCLUSIONS: SVNeoPP provides a reusable, traceable, and interpretable multi-dimensional evidence-driven framework for SV-derived neoantigens. As a complementary module to SNV/small-indel pipelines, it broadens the neoantigen candidate repertoire and generates ranked candidates with interpretable evidence to facilitate downstream prioritization and decision-making.

PMID:41892252 | PMC:PMC13024079 | DOI:10.3390/biology15060492

Suicidal Thoughts and Behaviors Among Chinese Adolescents in Relation to Negative Life Events, Internet Addiction, and Sexual Abuse: Cross-Sectional Study

Background: Increasing suicidal thoughts and behaviors (STB) among adolescents raise social concerns and have a well-recognized association with sexual abuse (SA). However, research regarding the mechanisms explaining the association between SA and STB remains limited. Objective: This study aims to examine the chained mediating effects of negative life events (NLE) and internet addiction (IA) between SA and STB among adolescents in China. Methods: This cross-sectional study used data from the Science Database of the People Mental Health survey conducted between March 2013 and December 2022 by the National Population Health Data Center of the National Research Institute for Family Planning. Through stratified sampling, 20,893 adolescents were recruited from 16 Chinese provinces. After excluding samples with missing relevant variables, 10,664 (55.89%; aged 16-17.9 y; n=5826, 54.63% women) adolescents were included in the final analysis. STB was the outcome variable, with NLE and IA as mediators, all assessed via a questionnaire that was uniformly administered by trained investigators in school settings. The Pearson χ test was used to analyze the association between SA and STB. Using a combination of multiple linear regression and bootstrap testing, the study constructed a chain mediation model to explore how SA influences STB in adolescents through NLE and IA. Results: The scores for SA, NLE, IA, and STB were 1.330 (SD 1.714), 51.960 (SD 23.822), 34.88 (SD 13.852), and 0.690 (SD 1.396), respectively. Multiple linear regression analysis indicated SA was associated with NLE (β=2.382, 95% CI 2.112‐2.653;

Spatial Omics in Gastrointestinal Oncology: Recent Advances, Therapeutic Insights, and Clinical Translation

J Cancer. 2026 Jan 30;17(3):515-523. doi: 10.7150/jca.127381. eCollection 2026.

ABSTRACT

Gastrointestinal (GI) cancers remain a leading cause of cancer-related morbidity and mortality worldwide, largely due to their molecular heterogeneity, complex tumor microenvironment (TME), and variable treatment responses. In recent years, the emergence of spatially resolved omics technologies-encompassing spatial transcriptomics, proteomics, metabolomics, and epigenomics-has revolutionized the ability to interrogate tumor architecture with unprecedented resolution. These methods enable precise mapping of cellular and molecular interactions within intact tissue contexts, thereby uncovering spatially defined niches that influence tumor progression, immune evasion, and therapeutic resistance. In GI malignancies such as colorectal, gastric, and esophageal cancers, spatial omics have provided critical insights into cancer-stromal-immune crosstalk, identified predictive biomarkers for immunotherapy and targeted agents, and guided the development of novel therapeutic strategies. This review synthesizes the latest advances in spatial omics applied to GI oncology over the past five years, with an emphasis on their integration into early diagnosis, treatment stratification, and real-time monitoring of therapeutic efficacy. We also discuss current challenges, including standardization, data integration, and clinical validation, as well as future directions for incorporating spatial profiling into routine oncology practice. By bridging the gap between bench discoveries and bedside applications, spatial omics hold transformative potential for achieving truly personalized treatment in gastrointestinal cancers.

PMID:41869445 | PMC:PMC13003551 | DOI:10.7150/jca.127381

Assembly of helper NLR resistosome clusters upon activation of a coiled-coil NLR

Nature, Published online: 11 March 2026; doi:10.1038/s41586-026-10215-1

SUMM2, a coiled-coil NLR, promotes the assembly of higher-order resistosome clusters to initiate cell death in plants.

Task learning increases information redundancy of neural responses in macaque visual cortex

arXiv:2603.07369v1 Announce Type: new Abstract: How does the brain optimize sensory information for decision-making in new tasks? One hypothesis suggests learning reduces redundancy in neural representations to improve efficiency, while another, based on Bayesian inference, predicts learning increases redundancy by distributing information across neurons. We tested these hypotheses by tracking population responses in macaque cortical area V4 as monkeys learned visual discrimination tasks. We found strong support for the Bayesian predictions: task learning increased redundancy in neural responses over weeks of training and within single trials. This redundancy did not reduce information but instead increased the information carried by individual neurons. These insights suggest sensory processing in the brain reflects a generative rather than discriminative inference process.

Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset

arXiv:2406.06512v2 Announce Type: replace-cross Abstract: The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previous state-of-the-art approaches for automated analysis leverage vision-language models (VLMs) that jointly model images and radiology reports. However, current medical VLMs are generally limited to 2D images and short reports. Here to overcome these shortcomings for abdominal CT interpretation, we introduce Merlin, a 3D VLM that learns from volumetric CT scans, electronic health record data and radiology reports. This approach is enabled by a multistage pretraining framework that does not require additional manual annotations. We trained Merlin using a high-quality clinical dataset of paired CT scans (>6 million images from 15,331 CT scans), diagnosis codes (>1.8 million codes) and radiology reports (>6 million tokens). We comprehensively evaluated Merlin on 6 task types and 752 individual tasks that covered diagnostic, prognostic and quality-related tasks. The non-adapted (off-the-shelf) tasks included zero-shot classification of findings (30 findings), phenotype classification (692 phenotypes) and zero-shot cross-modal retrieval (image-to-findings and image-to-impression). The model-adapted tasks included 5-year chronic disease prediction (6 diseases), radiology report generation and 3D semantic segmentation (20 organs). We validated Merlin at scale, with internal testing on 5,137 CT scans and external testing on 44,098 CT scans from 3 independent sites and 2 public datasets. The results demonstrated high generalization across institutions and anatomies. Merlin outperformed 2D VLMs, CT foundation models and off-the-shelf radiology models. We also release our trained models, code, and dataset, available at: https://github.com/StanfordMIMI/Merlin.
❌