❌

Normal view

Artificial Intelligence-Driven Multiomics and Clinical Investigation Identify Macrophage Migration Inhibitory Factor as a Pan-Cancer Biomarker

Phenomics. 2026 May 20;6(3):213-229. doi: 10.1007/s43657-026-00322-4. eCollection 2026 Jun.

ABSTRACT

Early cancer detection remains challenging due to the lack of reliable pan-cancer screening methods, particularly blood-based biomarkers. Using a novel three-tiered validation framework combining artificial intelligence (AI)-powered literature mining of 180,000 PubMed articles (1950-2024), multiomics integration across major databases, and extensive clinical validation, we identified macrophage migration inhibitory factor (MIF) as a promising blood-based biomarker for pan-cancer detection. Multiomics analysis revealed consistent MIF upregulation across 21 cancer types at the transcriptional level and across 12 cancer types at the protein level. Clinical validation in independent cohorts (n = 4,269) showed that serum MIF protein levels discriminated effectively between cancer patients and healthy controls (median AUC = 0.994) and between cancer and benign conditions (median AUC = 0.881). Notably, comparative analyses showed that MIF demonstrated superior or comparable performance to established cancer-specific markers, including AFP for hepatocellular carcinoma (MIF AUC = 0.885 vs. AFP AUC: 0.744-0.887) and CA125 for ovarian cancer (MIF AUC = 0.831 vs. CA125 AUC: 0.58-0.71). Meta-analysis of 28 cohorts (n = 5,347) confirmed the diagnostic efficacy of MIF (pooled AUC: 0.782). This cost-effective, blood-based ELISA approach establishes MIF as a valuable tool for broad applications in cancer screening.

SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s43657-026-00322-4.

PMID:42750739 | PMC:PMC13578188 | DOI:10.1007/s43657-026-00322-4

Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling

arXiv:2609.09691v1 Announce Type: cross Abstract: When training data are limited, increasing parameter count is not the only way to improve language-model performance. A small parameter set, when repeatedly applied, can also deliver comparable performance. We study Looped GPT-BERT in the BabyLM 2026 Strict-small setting, combining GPT-BERT's masked next-token and causal language-modeling objectives with depth-wise parameter sharing. We train on a preprocessed 7.48M-word English corpus and compare objective ratios, non-looped and looped architectures, and loop counts. Our final $4\times12$ model uses four physical layers for twelve recurrent traversals and contains 12.18M parameters. The BabyLM 2026 leaderboard reports an Overall Average of 35.42 and an NLP Average of 48.48. Compared with public BabyLM 10M Strict-small GPT-2 and GPT-BERT baselines, it achieves comparable performance on selected linguistic and downstream metrics, including BLiMP and GLUE, with fewer parameters. The loop ablations show that additional recurrent computation can improve training and preserve strong performance on selected linguistic tasks, whereas poorer performance on other tasks may reveal an inherent limitation of the looped design: using only a few physical layers restricts the model's representational space.

Activation of methionine metabolism mediated by HNF4α confers ferroptosis resistance in hepatocellular carcinoma

Cell Death Discovery, Published online: 26 May 2026; doi:10.1038/s41420-026-03165-0

Activation of methionine metabolism mediated by HNF4α confers ferroptosis resistance in hepatocellular carcinoma

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

arXiv:2603.29139v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled agentic systems that translate natural language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We also conduct a validity study with 12 SciVis experts to examine the agreement between human and LLM judges. Using this framework, we evaluate representative SciVis agents and general-purpose coding agents to establish initial baselines and reveal capability gaps. SciVisAgentBench is designed as a living benchmark to support systematic comparison, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is available at https://scivisagentbench.github.io/.

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

arXiv:2512.00818v2 Announce Type: replace Abstract: MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, splitting medical multimodal reasoning into fine-grained visual understanding and multi-step reasoning to enable targeted evaluation; 2) Challenging task design, with visual understanding across three key dimensions (small-object detection, fine-detail discrimination, spatial understanding) and reasoning covering four clinically relevant scenarios (temporal prediction, causal reasoning, long-tail generalization, multi-source integration); 3) Broad, high-quality data coverage, comprising 20,653 Visual Question Answering (VQA) pairs spanning 11 organ systems and 12 imaging modalities, validated via a rigorous two-stage (human expert + model-assisted) review to ensure clinical authenticity. We evaluate 18 state-of-the-art MLLMs with Med-CMR, revealing GPT-5 as the top-performing commercial model: 57.81 accuracy on multiple-choice questions (MCQs) and a 48.70 open-ended score, outperforming Gemini 2.5 Pro (49.87 MCQ accuracy, 45.98 open-ended score) and leading open-source model Qwen3-VL-235B-A22B (49.34 MCQ accuracy, 42.62 open-ended score). However, specialized medical MLLMs do not reliably outperform strong general models, and long-tail generalization emerges as the dominant failure mode. Med-CMR thus provides a stress test for visual-reasoning integration and rare-case robustness in medical MLLMs, and a rigorous yardstick for future clinical systems.
❌