❌

Reading view

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.
  •  

From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning

arXiv:2510.24528v1 Announce Type: new Abstract: The capability of in-context learning (ICL) enables large language models (LLMs) to perform novel tasks without parameter updates by conditioning on a few input-output examples. However, collecting high-quality examples for new or challenging tasks can be costly and labor-intensive. In this work, we propose a cost-efficient two-stage pipeline that reduces reliance on LLMs for data labeling. Our approach first leverages readily available cross-task examples to prompt an LLM and pseudo-label a small set of target task instances. We then introduce a graph-based label propagation method that spreads label information to the remaining target examples without additional LLM queries. The resulting fully pseudo-labeled dataset is used to construct in-task demonstrations for ICL. This pipeline combines the flexibility of cross-task supervision with the scalability of LLM-free propagation. Experiments across five tasks demonstrate that our method achieves strong performance while lowering labeling costs.
  •  

Multiple time points for detecting circulating tumor DNA to monitor the response to neoadjuvant therapy in breast cancer: a meta-analysis

BMC Cancer. 2025 Jan 22;25(1):115. doi: 10.1186/s12885-025-13526-0.

ABSTRACT

BACKGROUND: Not all breast cancer (BC) patients can benefit from neoadjuvant therapy (NAT). A poor response may result in patients missing the best opportunity for treatment, ultimately leading to a poor prognosis. Thus, to identify an effective predictor that can assess and predict patient response at early time points, we focused on circulating tumor DNA (ctDNA), which is a vital noninvasive liquid biopsy biomarker. We performed a meta-analysis to explore the predictive value of response by monitoring ctDNA at four time points of NAT using pathologic complete response (pCR) and residual cancer burden (RCB).

METHODS: By searching Embase, PubMed, the Cochrane Library, and the Web of Science until December 24, 2023, we selected studies concerning the relationship between ctDNA and response or prognosis. We analysed the results at the following various time points: baseline (T0), first cycle of NAT (T1), mid-treatment (MT), and end of NAT (EOT). pCR and RCB were used to evaluate the response as the primary endpoint. The secondary endpoint was to investigate the relationship between ctDNA and prognosis. Odds ratios (ORs) and hazard ratios (HRs) were used as effect indicators.

RESULTS: Thirteen reports from twelve studies were eligible for inclusion in this meta-analysis. The results demonstrated that ctDNA negativity was associated with pCR at T1 (OR = 0.34; 95% CI: 0.21-0.57), MT (OR = 0.35; 95% CI: 0.20-0.60), and EOT (OR = 0.38; 95% CI: 0.22-0.66). When RCB was used to evaluate responses, ctDNA negativity was associated with RCB-0/I at the MT (OR = 0.34; 95% CI: 0.21-0.55) and EOT (OR = 0.26; 95% CI: 0.15-0.46). Furthermore, ctDNA positivity at T1 predicted a worse prognosis for patients (HR = 2.73; 95% CI: 1.29-5.75). We also performed a subgroup analysis to more accurately assess the predictive value of ctDNA for triple-negative breast cancer.

CONCLUSIONS: Our meta-analysis suggested that the ctDNA status at the early stage of NAT can predict patient response, which provides evidence for adjusting personalized treatment strategies and improving patient survival.

PROSPERO REGISTRATION NUMBER: CRD42024496465.

PMID:39844103 | PMC:PMC11752932 | DOI:10.1186/s12885-025-13526-0

  •  
❌