❌

Reading view

Artificial Intelligence Applications in the Diagnosis, Treatment, and Prognosis of Hepatocellular Carcinoma

Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.

ABSTRACT

The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.

PMID:41472345 | DOI:10.5009/gnl250268

  •  

Multi-Omics and Functional Analysis of BFSP1 as a Prognostic and Therapeutic Target in Liver Hepatocellular Carcinoma

Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.

ABSTRACT

Background and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was performed across multiple platforms using datasets derived from The Cancer Genome Atlas. Differential gene expression, DNA methylation, copy number variation, immune cell infiltration, drug sensitivity, and co-expression networks were systematically examined. Functional enrichment analyses of protein-protein and gene-gene interaction networks were conducted using STRING and GeneMANIA. Additionally, short interfering RNA-mediated knockdown and wound-healing assays were performed in HepG2 cells to evaluate BFSP1 function in vitro. Results: The results showed that BFSP1 mRNA expression was significantly upregulated in tissues from LIHC patients. Elevated BFSP1 levels were associated with poorer prognostic patterns, which were further supported by detailed clinicopathological subgroup analyses. Furthermore, BFSP1 expression was correlated with promoter hypomethylation and associated with patterns of tumor-infiltrating immune cells, including specific immune cell subtypes such as M1 and M2 macrophages. Integrative analyses revealed strong associations between BFSP1 and drug sensitivity, as well as a regulatory network encompassing genes involved in the cell cycle, DNA repair, and metabolic processes. Functional knockdown of BFSP1 significantly reduced HepG2 cell migration in vitro, as assessed by wound healing assay, with decreased wound closure at 24 h (11.0% vs. 16.5%) and 48 h (7.4% vs. 12.5%) compared with the control (p < 0.05, n = 6 biological replicates). Conclusions: In conclusion, these findings suggest that BFSP1 functions as a multifaceted prognostic biomarker and a potential therapeutic target for LIHC.

PMID:41470198 | PMC:PMC12735119 | DOI:10.3390/medicina61122196

  •  

Artificial Intelligence Applications in the Diagnosis, Treatment, and Prognosis of Hepatocellular Carcinoma

Gut Liver. 2025 Dec 31. doi: 10.5009/gnl250268. Online ahead of print.

ABSTRACT

The global burden of hepatocellular carcinoma (HCC) has shifted from viral to nonviral etiologies. However, successful antiviral therapy does not fully eliminate the risk of HCC, underscoring the demand for more effective surveillance strategies. Current screening methods, such as semiannual ultrasonography and the measurement of Ξ±-fetoprotein levels, offer suboptimal sensitivity for early detection. A cost-effective, reliable surveillance approach remains an unmet need. The Barcelona Clinic Liver Cancer staging system provides a framework to guide HCC therapy; yet, some gray zone exists, particularly for patients with intermediate-stage disease. Although tyrosine kinase inhibitors and immunotherapies have transformed the therapeutic landscape, their efficacies vary among patients, highlighting the necessity for personalized treatment strategies. In response to these challenges, artificial intelligence (AI) approaches have emerged as transformative tools in healthcare. By processing complex, nonlinear relationships and uncovering hidden patterns in clinical data, AI methods offer capabilities beyond those of traditional statistical methods. Furthermore, AI-driven multi-omics analysis holds promise for identifying novel biomarkers, thereby advancing precision medicine for HCC patients. This review introduces the potential of AI applications in enhancing the diagnosis, treatment, and prognosis of HCC.

PMID:41472345 | DOI:10.5009/gnl250268

  •  

Multi-Omics and Functional Analysis of BFSP1 as a Prognostic and Therapeutic Target in Liver Hepatocellular Carcinoma

Medicina (Kaunas). 2025 Dec 11;61(12):2196. doi: 10.3390/medicina61122196.

ABSTRACT

Background and Objectives: Although beaded filament structural protein 1 (BFSP1) may be involved in oncogenic mechanisms, its clinical relevance and functional role in liver hepatocellular carcinoma (LIHC) remain unclear. This study examined the prognostic significance, regulatory mechanisms, and potential therapeutic implications of BFSP1 in LIHC. Materials and Methods: Comprehensive bioinformatics analysis was performed across multiple platforms using datasets derived from The Cancer Genome Atlas. Differential gene expression, DNA methylation, copy number variation, immune cell infiltration, drug sensitivity, and co-expression networks were systematically examined. Functional enrichment analyses of protein-protein and gene-gene interaction networks were conducted using STRING and GeneMANIA. Additionally, short interfering RNA-mediated knockdown and wound-healing assays were performed in HepG2 cells to evaluate BFSP1 function in vitro. Results: The results showed that BFSP1 mRNA expression was significantly upregulated in tissues from LIHC patients. Elevated BFSP1 levels were associated with poorer prognostic patterns, which were further supported by detailed clinicopathological subgroup analyses. Furthermore, BFSP1 expression was correlated with promoter hypomethylation and associated with patterns of tumor-infiltrating immune cells, including specific immune cell subtypes such as M1 and M2 macrophages. Integrative analyses revealed strong associations between BFSP1 and drug sensitivity, as well as a regulatory network encompassing genes involved in the cell cycle, DNA repair, and metabolic processes. Functional knockdown of BFSP1 significantly reduced HepG2 cell migration in vitro, as assessed by wound healing assay, with decreased wound closure at 24 h (11.0% vs. 16.5%) and 48 h (7.4% vs. 12.5%) compared with the control (p < 0.05, n = 6 biological replicates). Conclusions: In conclusion, these findings suggest that BFSP1 functions as a multifaceted prognostic biomarker and a potential therapeutic target for LIHC.

PMID:41470198 | PMC:PMC12735119 | DOI:10.3390/medicina61122196

  •  

Effectiveness of Digital Interventions for Low-Income, Food-Insecure Populations: Natural Language Processing Study of WIC Smartphone App User Reviews, 2013-2024

Background: The Special Supplemental Nutrition Program for Women, Infants, and Children (WIC) is a federal nutrition assistance program for low-income, food-insecure mothers and young children in the United States. Despite its intended goals, many eligible individuals forgo WIC benefits, in part due to administrative burden – defined as the complex, often frustrating processes encountered when navigating public benefits programs. In response, a range of digital interventions and policy waivers were introduced during the COVID-19 pandemic, but their effectiveness in reducing barriers remains unclear. Objective: Drawing from administrative burden theory and human-computer interaction (HCI) research, this study examined user reviews of WIC smartphone applications (WIC Apps) utilized by local agencies. Specifically, it investigated (a) how obstacles to WIC access manifested in daily app use, (b) how user experiences shifted after the onset of the COVID-19 pandemic, and (c) how these changes were associated with app ratings. Methods: An original dataset of user reviews (Nreview = 28,212) was compiled for 26 WIC Apps between 2013 and 2024. Structural topic modeling identified eight key themes, and sentiment was examined with RoBERTa (Robustly Optimized Bidirectional Encoder Representations from Transformers Pretraining Approach). Analyses compared topic prevalence and sentiment distributions before and after COVID-19. Mixed-effects models examined the relationship between topics, sentiment, and app ratings. Results: Technical concerns related to account authentication and login, document upload, and app updates were among the most prevalent themes. These issues were typically expressed with negative sentiment and appeared more frequently in pre-COVID-19 reviews than in post-COVID-19 reviews. Although reliability problems (e.g., outages, maintenance) persisted, post-COVID-19 reviews increasingly emphasized features that facilitated program tracking, shopping and benefit redemption, and general ease of use, which were generally described with positive sentiment. Mixed-effects analyses indicated that the post-COVID-19 topics were significantly associated with higher app ratings (program tracking: B = 0.21, SE = 0.06, P = .001, shopping and redemption: B = 0.18, SE = 0.07, P = .01, ease of use: B = 0.10, SE = 0.05, P = .04), whereas pre-COVID-19 concerns were not associated with ratings (Ps > .05). When sentiment was added to the mixed-effect model, it became the dominant factor: negative sentiment was associated with lower ratings (B = -1.71, SE = 0.03, P .05), suggesting that sentiment contributed to much of the variance previously linked to topics. Conclusions: User-centered digital interventions, such as WIC Apps, have potential to support WIC access and participation.
  •  

PIC-SURE: an open-source platform for integrating clinical and genomic data

npj Digital Medicine, Published online: 30 December 2025; doi:10.1038/s41746-025-02284-9

PIC-SURE: an open-source platform for integrating clinical and genomic data
  •  

Bidirectional RAG: Safe Self-Improving Retrieval-Augmented Generation Through Multi-Stage Validation

arXiv:2512.22199v1 Announce Type: new Abstract: Retrieval-Augmented Generation RAG systems enhance large language models by grounding responses in external knowledge bases, but conventional RAG architectures operate with static corpora that cannot evolve from user interactions. We introduce Bidirectional RAG, a novel RAG architecture that enables safe corpus expansion through validated write back of high quality generated responses. Our system employs a multi stage acceptance layer combining grounding verification (NLI based entailment, attribution checking, and novelty detection to prevent hallucination pollution while enabling knowledge accumulation. Across four datasets Natural Questions, TriviaQA, HotpotQA, Stack Overflow with three random seeds 12 experiments per system, Bidirectional RAG achieves 40.58% average coverage nearly doubling Standard RAG 20.33% while adding 72% fewer documents than naive write back 140 vs 500. Our work demonstrates that self improving RAG is feasible and safe when governed by rigorous validation, offering a practical path toward RAG systems that learn from deployment.
  •  

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.
  •  

DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior

arXiv:2512.22470v1 Announce Type: new Abstract: The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grained assessment of manipulative content in LLM outputs across seven harm categories: Legal/Power, Psychological, Emotional, Physical, Autonomy, Economic, and Societal Harm. Our framework implements a four-layer analytical pipeline comprising Multi-Granular Detection (MGD), Multi-Scale Intent Analysis (MSIAN), Threat Harmonization Protocol (THP), and Deep Contextual Risk Alignment (DCRA). The dataset contains 401 meticulously curated examples with instruction-response pairs and expert annotations. Through evaluation of state-of-the-art models including GPT-4, Claude 3.5, and LLaMA-3-70B, we observe significant performance disparities (65.2\%--89.7\%) and consistent weaknesses in detecting autonomy-undermining patterns. DarkPatterns-LLM establishes the first standardized, multi-dimensional benchmark for manipulation detection in LLMs, offering actionable diagnostics toward more trustworthy AI systems.
  •  

Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI

arXiv:2512.22568v1 Announce Type: new Abstract: The phenomenal advances in large language models (LLMs) and other foundation models over the past few years have been based on optimizing large-scale transformer models on the surprisingly simple objective of minimizing next-token prediction loss, a form of predictive coding that is also the backbone of an increasingly popular model of brain function in neuroscience and cognitive science. However, current foundation models ignore three other important components of state-of-the-art predictive coding models: tight integration of actions with generative models, hierarchical compositional structure, and episodic memory. We propose that to achieve safe, interpretable, energy-efficient, and human-like AI, foundation models should integrate actions, at multiple scales of abstraction, with a compositional generative architecture and episodic memory. We present recent evidence from neuroscience and cognitive science on the importance of each of these components. We describe how the addition of these missing components to foundation models could help address some of their current deficiencies: hallucinations and superficial understanding of concepts due to lack of grounding, a missing sense of agency/responsibility due to lack of control, threats to safety and trustworthiness due to lack of interpretability, and energy inefficiency. We compare our proposal to current trends, such as adding chain-of-thought (CoT) reasoning and retrieval-augmented generation (RAG) to foundation models, and discuss new ways of augmenting these models with brain-inspired components. We conclude by arguing that a rekindling of the historically fruitful exchange of ideas between brain science and AI will help pave the way towards safe and interpretable human-centered AI.
  •  

Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities

arXiv:2512.23508v1 Announce Type: new Abstract: How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down when a shutdown button is pressed; neither try to prevent nor cause the pressing of the shutdown button; and otherwise accomplish their task competently. In this paper, we show that addressing these challenges requires AI agents that can reason under uncertainty and handle both incomplete and non-Archimedean preferences.
  •  

Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments

arXiv:2512.22149v1 Announce Type: cross Abstract: Multi-agent systems powered by large language models have emerged as a promising paradigm for solving complex reasoning tasks through collaborative intelligence. However, efficiently deploying these systems on serverless GPU platforms presents significant resource allocation challenges due to heterogeneous agent workloads, varying computational demands, and the need for cost-effective scaling. This paper presents an adaptive GPU resource allocation framework that achieves 85\% latency reduction compared to round-robin scheduling while maintaining comparable throughput to static allocation, using an $O(N)$ complexity algorithm for real-time adaptation. Our approach dynamically allocates GPU resources based on workload characteristics, agent priorities, and minimum resource requirements, enabling efficient utilization while maintaining quality of service. The framework addresses three key challenges: (1) heterogeneous computational demands across lightweight coordinators and heavyweight specialists, (2) dynamic workload fluctuations requiring millisecond-scale reallocation, and (3) capacity constraints in serverless environments. Through comprehensive simulations modeling realistic multi-agent workflows with four heterogeneous agents, we demonstrate that adaptive allocation outperforms static equal and round-robin strategies across latency, cost, and GPU utilization metrics. The framework provides a practical solution for deploying cost-efficient multi-agent AI systems on serverless GPU infrastructure.
  •  

Interpretable Link Prediction in AI-Driven Cancer Research: Uncovering Co-Authorship Patterns

arXiv:2512.22181v1 Announce Type: cross Abstract: Artificial intelligence (AI) is transforming cancer diagnosis and treatment. The intricate nature of this disease necessitates the collaboration of diverse stakeholders with varied expertise to ensure the effectiveness of cancer research. Despite its importance, forming effective interdisciplinary research teams remains challenging. Understanding and predicting collaboration patterns can help researchers, organizations, and policymakers optimize resources and foster impactful research. We examined co-authorship networks as a proxy for collaboration within AI-driven cancer research. Using 7,738 publications (2000-2017) from Scopus, we constructed 36 overlapping co-authorship networks representing new, persistent, and discontinued collaborations. We engineered both attribute-based and structure-based features and built four machine learning classifiers. Model interpretability was performed using Shapley Additive Explanations (SHAP). Random forest achieved the highest recall for all three types of examined collaborations. The discipline similarity score emerged as a crucial factor, positively affecting new and persistent patterns while negatively impacting discontinued collaborations. Additionally, high productivity and seniority were positively associated with discontinued links. Our findings can guide the formation of effective research teams, enhance interdisciplinary cooperation, and inform strategic policy decisions.
  •  

Enhancing Medical Data Analysis through AI-Enhanced Locally Linear Embedding: Applications in Medical Point Location and Imagery

arXiv:2512.22182v1 Announce Type: cross Abstract: The rapid evolution of Artificial intelligence in healthcare has opened avenues for enhancing various processes, including medical billing and transcription. This paper introduces an innovative approach by integrating AI with Locally Linear Embedding (LLE) to revolutionize the handling of high-dimensional medical data. This AI-enhanced LLE model is specifically tailored to improve the accuracy and efficiency of medical billing systems and transcription services. By automating these processes, the model aims to reduce human error and streamline operations, thereby facilitating faster and more accurate patient care documentation and financial transactions. This paper provides a comprehensive mathematical model of AI-enhanced LLE, demonstrating its application in real-world healthcare scenarios through a series of experiments. The results indicate a significant improvement in data processing accuracy and operational efficiency. This study not only underscores the potential of AI-enhanced LLE in medical data analysis but also sets a foundation for future research into broader healthcare applications.
  •  

Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening

arXiv:2512.22242v1 Announce Type: cross Abstract: Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p
  •  

Interpretable Perturbation Modeling Through Biomedical Knowledge Graphs

arXiv:2512.22251v1 Announce Type: cross Abstract: Understanding how small molecules perturb gene expression is essential for uncovering drug mechanisms, predicting off-target effects, and identifying repurposing opportunities. While prior deep learning frameworks have integrated multimodal embeddings into biomedical knowledge graphs (BKGs) and further improved these representations through graph neural network message-passing paradigms, these models have been applied to tasks such as link prediction and binary drug-disease association, rather than the task of gene perturbation, which may unveil more about mechanistic transcriptomic effects. To address this gap, we construct a merged biomedical graph that integrates (i) PrimeKG++, an augmentation of PrimeKG containing semantically rich embeddings for nodes with (ii) LINCS L1000 drug and cell line nodes, initialized with multimodal embeddings from foundation models such as MolFormerXL and BioBERT. Using this heterogeneous graph, we train a graph attention network (GAT) with a downstream prediction head that learns the delta expression profile of over 978 landmark genes for a given drug-cell pair. Our results show that our framework outperforms MLP baselines for differentially expressed genes (DEG) -- which predict the delta expression given a concatenated embedding of drug features, target features, and baseline cell expression -- under the scaffold and random splits. Ablation experiments with edge shuffling and node feature randomization further demonstrate that the edges provided by biomedical KGs enhance perturbation-level prediction. More broadly, our framework provides a path toward mechanistic drug modeling: moving beyond binary drug-disease association tasks to granular transcriptional effects of therapeutic intervention.
  •  

LLM-Guided Exemplar Selection for Few-Shot Wearable-Sensor Human Activity Recognition

arXiv:2512.22385v1 Announce Type: cross Abstract: In this paper, we propose an LLM-Guided Exemplar Selection framework to address a key limitation in state-of-the-art Human Activity Recognition (HAR) methods: their reliance on large labeled datasets and purely geometric exemplar selection, which often fail to distinguish similar weara-ble sensor activities such as walking, walking upstairs, and walking downstairs. Our method incorporates semantic reasoning via an LLM-generated knowledge prior that captures feature importance, inter-class confusability, and exemplar budget multipliers, and uses it to guide exemplar scoring and selection. These priors are combined with margin-based validation cues, PageRank centrality, hubness penalization, and facility-location optimization to obtain a compact and informative set of exemplars. Evaluated on the UCI-HAR dataset under strict few-shot conditions, the framework achieves a macro F1-score of 88.78%, outperforming classical approaches such as random sampling, herding, and $k$-center. The results show that LLM-derived semantic priors, when integrated with structural and geometric cues, provide a stronger foundation for selecting representative sensor exemplars in few-shot wearable-sensor HAR.
  •  

AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents

arXiv:2512.22387v1 Announce Type: cross Abstract: The rise of Large Language Models (LLMs) as coding agents promises to accelerate software development, but their impact on generated code reproducibility remains largely unexplored. This paper presents an empirical study investigating whether LLM-generated code can be executed successfully in a clean environment with only OS packages and using only the dependencies that the model specifies. We evaluate three state-of-the-art LLM coding agents (Claude Code, OpenAI Codex, and Gemini) across 300 projects generated from 100 standardized prompts in Python, JavaScript, and Java. We introduce a three-layer dependency framework (distinguishing between claimed, working, and runtime dependencies) to quantify execution reproducibility. Our results show that only 68.3% of projects execute out-of-the-box, with substantial variation across languages (Python 89.2%, Java 44.0%). We also find a 13.5 times average expansion from declared to actual runtime dependencies, revealing significant hidden dependencies.
  •  

Harnessing Large Language Models for Biomedical Named Entity Recognition

arXiv:2512.22738v1 Announce Type: cross Abstract: Background and Objective: Biomedical Named Entity Recognition (BioNER) is a foundational task in medical informatics, crucial for downstream applications like drug discovery and clinical trial matching. However, adapting general-domain Large Language Models (LLMs) to this task is often hampered by their lack of domain-specific knowledge and the performance degradation caused by low-quality training data. To address these challenges, we introduce BioSelectTune, a highly efficient, data-centric framework for fine-tuning LLMs that prioritizes data quality over quantity. Methods and Results: BioSelectTune reformulates BioNER as a structured JSON generation task and leverages our novel Hybrid Superfiltering strategy, a weak-to-strong data curation method that uses a homologous weak model to distill a compact, high-impact training dataset. Conclusions: Through extensive experiments, we demonstrate that BioSelectTune achieves state-of-the-art (SOTA) performance across multiple BioNER benchmarks. Notably, our model, trained on only 50% of the curated positive data, not only surpasses the fully-trained baseline but also outperforms powerful domain-specialized models like BioMedBERT.
  •  
❌