Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
arXiv:2512.22334v1 Announce Type: new Abstract: We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Gene
-
cs.AI, q-bio.NC updates on arXiv.org
-
DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
arXiv:2512.22470v1 Announce Type: new Abstract: The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary labels and fail to capture the nuanced psychological and social mechanisms constituting manipulation. We introduce \textbf{DarkPatterns-LLM}, a comprehensive benchmark dataset and diagnostic framework for fine-grai
DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
-
cs.AI, q-bio.NC updates on arXiv.org
-
Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
arXiv:2512.22568v1 Announce Type: new Abstract: The phenomenal advances in large language models (LLMs) and other foundation models over the past few years have been based on optimizing large-scale transformer models on the surprisingly simple objective of minimizing next-token prediction loss, a form of predictive coding that is also the backbone of an increasingly popular model of brain function in neuroscience and cognitive science. However, current foundation models ignore three other impor
Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
-
cs.AI, q-bio.NC updates on arXiv.org
-
Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities
arXiv:2512.23508v1 Announce Type: new Abstract: How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that helps a human to maximise their utility function(s). However, only the human knows these function(s); the AI assistant must learn them. The shutdown problem instead concerns designing AI agents that: shut down w
Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities
-
cs.AI, q-bio.NC updates on arXiv.org
-
Enhancing Medical Data Analysis through AI-Enhanced Locally Linear Embedding: Applications in Medical Point Location and Imagery
arXiv:2512.22182v1 Announce Type: cross Abstract: The rapid evolution of Artificial intelligence in healthcare has opened avenues for enhancing various processes, including medical billing and transcription. This paper introduces an innovative approach by integrating AI with Locally Linear Embedding (LLE) to revolutionize the handling of high-dimensional medical data. This AI-enhanced LLE model is specifically tailored to improve the accuracy and efficiency of medical billing systems and transc
Enhancing Medical Data Analysis through AI-Enhanced Locally Linear Embedding: Applications in Medical Point Location and Imagery
-
cs.AI, q-bio.NC updates on arXiv.org
-
When Algorithms Manage Humans: A Double Machine Learning Approach to Estimating Nonlinear Effects of Algorithmic Control on Gig Worker Performance and Wellbeing
arXiv:2512.22290v1 Announce Type: cross Abstract: A central question for the future of work is whether person centered management can survive when algorithms take on managerial roles. Standard tools often miss what is happening because worker responses to algorithmic systems are rarely linear. We use a Double Machine Learning framework to estimate a moderated mediation model without imposing restrictive functional forms. Using survey data from 464 gig workers, we find a clear nonmonotonic patte
When Algorithms Manage Humans: A Double Machine Learning Approach to Estimating Nonlinear Effects of Algorithmic Control on Gig Worker Performance and Wellbeing
-
cs.AI, q-bio.NC updates on arXiv.org
-
AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents
arXiv:2512.22387v1 Announce Type: cross Abstract: The rise of Large Language Models (LLMs) as coding agents promises to accelerate software development, but their impact on generated code reproducibility remains largely unexplored. This paper presents an empirical study investigating whether LLM-generated code can be executed successfully in a clean environment with only OS packages and using only the dependencies that the model specifies. We evaluate three state-of-the-art LLM coding agents (C
AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents
-
cs.AI, q-bio.NC updates on arXiv.org
-
The body is not there to compute: Comment on "Informational embodiment: Computational role of information structure in codes and robots" by Pitti et al
arXiv:2512.22868v1 Announce Type: cross Abstract: Applying the lens of computation and information has been instrumental in driving the technological progress of our civilization as well as in empowering our understanding of the world around us. The digital computer was and for many still is the leading metaphor for how our mind operates. Information theory (IT) has also been important in our understanding of how nervous systems encode and process information. The target article deploys informa
The body is not there to compute: Comment on "Informational embodiment: Computational role of information structure in codes and robots" by Pitti et al
-
cs.AI, q-bio.NC updates on arXiv.org
-
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
arXiv:2512.23029v1 Announce Type: cross Abstract: The proliferation of Large Language Models (LLMs) has been accompanied by a reliance on cloud-based, proprietary systems, raising significant concerns regarding data privacy, operational sovereignty, and escalating costs. This paper investigates the feasibility of deploying a high-performance, private LLM inference server at a cost accessible to Small and Medium Businesses (SMBs). We present a comprehensive benchmarking analysis of a locally hos
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
-
cs.AI, q-bio.NC updates on arXiv.org
-
Multi-agent Self-triage System with Medical Flowcharts
arXiv:2511.12439v2 Announce Type: replace Abstract: Online health resources and large language models (LLMs) are increasingly used as a first point of contact for medical decision-making, yet their reliability in healthcare remains limited by low accuracy, lack of transparency, and susceptibility to unverified information. We introduce a proof-of-concept conversational self-triage system that guides LLMs with 100 clinically validated flowcharts from the American Medical Association, providing a
Multi-agent Self-triage System with Medical Flowcharts
-
cs.AI, q-bio.NC updates on arXiv.org
-
Taming Data Challenges in ML-based Security Tasks: Lessons from Integrating Generative AI
arXiv:2507.06092v3 Announce Type: replace-cross Abstract: Machine learning-based supervised classifiers are widely used for security tasks, and their improvement has been largely focused on algorithmic advancements. We argue that data challenges that negatively impact the performance of these classifiers have received limited attention. We address the following research question: Can developments in Generative AI (GenAI) address these data challenges and improve classifier performance? We propo
Taming Data Challenges in ML-based Security Tasks: Lessons from Integrating Generative AI
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generating Verifiable Chain of Thoughts from Exection-Traces
arXiv:2512.00127v2 Announce Type: replace-cross Abstract: Teaching language models to reason about code execution remains a fundamental challenge. While Chain-of-Thought (CoT) prompting has shown promise, current synthetic training data suffers from a critical weakness: the reasoning steps are often plausible-sounding explanations generated by teacher models, not verifiable accounts of what the code actually does. This creates a troubling failure mode where models learn to mimic superficially c
Generating Verifiable Chain of Thoughts from Exection-Traces
-
Journal of Medical Internet Research
- Correction: Effects of Internet-Based Cognitive Behavioral Therapy in Routine Care for Adults in Treatment for Depression and Anxiety: Systematic Review and Meta-Analysis
-
Cell Death Discovery nature.com science feeds
-
Modeling hepatocellular carcinoma and its microenvironment on a chip
Cell Death Discovery, Published online: 29 December 2025; doi:10.1038/s41420-025-02917-8Modeling hepatocellular carcinoma and its microenvironment on a chip
Modeling hepatocellular carcinoma and its microenvironment on a chip
Cell Death Discovery, Published online: 29 December 2025; doi:10.1038/s41420-025-02917-8
Modeling hepatocellular carcinoma and its microenvironment on a chip-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Metabolic signatures in gastroenteropancreatic neuroendocrine neoplasms: unraveling diagnostic and prognostic insights
Front Endocrinol (Lausanne). 2025 Dec 11;16:1676021. doi: 10.3389/fendo.2025.1676021. eCollection 2025.ABSTRACTGastroenteropancreatic neuroendocrine neoplasms (GEP-NENs) are a heterogeneous group of tumors characterized by diverse biological behaviors and variable clinical outcomes. Recent advances have highlighted the important role of metabolic reprogramming in tumorigenesis, progression, and therapeutic resistance in GEP-NENs. In this review, we synthesize the current evidence on metabolic bi
Metabolic signatures in gastroenteropancreatic neuroendocrine neoplasms: unraveling diagnostic and prognostic insights
Front Endocrinol (Lausanne). 2025 Dec 11;16:1676021. doi: 10.3389/fendo.2025.1676021. eCollection 2025.
ABSTRACT
Gastroenteropancreatic neuroendocrine neoplasms (GEP-NENs) are a heterogeneous group of tumors characterized by diverse biological behaviors and variable clinical outcomes. Recent advances have highlighted the important role of metabolic reprogramming in tumorigenesis, progression, and therapeutic resistance in GEP-NENs. In this review, we synthesize the current evidence on metabolic biomarkers and altered metabolic pathways-particularly those involving glucose, lipid, and amino acid metabolism. Key biomarkers such as GLUT-1, FASN, and enzymes involved in ferroptosis, cholesterol biosynthesis, and amino acid catabolism demonstrate strong associations with tumor aggressiveness, hypoxia, and mTOR signaling. Moreover, metabolomic profiling and functional studies suggest that metabolic markers may inform prognosis and predict response to targeted therapies such as Everolimus. Although promising, the clinical translation of these markers is still limited and requires further validation in large, subtype-specific cohorts. Our findings highlight the importance of integrating metabolic profiling into the diagnostic and therapeutic landscape of GEP-NENs. Future research should prioritize biomarker standardization, multi-omics integration, and the development of metabolism-based therapeutic strategies tailored to tumor subtype and differentiation grade.
PMID:41458541 | PMC:PMC12738315 | DOI:10.3389/fendo.2025.1676021