Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
TABQAWORLD: Optimizing Multimodal Reasoning for Multi-Turn Table Question Answering
arXiv:2604.03393v1 Announce Type: new Abstract: Multimodal reasoning has emerged as a powerful framework for enhancing reasoning capabilities of reasoning models. While multi-turn table reasoning methods have improved reasoning accuracy through tool use and reward modeling, they rely on fixed text serialization for table state readouts. This introduces representation errors in table encoding that significantly accumulate over multiple turns. Such accumulation is alleviated by tabular grounding
-
cs.AI, q-bio.NC updates on arXiv.org
-
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
arXiv:2604.03893v1 Announce Type: new Abstract: Breakthroughs in frontier theory often depend on the combination of concrete diagrammatic notations with rigorous logic. While multimodal large language models (MLLMs) show promise in general scientific tasks, current benchmarks often focus on local information extraction rather than the global structural logic inherent in formal scientific notations. In this work, we introduce FeynmanBench, the first benchmark centered on Feynman diagram tasks. I
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
-
cs.AI, q-bio.NC updates on arXiv.org
-
RDEx-CMOP: Feasibility-Aware Indicator-Guided Differential Evolution for Fixed-Budget Constrained Multiobjective Optimization
arXiv:2604.03708v1 Announce Type: cross Abstract: Constrained multiobjective optimisation requires fast feasibility attainment together with stable convergence and diversity preservation under strict evaluation budgets. This report documents RDEx-CMOP, the differential evolution variant used in the IEEE CEC 2025 numerical optimisation competition (C06 special session) constrained multiobjective track. RDEx-CMOP integrates an {\epsilon}-level feasibility schedule, a SPEA2-style indicator-driven
RDEx-CMOP: Feasibility-Aware Indicator-Guided Differential Evolution for Fixed-Budget Constrained Multiobjective Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
arXiv:2604.03750v1 Announce Type: cross Abstract: Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities. It supports critical tasks such as vulnerability discovery and malware analysis. Despite its importance, RE remains labor-intensive and requires substantial expertise, making large language models (LLMs) a potential solution for automating the process. However, their capabilities fo
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
-
cs.AI, q-bio.NC updates on arXiv.org
-
Individual and Combined Effects of English as a Second Language and Typos on LLM Performance
arXiv:2604.04723v1 Announce Type: cross Abstract: Large language models (LLMs) are used globally, and because much of their training data is in English, they typically perform best on English inputs. As a result, many non-native English speakers interact with them in English as a second language (ESL), and these inputs often contain typographical errors. Prior work has largely studied the effects of ESL variation and typographical errors separately, even though they often co-occur in real-world
Individual and Combined Effects of English as a Second Language and Typos on LLM Performance
-
cs.AI, q-bio.NC updates on arXiv.org
-
Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration
arXiv:2602.03151v2 Announce Type: replace Abstract: Vision Language Model (VLM) typically assume complete modality input during inference. However, their effectiveness drops sharply when certain modalities are unavailable or incomplete. Current research on missing modality primarily faces two dilemmas: Prompt-based methods struggle to restore missing yet indispensable features and degrade the generalizability of VLM. Imputation-based approaches, lacking effective guidance, are prone to generati
Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generate Then Correct: Single Shot Global Correction for Aspect Sentiment Quad Prediction
arXiv:2603.13777v2 Announce Type: replace-cross Abstract: Aspect-based sentiment analysis (ABSA) extracts aspect-level sentiment signals from user-generated text, supports product analytics, experience monitoring, and public-opinion tracking, and is central to fine-grained opinion mining. A key challenge in ABSA is aspect sentiment quad prediction (ASQP), which requires identifying four elements: the aspect term, the aspect category, the opinion term, and the sentiment polarity. However, existi
Generate Then Correct: Single Shot Global Correction for Aspect Sentiment Quad Prediction
-
cs.AI, q-bio.NC updates on arXiv.org
-
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
arXiv:2604.02029v1 Announce Type: new Abstract: Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an increasing body of work shows that many critical internal processes are more naturally carried out in continuous latent space than in human-readable verbal traces. This shift is driven by the structural limitations of explicit-space computation, including linguistic redunda
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
-
cs.AI, q-bio.NC updates on arXiv.org
-
Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models
arXiv:2604.02236v1 Announce Type: new Abstract: Emotional tone is pervasive in human communication, yet its influence on large language model (LLM) behaviour remains unclear. Here, we examine how first-person emotional framing in user-side queries affect LLM performance across six benchmark domains, including mathematical reasoning, medical question answering, reading comprehension, commonsense reasoning and social inference. Across models and tasks, static emotional prefixes usually produce on
Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Predicting LLM Output Length via Entropy-Guided Representations
arXiv:2602.11812v2 Announce Type: replace Abstract: The long-tailed distribution of sequence lengths in LLM serving and reinforcement learning (RL) sampling causes significant computational waste due to excessive padding in batched inference. Existing methods rely on auxiliary models for static length prediction, but they incur high overhead, generalize poorly, and fail in stochastic "one-to-many" sampling scenarios. We introduce a lightweight framework that reuses the main model's internal hid
Predicting LLM Output Length via Entropy-Guided Representations
-
cs.AI, q-bio.NC updates on arXiv.org
-
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
arXiv:2603.29112v1 Announce Type: new Abstract: We introduce GISTBench, a benchmark for evaluating Large Language Models' (LLMs) ability to understand users from their interaction histories in recommendation systems. Unlike traditional RecSys benchmarks that focus on item prediction accuracy, our benchmark evaluates how well LLMs can extract and verify user interests from engagement data. We propose two novel metric families: Interest Groundedness (IG), decomposed into precision and recall comp
GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
-
cs.AI, q-bio.NC updates on arXiv.org
-
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
arXiv:2603.29386v1 Announce Type: cross Abstract: The rapid democratization of prompt-based AI image editing has recently exacerbated the risks associated with malicious content fabrication and misinformation. However, forgery localization methods targeting these emerging editing techniques remain significantly under-explored. To bridge this gap, we first introduce a fully automated mask annotating framework that leverages keypoint alignment and semantic space similarity to generate precise gro
PromptForge-350k: A Large-Scale Dataset and Contrastive Framework for Prompt-Based AI Image Forgery Localization
-
cs.AI, q-bio.NC updates on arXiv.org
-
Memory Bear AI Memory Science Engine for Multimodal Affective Intelligence: A Technical Report
arXiv:2603.22306v1 Announce Type: new Abstract: Affective judgment in real interaction is rarely a purely local prediction problem. Emotional meaning often depends on prior trajectory, accumulated context, and multimodal evidence that may be weak, noisy, or incomplete at the current moment. Although multimodal emotion recognition (MER) has improved the integration of text, speech, and visual signals, many existing systems remain optimized for short-range inference and provide limited support fo
Memory Bear AI Memory Science Engine for Multimodal Affective Intelligence: A Technical Report
-
cs.AI, q-bio.NC updates on arXiv.org
-
Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning
arXiv:2603.22317v1 Announce Type: cross Abstract: Graph-structured data typically exhibits complex topological heterogeneity, making it difficult to model accurately within a single Riemannian manifold. While emerging mixed-curvature methods attempt to capture such diversity, they often rely on implicit, task-driven routing that lacks fundamental geometric grounding. To address this challenge, we propose a Geometric Mixture-of-Experts framework (GeoMoE) that adaptively fuses node representation
Geometric Mixture-of-Experts with Curvature-Guided Adaptive Routing for Graph Representation Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
arXiv:2302.10426v3 Announce Type: replace Abstract: Trustworthy process monitoring seeks to build an accurate and interpretable monitoring framework, which is critical for ensuring the safety of energy conversion plant (ECP) that operates under extreme working conditions such as high pressure and temperature. Contemporary self-attentive models, however, fall short in this domain for two main reasons. First, they rely on step-wise correlations that fail to involve physically meaningful semantics
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
-
Omics in Hepatocellular
-
Comment on "Integrated multi-Omics and network toxicology elucidate the multi-target mechanisms of environmental hormones in driving hepatocellular carcinoma"
Ecotoxicol Environ Saf. 2026 Apr 1;314:120063. doi: 10.1016/j.ecoenv.2026.120063. Epub 2026 Mar 23.NO ABSTRACTPMID:41875554 | DOI:10.1016/j.ecoenv.2026.120063
Comment on "Integrated multi-Omics and network toxicology elucidate the multi-target mechanisms of environmental hormones in driving hepatocellular carcinoma"
Ecotoxicol Environ Saf. 2026 Apr 1;314:120063. doi: 10.1016/j.ecoenv.2026.120063. Epub 2026 Mar 23.
NO ABSTRACT
PMID:41875554 | DOI:10.1016/j.ecoenv.2026.120063
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
Comment on "Integrated multi-Omics and network toxicology elucidate the multi-target mechanisms of environmental hormones in driving hepatocellular carcinoma"
Ecotoxicol Environ Saf. 2026 Mar 23;314:120063. doi: 10.1016/j.ecoenv.2026.120063. Online ahead of print.NO ABSTRACTPMID:41875554 | DOI:10.1016/j.ecoenv.2026.120063
Comment on "Integrated multi-Omics and network toxicology elucidate the multi-target mechanisms of environmental hormones in driving hepatocellular carcinoma"
Ecotoxicol Environ Saf. 2026 Mar 23;314:120063. doi: 10.1016/j.ecoenv.2026.120063. Online ahead of print.
NO ABSTRACT
PMID:41875554 | DOI:10.1016/j.ecoenv.2026.120063
-
(Multiomics OR Omics) AND (Lung OR gastric OR Hepatocellular)
-
NUP85 as a Pan-Cancer Immune Biomarker: Integrated Multi Omics and Functional Analyses Reveal Its Role in Tumor Prognosis
Immunotargets Ther. 2026 Mar 17;15:541852. doi: 10.2147/ITT.S541852. eCollection 2026.ABSTRACTPURPOSE: NUP85 encodes protein components of the Nup107-160 subunit of the nuclear pore complex, belonging to the Nucleoporins (NUPs) family, potentially implicating its role in human cancer. This study aims to elucidate the potential involvement of NUP85 in cancer pathogenesis.METHODS: Leveraging data from The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), Clinical Proteomic Tumor Analy
NUP85 as a Pan-Cancer Immune Biomarker: Integrated Multi Omics and Functional Analyses Reveal Its Role in Tumor Prognosis
Immunotargets Ther. 2026 Mar 17;15:541852. doi: 10.2147/ITT.S541852. eCollection 2026.
ABSTRACT
PURPOSE: NUP85 encodes protein components of the Nup107-160 subunit of the nuclear pore complex, belonging to the Nucleoporins (NUPs) family, potentially implicating its role in human cancer. This study aims to elucidate the potential involvement of NUP85 in cancer pathogenesis.
METHODS: Leveraging data from The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), Clinical Proteomic Tumor Analysis Consortium (CPTAC), Cancer Cell Line Encyclopedia (CCLE), Human Protein Atlas (HPA), Gene Expression Profiling Interactive Analysis (GEPIA), CellMiner, and GeneMANIA databases, we investigated the role of NUP85 across various tumors. Correlations between NUP85 expression and pathological stage, histological grade, survival, immune infiltration, tumor mutational burden (TMB), microsatellite instability (MSI), drug resistance, DNA methylation, copy number variation (CNV), and single-cell expression were analyzed. Gene functional enrichment analysis was conducted to explore NUP85-associated pathways. Molecular biology experiments including Western blotting, flow cytometry, trans-well migration, and invasion assays were performed to validate NUP85's oncogenic role in lung adenocarcinoma (LUAD) and oral squamous cell carcinoma (OSCC) cell lines.
RESULTS: Our findings reveal up-regulated expression of NUP85 in most tumor tissues, with significant correlations observed with pathological stage, survival, immune infiltration, TMB, MSI, drug resistance, DNA methylation, and CNV. Molecular biology experiments confirm NUP85's tumor-promoting role in LUAD and OSCC cell lines. Single-cell sequencing data suggest elevated NUP85 expression primarily in proliferative T cells (Tprolif).
CONCLUSION: NUP85 emerges as a potential tumor marker associated with tumor immunity and poor prognosis. These insights offer avenues for the development of novel therapeutic targets and anti-neoplastic drugs.
PMID:41869435 | PMC:PMC13005628 | DOI:10.2147/ITT.S541852
-
cs.AI, q-bio.NC updates on arXiv.org
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
arXiv:2602.12670v3 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment LLM agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark of 86 tasks across 11 domains paired with curated Skills and deterministic verifiers. Each task is evaluated under three conditions: no Skills, curated Skills, and self-generated Skills. We test 7 agent-model configurat
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
-
cs.AI, q-bio.NC updates on arXiv.org
-
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation
arXiv:2601.15369v2 Announce Type: replace-cross Abstract: This paper presents a family of advanced vision encoder, named OpenVision 3, that learns a single, unified visual representation that can serve both image understanding and image generation. Our core architecture is simple: we feed VAE-compressed image latents to a ViT encoder and train its output to support two complementary roles. First, the encoder output is passed to the ViT-VAE decoder to reconstruct the original image, encouraging