Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
arXiv:2609.10335v1 Announce Type: new Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when equipped with specialized modules, can rival state-of-the-art LMMs on complex geometry problems. Our framework integrates a Geomet
-
cs.AI, q-bio.NC updates on arXiv.org
-
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
arXiv:2605.23965v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing evaluations rely on static benchmarks, which fail to assess robustness under logically equivalent transformations and often overestimate reasoning capability. We propose LGMT (Logic-Grounded Metamorphic Testing), an oracle-free framework that leverages first-order logic (FOL) to evaluate LLM reasoning. By deriv
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
-
cs.AI, q-bio.NC updates on arXiv.org
-
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS
arXiv:2605.25389v1 Announce Type: cross Abstract: While Large Language Model-based Multi-Agent Systems (LLM-MAS) demonstrate remarkable capabilities in solving complex tasks by orchestrating specialized agents and external tools, the implicit trust in tool outputs creates a critical attack surface. Existing tool attacks are limited by domain specificity or fixed and static templates. To address these challenges, we propose Evo-Attacker, which formulates the tool attack as a self-evolving, memor
Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS
-
cs.AI, q-bio.NC updates on arXiv.org
-
Geometric Flow Matching for Molecular Conformation Generation via Manifold Decomposition
arXiv:2605.25577v1 Announce Type: cross Abstract: The generation of accurate 3D molecular conformations is a pivotal challenge in computational chemistry and drug discovery. Recently, diffusion and flow matching models have achieved remarkable success. However, there is a critical misalignment between their mathematical formulation and the physical reality of molecules. Existing approaches predominantly treat molecules as unstructured point clouds in Cartesian space, overlooking the intrinsic h
Geometric Flow Matching for Molecular Conformation Generation via Manifold Decomposition
-
cs.AI, q-bio.NC updates on arXiv.org
-
A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
arXiv:2605.13850v2 Announce Type: replace Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus on execution topology -- how data flows -- while cognitive science surveys focus on cognitive function -- what the agent does. Neither axis alone disambiguates architecturally distinct systems: the same Orchestrator-Workers topology can implement Plan-and-Execute, Hierarchical Delegation, or Adv
A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Paper to Program: A Multi-Stage LLM-Assisted Workflow for Accelerating Quantum Many-Body Algorithm Development
arXiv:2604.04089v1 Announce Type: cross Abstract: Translating quantum many-body theory into scalable software traditionally requires months of effort. Zero-shot generation of tensor network algorithms by Large Language Models (LLMs) frequently fails due to spatial reasoning errors and memory bottlenecks. We resolve this using a multi-stage workflow that mimics a physics research group. By generating a mathematically rigorous LaTeX specification as an intermediate blueprint, we constrain the cod
From Paper to Program: A Multi-Stage LLM-Assisted Workflow for Accelerating Quantum Many-Body Algorithm Development
-
cs.AI, q-bio.NC updates on arXiv.org
-
One Model for All: Multi-Objective Controllable Language Models
arXiv:2604.04497v1 Announce Type: cross Abstract: Aligning large language models (LLMs) with human preferences is critical for enhancing LLMs' safety, helpfulness, humor, faithfulness, etc. Current reinforcement learning from human feedback (RLHF) mainly focuses on a fixed reward learned from average human ratings, which may weaken the adaptability and controllability of varying preferences. However, creating personalized LLMs requires aligning LLMs with individual human preferences, which is n
One Model for All: Multi-Objective Controllable Language Models
-
cs.AI, q-bio.NC updates on arXiv.org
-
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
arXiv:2603.29252v1 Announce Type: cross Abstract: Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and proposed a novel and training-free approach, termed \emph{Flexible Memory} (\textbf{FlexMem}). In principle, FlexMem aims to mimic human behavior of video watching, \emph{i.e.}, continually watching video content and recalling the mos
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
-
cs.AI, q-bio.NC updates on arXiv.org
-
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
arXiv:2603.29888v1 Announce Type: cross Abstract: In collaboration with Alibaba, this study leverages a large-scale field experiment to assess the impact of a generative AI assistant on worker performance in e-commerce after-sales service. Human agents providing digital chat support were randomly assigned with access to a gen AI assistant that offered two core functions: diagnosis of customer issues and solution proposals, presented as text messages. Agents retained discretion to adopt, modify,
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
-
Omics in Hepatocellular
-
Proposed Role of Circadian Clock Genes in Pathogenesis of HCC: Molecular Subtyping and Characterization
Biomedicines. 2026 Mar 12;14(3):645. doi: 10.3390/biomedicines14030645.ABSTRACTBackground: Hepatocellular carcinoma (HCC) stands as a prevalent global health issue with increasing incidence and mortality rates. Hepatocellular carcinoma (HCC) exhibits profound molecular and clinical heterogeneity, which limits the effectiveness of current therapeutic strategies. Circadian rhythm disruption has been implicated in metabolic reprogramming, proliferation, and immune modulation in cancer, but its role
Proposed Role of Circadian Clock Genes in Pathogenesis of HCC: Molecular Subtyping and Characterization
Biomedicines. 2026 Mar 12;14(3):645. doi: 10.3390/biomedicines14030645.
ABSTRACT
Background: Hepatocellular carcinoma (HCC) stands as a prevalent global health issue with increasing incidence and mortality rates. Hepatocellular carcinoma (HCC) exhibits profound molecular and clinical heterogeneity, which limits the effectiveness of current therapeutic strategies. Circadian rhythm disruption has been implicated in metabolic reprogramming, proliferation, and immune modulation in cancer, but its role in shaping HCC heterogeneity remains poorly defined. Methods: Four public HCC transcriptomic cohorts (TCGA-LIHC, CHCC, LIRI, LICA) were integrated using RMA normalization and ComBat for batch correction. Consensus clustering based on 31 core circadian clock genes (CCGs) identified robust molecular subtypes. Multi-omics characterization-including genomic alterations, pathway activity (GSEA/GSVA), immune microenvironment profiling (CIBERSORT, EPIC, MCP-counter, xCell), and drug-sensitivity prediction (pRRophetic/oncoPredict)-was performed to delineate subtype-specific biological properties. A nine-gene CCG-based RiskScore model was constructed using LASSO Cox regression to internally validate subtype robustness and intra-subtype risk stratification. Results: Using consensus clustering of 31 core CCGs in TCGA-LIHC and three independent validation cohorts (CHCC, LIRI, LICA), we identified three reproducible subtypes-Cluster-1 (metabolic-quiescent), Cluster-2 (transition-intermediate), and Cluster-3 (proliferation-inflammatory)-which were recapitulated across cohorts and showed distinct overall survival (Cluster-3 worst; log-rank p values significant across datasets). Multi-omic characterization revealed that Cluster-3 exhibits the highest tumor mutational burden and CNV burden with enrichment of TP53/AXIN1/TERT alterations, strong activation of cell-cycle, E2F, and G2M programs, and an immune-hot yet immunosuppressed microenvironment enriched for TAMs, Tregs and MDSCs. By contrast, Cluster-1 shows relative genomic stability, dominant hepatic metabolic signatures (fatty-acid oxidation, bile-acid and xenobiotic metabolism) and an immune-cold phenotype. Single-cell mapping linked ALAS1 expression to malignant hepatocytes predominating in Cluster-1, whereas NONO and CSNK1D localized to stromal (CAFs/TECs) and both malignant/immune compartments respectively in Cluster-3, providing a cellular mechanism for subtype-specific metabolism, angiogenesis and immune modulation. Finally, a nine-gene CCG-based RiskScore validated prognostic stratification and drug-sensitivity predictions indicated subtype-specific therapeutic vulnerabilities (notably increased predicted TKI sensitivity in Cluster-3). Conclusion: In conclusion, this study proposes a robust circadian rhythm-based molecular classification of hepatocellular carcinoma, revealing three biologically and clinically distinct subtypes characterized by divergent genomic alterations, metabolic programs, immune microenvironment states, and prognostic patterns. By integrating bulk and single-cell transcriptomic data, we identify subtype-specific roles of key circadian regulators-including ALAS1, NONO, and CSNK1D-in shaping tumor metabolism, proliferation, stromal remodeling, and immune suppression. These findings highlight circadian dysregulation as a potential upstream factor associated with HCC heterogeneity and provide a conceptual framework for developing subtype-tailored mechanistic studies and circadian-informed therapeutic strategies.
PMID:41898292 | PMC:PMC13024568 | DOI:10.3390/biomedicines14030645
-
cs.AI, q-bio.NC updates on arXiv.org
-
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
arXiv:2603.22911v1 Announce Type: cross Abstract: Due to the great saving of computation and memory overhead, token compression has become a research hot-spot for MLLMs and achieved remarkable progress in image-language tasks. However, for the video, existing methods still fall short of high-ratio token compression. We attribute this shortcoming to the insufficient modeling of temporal and continual video content, and propose a novel and training-free token pruning method for video MLLMs, terme
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
-
cs.AI, q-bio.NC updates on arXiv.org
-
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
arXiv:2510.00415v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) and agent system designs have empowered agents with unprecedented levels of capability. However, existing agent benchmarks are showing a trend of rapid ceiling-hitting by newly developed agents, making it difficult to meet the demands for evaluating agent abilities. To address this problem, we propose the Trajectory-based Validated-by-Reproducing Agent-benchmark Complexity Evolution (TRACE) frame
Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
-
cs.AI, q-bio.NC updates on arXiv.org
-
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
arXiv:2603.08652v1 Announce Type: new Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning. However, existing CoT-based T2I methods largely rely on abstract natural-language planning, which lacks the precision required for complex spatial layouts, structured visual elements, and dense textual content. In this work, we propose CoCo (Code-as-CoT), a cod
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
arXiv:2509.25541v2 Announce Type: replace-cross Abstract: Although reinforcement learning (RL) has emerged as a promising approach for improving vision-language models (VLMs) and multimodal large language models (MLLMs), current methods rely heavily on manually curated datasets and costly human verification, which limits scalable self-improvement in multimodal systems. To address this challenge, we propose Vision-Zero, a label-free, domain-agnostic multi-agent self-play framework for self-evolv
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
-
cs.AI, q-bio.NC updates on arXiv.org
-
RubricBench: Aligning Model-Generated Rubrics with Human Standards
arXiv:2603.01562v2 Announce Type: replace Abstract: As Large Language Model (LLM) alignment evolves from simple completions to complex, highly sophisticated generation, Reward Models are increasingly shifting toward rubric-guided evaluation to mitigate surface-level biases. However, the community lacks a unified benchmark to assess this evaluation paradigm, as existing benchmarks lack both the discriminative complexity and the ground-truth rubric annotations required for rigorous analysis. To b
RubricBench: Aligning Model-Generated Rubrics with Human Standards
-
cs.AI, q-bio.NC updates on arXiv.org
-
TxRay: Agentic Postmortem of Live Blockchain Attacks
arXiv:2602.01317v5 Announce Type: replace-cross Abstract: Decentralized Finance (DeFi) has turned blockchains into financial infrastructure, allowing anyone to trade, lend, and build protocols without intermediaries, but this openness exposes pools of value controlled by code. Within five years, the DeFi ecosystem has lost over 15.75B USD to reported exploits. Many exploits arise from permissionless opportunities that any participant can trigger using only public state and standard interfaces,
TxRay: Agentic Postmortem of Live Blockchain Attacks
-
cs.AI, q-bio.NC updates on arXiv.org
-
Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
arXiv:2602.14299v1 Announce Type: cross Abstract: As large language model agents increasingly populate networked environments, a fundamental question arises: do artificial intelligence (AI) agent societies undergo convergence dynamics similar to human social systems? Lately, Moltbook approximates a plausible future scenario in which autonomous agents participate in an open-ended, continuously evolving online society. We present the first large-scale systemic diagnosis of this AI agent society.