❌

Normal view

UniCA: Unified Covariate Adaptation for Time Series Foundation Model

arXiv:2506.22039v2 Announce Type: replace-cross Abstract: Time Series Foundation Models (TSFMs) have achieved remarkable success through large-scale pretraining. However, their design primarily targets real-valued series, limiting their ability to handle general forecasting tasks involving diverse and often heterogeneous covariates -- such as categorical variables and multimodal data (e.g., images, text) -- which are typically task-specific and difficult to leverage during pretraining. To address this gap, we propose Unified Covariate Adaptation (UniCA), a framework to bridge TSFMs with general covariate-aware forecasting. UniCA first performs covariate homogenization to transform heterogeneous covariates into high-level homogeneous series representations and then fuses them via a unified attention-based fusion mechanism. UniCA is compatible and universal for adaptation with both homogeneous and heterogeneous covariates, incorporating extra covariate information while preserving the generalization ability of TSFMs.Extensive experiments on multiple unimodal and multimodal covariate-aware forecasting benchmarks demonstrate the superiority of UniCA, highlighting the promise of covariate-aware TSFM adaptation in real-world forecasting scenarios.Code: https://github.com/hanlu-nju/UniCA.

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

Towards AI Search Paradigm

arXiv:2506.17188v2 Announce Type: replace-cross Abstract: In this paper, we introduce the AI Search Paradigm, a comprehensive blueprint for next-generation search systems capable of emulating human information processing and decision-making. The paradigm employs a modular architecture of four LLM-powered agents (Master, Planner, Executor and Writer) that dynamically adapt to the full spectrum of information needs, from simple factual queries to complex multi-stage reasoning tasks. These agents collaborate dynamically through coordinated workflows to evaluate query complexity, decompose problems into executable plans, and orchestrate tool usage, task execution, and content synthesis. We systematically present key methodologies for realizing this paradigm, including task planning and tool integration, execution strategies, aligned and robust retrieval-augmented generation, and efficient LLM inference, spanning both algorithmic techniques and infrastructure-level optimizations. By providing an in-depth guide to these foundational components, this work aims to inform the development of trustworthy, adaptive, and scalable AI search systems.

Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills

arXiv:2512.16301v3 Announce Type: replace Abstract: Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.

FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels

arXiv:2511.02872v4 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have demonstrated impressive capabilities in formal theorem proving, particularly on contest-based mathematical benchmarks like the IMO. However, these contests do not reflect the depth, breadth, and abstraction of modern mathematical research. To bridge this gap, we introduce FATE (Formal Algebra Theorem Evaluation), a new benchmark series in formal algebra designed to chart a course toward advanced mathematical reasoning. We present two new components, FATE-H and FATE-X, each with 100 problems in abstract and commutative algebra. The FATE series spans a difficulty spectrum from undergraduate exercises to problems exceeding PhD qualifying exams. Notably, FATE-X is the first formal benchmark to surpass both PhD-level exam difficulty and the coverage of the Mathlib library. Our evaluations of state-of-the-art LLM provers on this new benchmark reveal a stark performance gap compared to contest math: the best model achieves only 3% (pass@64) accuracy on FATE-H and 0% on FATE-X. Our two-stage evaluation reveals that models' natural-language reasoning is notably more accurate than their ability to formalize this reasoning. We systematically classify the common errors that arise during this formalization process. Furthermore, a comparative study shows that a specialized prover can exhibit less effective reflection than general-purpose models, reducing its accuracy at the natural-language stage. We believe FATE provides a robust and challenging benchmark that establishes essential checkpoints on the path toward research-level formal mathematical reasoning.

VANGUARD: Vehicle-Anchored Ground Sample Distance Estimation for UAVs in GPS-Denied Environments

arXiv:2603.04277v1 Announce Type: cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene. As LLM/VLM-based planners are increasingly adopted as high-level agents for embodied systems, their ability to reason about physical dimensions becomes safety-critical -- yet our experiments show that five state-of-the-art VLMs suffer from spatial scale hallucinations, with median area estimation errors exceeding 50%. We propose VANGUARD, a lightweight, deterministic Geometric Perception Skill designed as a callable tool that any LLM-based agent can invoke to recover Ground Sample Distance (GSD) from ubiquitous environmental anchors: small vehicles detected via oriented bounding boxes, whose modal pixel length is robustly estimated through kernel density estimation and converted to GSD using a pre-calibrated reference length. The tool returns both a GSD estimate and a composite confidence score, enabling the calling agent to autonomously decide whether to trust the measurement or fall back to alternative strategies. On the DOTA~v1.5 benchmark, VANGUARD achieves 6.87% median GSD error on 306~images. Integrated with SAM-based segmentation for downstream area measurement, the pipeline yields 19.7% median error on a 100-entry benchmark -- with 2.6x lower category dependence and 4x fewer catastrophic failures than the best VLM baseline -- demonstrating that equipping agents with deterministic geometric tools is essential for safe autonomous spatial reasoning.

From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs

arXiv:2603.03112v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are fundamental for modeling complex physical systems, yet classical numerical solvers face prohibitive computational costs in high-dimensional and multi-scale regimes. While Transformer-based neural operators have emerged as powerful data-driven alternatives, they conventionally treat all discretized spatial points as uniform, independent tokens. This monolithic approach ignores the intrinsic scale separation of physical fields, applying computationally prohibitive global attention that redundantly mixes smooth large-scale dynamics with high-frequency fluctuations. Rethinking Transformers through the lens of complex dynamics, we propose DynFormer, a novel dynamics-informed neural operator. Rather than applying a uniform attention mechanism across all scales, DynFormer explicitly assigns specialized network modules to distinct physical scales. It leverages a Spectral Embedding to isolate low-frequency modes, enabling a Kronecker-structured attention mechanism to efficiently capture large-scale global interactions with reduced complexity. Concurrently, we introduce a Local-Global-Mixing transformation. This module utilizes nonlinear multiplicative frequency mixing to implicitly reconstruct the small-scale, fast-varying turbulent cascades that are slaved to the macroscopic state, without incurring the cost of global attention. Integrating these modules into a hybrid evolutionary architecture ensures robust long-term temporal stability. Extensive memory-aligned evaluations across four PDE benchmarks demonstrate that DynFormer achieves up to a 95% reduction in relative error compared to state-of-the-art baselines, while significantly reducing GPU memory consumption. Our results establish that embedding first-principles physical dynamics into Transformer architectures yields a highly scalable, theoretically grounded blueprint for PDE surrogate modeling.

Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy

arXiv:2510.08646v2 Announce Type: replace-cross Abstract: Safety alignment of large language models currently faces a central challenge: existing alignment techniques often prioritize mitigating responses to harmful prompts at the expense of overcautious behavior, leading models to incorrectly refuse benign requests. A key goal of safe alignment is therefore to improve safety while simultaneously minimizing false refusals. In this work, we introduce Energy Landscape Steering (ELS), a novel, fine-tuning free framework designed to resolve this challenge through dynamic, inference-time intervention. We train a lightweight external Energy-Based Model (EBM) to assign high energy to undesirable states (false refusal or jailbreak) and low energy to desirable states (helpful response or safe reject). During inference, the EBM maps the LLM's internal activations to an energy landscape, and we use the gradient of the energy function to steer the hidden states toward low-energy regions in real time. This dynamically guides the model toward desirable behavior without modifying its parameters. By decoupling behavioral control from the model's core knowledge, ELS provides a flexible and computationally efficient solution. Extensive experiments across diverse models demonstrate its effectiveness, raising compliance on the ORB-H benchmark from 57.3 percent to 82.6 percent while maintaining baseline safety performance. Our work establishes a promising paradigm for building LLMs that simultaneously achieve high safety and low false refusal rates.

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

arXiv:2602.13954v1 Announce Type: cross Abstract: We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

arXiv:2510.10689v2 Announce Type: replace Abstract: Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.

An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

arXiv:2506.20430v3 Announce Type: replace-cross Abstract: Rare diseases affect over 300 million individuals worldwide, yet timely and accurate diagnosis remains an urgent challenge. Patients often endure a prolonged diagnostic odyssey exceeding five years, marked by repeated referrals, misdiagnoses, and unnecessary interventions, leading to delayed treatment and substantial emotional and economic burdens. Here we present DeepRare, a multi-agent system for rare disease differential diagnosis decision support powered by large language models, integrating over 40 specialized tools and up-to-date knowledge sources. DeepRare processes heterogeneous clinical inputs, including free-text descriptions, structured Human Phenotype Ontology terms, and genetic testing results, to generate ranked diagnostic hypotheses with transparent reasoning linked to verifiable medical evidence. Evaluated across nine datasets from literature, case reports and clinical centres across Asia, North America and Europe spanning 14 medical specialties, DeepRare demonstrates exceptional performance on 3,134 diseases. In human-phenotype-ontology-based tasks, it achieves an average Recall@1 of 57.18%, outperforming the next-best method by 23.79%; in multi-modal tests, it reaches 69.1% compared with Exomiser's 55.9% on 168 cases. Expert review achieved 95.4% agreement on its reasoning chains, confirming their validity and traceability. Our work not only advances rare disease diagnosis but also demonstrates how the latest powerful large-language-model-driven agentic systems can reshape current clinical workflows.
❌