❌

Reading view

SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning

arXiv:2609.12599v1 Announce Type: cross Abstract: Multi-task learning (MTL) requires navigating unavoidable trade-offs among competing objectives. This paradigm is frequently formulated as multi-objective optimization (MOO), where the scalarization is favored to reduce an MOO problem to a single objective. We empirically find that existing merit-function-based scalarization approaches are sensitive to the relative scales of different objectives in practical MTL, where task losses commonly differ by orders of magnitude. The optimization process often favors objectives with larger scales even though the underlying Pareto optimal solutions remains invariant to rescaling (i.e., multiplying an objective by a positive constant). To address this issue, we propose Scale-Invariant Merit-function-based Scalarization (SIMS) for MTL. Specifically, SIMS adopts a transformation-induced merit function to convert the MOO problem of MTL to a single objective that renders optimization invariant to the magnitudes of losses. Theoretically, we prove that the requirement for scale invariance uniquely determines this transformation to be logarithmic. We further show that this general transformation-induced merit function preserves weak Pareto optimality and admits a smooth surrogate with controllable approximation error. Extensive experiments on representative multi-task benchmarks demonstrate that SIMS consistently outperforms existing scalarization methods and achieves state-of-the-art performance.
  •  

Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval

arXiv:2510.20486v2 Announce Type: replace-cross Abstract: Imbalanced labels can cause frequent samples to dominate AI-based quantitative remote sensing, degrading rare-event retrieval. In rain-rate retrieval based on satellite infrared brightness temperatures, this imbalance leads to systematic underestimation of rare high-intensity rainfall. In this study, Hurdle-Retrieval Model Imbalanced Learning (RMIL) is proposed. Following a divide-and-conquer strategy, Hurdle-RMIL separates zero inflation from the long-tailed distribution of positive rain. A hurdle model handles zero inflation, whereas RMIL exploits invariance under fixed observation conditions of the rainfall-to-satellite forward process to derive a Bayes-based transformation linking conditional distributions under naturally long-tailed and hypothetical balanced rainfall. This transformation enables the balanced-distribution model to be learned from natural samples without constructing a balanced dataset. Comparisons with conventional learning, classification-regression modeling, cost-sensitive learning, and generative learning using test data from multiple regions in China show that Hurdle-RMIL mitigates systematic underestimation and improves detection of rare high-intensity and extreme rainfall without markedly degrading lower-threshold accuracy. At 0.1-10 mm per hour, its root mean square error remains close to those of the best baselines, and it yields the highest equitable threat score (ETS) at most evaluated thresholds, with its advantage becoming more pronounced at high thresholds. At 30 mm per hour, its ETS is 0.051 versus 0.015 for the best baseline, and its mean error is -25.41 mm per hour versus -28.98 mm per hour. Case studies further show improved representations of rainfall intensity and spatial extent, demonstrating that Hurdle-RMIL effectively addresses rainfall-distribution imbalance and improves the retrieval of rare high-intensity rainfall.
  •  

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

arXiv:2608.18234v3 Announce Type: replace-cross Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.
  •  

Lineage-specific pulmonary transcriptome landscape of coronavirus infection unveils universal immunotherapy for viral pneumonia

In the infection courses of different SARS-CoV-2 variants, disease outcomes and signatures were delineated by physiological changes, viral load, pathology, and pulmonary transcriptome analysis. This multi-dimensional landscape of disease outcomes and underlying mechanisms might provide important clues for immunotherapy of SARS-CoV-2 infection and pneumonia caused by other respiratory viruses.
  •  

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

arXiv:2410.20487v5 Announce Type: replace-cross Abstract: Experience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-dimensional state spaces. To address this limitation, we propose a novel approach, Efficient Diversity-based Experience Replay (EDER). EDER employs a determinantal point process to model the diversity between samples and prioritizes replay based on the diversity between samples. To further enhance learning efficiency, we incorporate Cholesky decomposition for handling large state spaces in realistic environments. Additionally, rejection sampling is applied to select samples with higher diversity, thereby improving overall learning efficacy. Extensive experiments are conducted on robotic manipulation tasks in MuJoCo, Atari games, and realistic indoor environments in Habitat. The results demonstrate that our approach not only significantly improves learning efficiency but also achieves superior performance in high-dimensional, realistic environments.
  •  

Targeting peripheral 5-HT2AR enhances antitumor immunity in colorectal cancer

By selectively targeting peripheral 5-HT2AR without inducing psychedelic effects, a non-brain-penetrant agonist boosts antitumor CD8+ T cell immunity and improves immunotherapy responses in preclinical models of colorectal cancer.
  •  

Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models

arXiv:2605.25681v1 Announce Type: cross Abstract: Designing a single molecule that modulates two targets is a promising strategy for polypharmacology, but it remains substantially harder than standard single-target generation because one candidate must satisfy two binding requirements while preserving drug-likeness and synthesizability. Existing dual-target generative methods typically introduce dual-target capability by either retraining the generator or intervening in the diffusion process during sampling. The former can be costly and difficult to stabilize when dual-target supervision is sparse, while the latter may be sensitive to denoising-time target balancing and competing update directions. These limitations motivate a generator-preserving alternative that keeps the pretrained prior intact: can dual-target candidates instead be recovered from the input space of a frozen single-target diffusion model, without modifying its parameters or denoising dynamics? We formulate this task as a constrained multi-objective optimization problem and propose REUSE, a hierarchical evolutionary input-space search framework that combines pair-conditioned exploration with structured multi-stage selection to enforce dual-target affinity, chemical quality, and diversity. Experiments show that, compared with methods that modify the diffusion process, REUSE consistently improves dual-target affinity and balance, achieving a 20.9-percentage-point gain in Dual High Affinity over the strongest prior baseline while maintaining competitive molecular quality.
  •  

DeepEN: A Deep Reinforcement Learning Framework for Personalized Enteral Nutrition in Critical Care

arXiv:2510.08350v3 Announce Type: replace-cross Abstract: Objective: Enteral nutrition (EN) delivery in the ICU remains suboptimal due to limited personalization and uncertainty regarding appropriate calorie, protein, and fluid targets under dynamic metabolic demands. We introduce DeepEN, a reinforcement learning (RL) framework for personalized EN optimization using electronic health record data. Methods: DeepEN was trained on over 11,000 ICU patients from MIMIC-IV to generate 4-hourly, patient-specific caloric, protein, and fluid targets. The state representation incorporated demographics, comorbidities, vital signs, laboratory values, and recent interventions. A physiologically aligned reward framework balanced biomarker stability with long-term survival. Policy learning employed a dueling double deep Q-network with Conservative Q-Learning regularization to enable safe offline training. Results: DeepEN achieved the highest estimated policy value ($V^\pi = 9.48$) and the lowest calibrated mortality (18.8 +/- 1.0%), representing a 4.0 percentage-point absolute reduction compared with clinician practice (22.8%). The policy also demonstrated superior metabolic stability, achieving the highest proportion of glucose, phosphate, and sodium values within target range. Furthermore, deviation from the DeepEN policy was independently associated with increased mortality and biomarker instability, whereas deviation from a random policy showed no such association. Interpretability analyses further indicated that recommendations were conditioned on physiologically relevant markers of organ function and metabolic status rather than static dosing heuristics. Conclusion: DeepEN demonstrates the feasibility of conservative offline RL for safe, individualized EN optimization, highlighting the potential of data-driven personalization to complement guideline-based approaches in critical care.
  •  

Characterization of dysbiosis patterns in gut microbiota of digestive system cancers: an umbrella review

Front Microbiol. 2026 Apr 28;17:1782471. doi: 10.3389/fmicb.2026.1782471. eCollection 2026.

ABSTRACT

Digestive system cancers (DSCs) represent a substantial global health burden. In recent years, the role of gut microbiota in the DSCs has garnered considerable attention, but its change pattern during tumor progression and the specific mechanisms are still not fully understood. We conducted a comprehensive systematic review to characterize patterns of gut microbiota dysbiosis across different DSC types and assess their clinical significance. We systematically searched four English and three Chinese databases up to January 2025 to identify systematic reviews focused on the dynamic characteristics of the gut microbiota during gastrointestinal tumorigenesis. Microbiota biodiversity and taxonomic composition were extracted to identify specific signatures associated with DSCs. The ROBIS tool was used to evaluate the methodological quality of the included studies. Ultimately, 59 studies involving six distinct DSC types were included. Data synthesis and comparison revealed distinct microbiota profiles across DSCs. At the phylum level, Bacillota was decreased in esophageal cancer (EC) and pancreatic ductal adenocarcinoma (PDAC), Pseudomonadota was augmented in EC but exhibited divergent trajectories in colorectal cancer (CRC) and PDAC. Genus-level analyses revealed Veillonella enrichment in EC and PDAC, and Fusobacterium outgrowth in EC, gastric cancer (GC) and CRC. Parvimonas and Streptococcus showed a concordant ascending trend in GC and CRC. Prevotella was overrepresented in EC and GC. This synthesis delineates a qualitative landscape of gut microbiota imbalances associated with various DSCs, highlighting the potential for these microbial shifts to serve as markers for early detection and targeted therapy. Multiomics integration and prospective cohort studies should be prioritized to accelerate clinical translation.

PMID:42131199 | PMC:PMC13161176 | DOI:10.3389/fmicb.2026.1782471

  •  

FDX1 as a predictive biomarker and therapeutic target for lymph node metastasis in gastric cancer

Clin Exp Med. 2026 May 10. doi: 10.1007/s10238-026-02160-0. Online ahead of print.

ABSTRACT

The prognostic values of cuproptosis-related genes (CRGs) in gastric cancer with lymph node metastasis (GCLM), especially in the tumor immune microenvironment (TIME), remain unclear. We analyzed the expression, mutation, immunity, drug sensitivity, and prognostic value of CRGs in GCLM using TCGA and GEO cohorts. Consensus clustering was performed to identify CRG subtypes, with differences characterized by multi-omics analysis. A CRG-based prognostic risk score and immune score were constructed for individualized assessment, and the role of CRGs was validated through in vitro and in vivo experiments. Consensus clustering revealed that CRGs were significantly enriched in biological processes related to mitosis and energy metabolism, as well as in immune-related and cancer-associated pathways. Four distinct CRG subtypes were identified, showing marked differences in expression profiles, prognosis, genetic alterations, TIME, and chemotherapeutic drug sensitivity. We developed an exploratory CRG-based prognostic risk score for preliminary individualized assessment, and the functional relevance of CRGs in GCLM was further validated through in vitro experiments. Among these, FDX1, LIAS, DLAT, MTF1, and GLS were identified as key determinants of overall survival in patients with GCLM, with FDX1 emerging as a potential independent prognostic factor. Notably, upregulation of FDX1 significantly suppressed lymph node metastasis of gastric cancer cells in a mouse popliteal lymph node metastasis model. Our data uncovers FDX1 might be a potential favorable prognostic factors in GCLM patients. These findings may improve our understanding of CRGs in GCLM and provide new in-sights for assessing prognosis and developing more effective treatment strategies.

PMID:42107026 | DOI:10.1007/s10238-026-02160-0

  •  

Stabilizing Rubric Integration Training via Decoupled Advantage Normalization

arXiv:2603.26535v2 Announce Type: replace Abstract: We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, treating all correct responses identically regardless of reasoning quality, and gradually lose the advantage signal as groups become uniformly correct. Process reward models (PRM) offer richer supervision, but directly using PRM scores causes reward hacking, where models exploit verbosity to inflate scores while accuracy collapses. PAPO resolves both by composing the advantage from an outcome component Aout, derived from ORM and normalized over all responses, and a process component Aproc, derived from a rubric-based PRM and normalized exclusively among correct responses. This decoupled design ensures that Aout anchors training on correctness while Aproc differentiates reasoning quality without distorting the outcome signal. Experiments across multiple model scales and six benchmarks demonstrate that PAPO consistently outperforms ORM, reaching 51.3% vs.\ 46.3% on OlympiadBench while continuing to improve as ORM plateaus and declines.
  •  

Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents

arXiv:2602.02050v3 Announce Type: replace Abstract: Tool-using agents based on Large Language Models (LLMs) excel in tasks such as mathematical reasoning and multi-hop question answering. However, in long trajectories, agents often trigger excessive and low-quality tool calls, increasing latency and degrading inference performance, making managing tool-use behavior challenging. In this work, we conduct entropy-based pilot experiments and observe a strong positive correlation between entropy reduction and high-quality tool calls. Building on this finding, we propose using entropy reduction as a supervisory signal and design two reward strategies to address the differing needs of optimizing tool-use behavior. Sparse outcome rewards provide coarse, trajectory-level guidance to improve efficiency, while dense process rewards offer fine-grained supervision to enhance performance. Experiments across diverse domains show that both reward designs improve tool-use behavior: the former reduces tool calls by 72.07% compared to the average of baselines, while the latter improves performance by 22.27%. These results position entropy reduction as a key mechanism for enhancing tool-use behavior, enabling agents to be more adaptive in real-world applications.
  •  

Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation

arXiv:2602.07023v2 Announce Type: replace-cross Abstract: Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. Investors are typically classified as fundamental or technical traders, but most simulations fix strategies at initialization, failing to reflect real-world trading dynamics. In this work, we assess whether agents' strategy switching aligns with financial theory, providing a framework for this evaluation. We operationalize four behavioral-finance drivers-loss aversion, herding, wealth differentiation, and price misalignment-as personality traits set via prompting and stored long-term. In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Our results show that recent LLMs' switching behavior is only partially consistent with behavioral-finance theories, highlighting the need for further refinement in aligning agent behavior with financial theory.
  •  

Fluid-Derived Organoids from Pleural Effusion and Ascites: Emerging Models for Drug Resistance and Personalized Oncology

J Cancer. 2026 Mar 4;17(3):614-625. doi: 10.7150/jca.127511. eCollection 2026.

ABSTRACT

Malignant pleural effusion (MPE) and malignant ascites (MA) are common complications in advanced-stage cancers, often signifying disease progression and resistance to treatment. Compared to tissue biopsies or surgical specimens, materials derived from effusions offer advantages such as minimal invasiveness, ease of accessibility, and the feasibility of repeated collection during therapeutic interventions. Organoids generated from tumor cells in effusions, termed fluid-derived organoids (FDOs), have demonstrated the ability to maintain genetic heterogeneity and accurately replicate patient-specific tumor phenotypes. These characteristics position FDOs as promising models for investigating drug resistance mechanisms and informing personalized oncology strategies. In the context of lung cancer, organoids derived from pleural effusions have been employed to study acquired resistance to epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors and immunotherapy. Similarly, in ovarian and gastrointestinal cancers, organoids derived from ascites have proven to be valuable platforms for examining chemotherapy resistance and conducting drug sensitivity testing. FDOs have shown significant potential for translational applications by effectively correlating ex vivo drug responses with clinical outcomes, thus facilitating real-time monitoring of resistance evolution. However, several challenges remain, such as achieving culture standardization, maintaining the integrity of tumor microenvironment components, and integrating with multi-omics approaches. This review provides a comprehensive overview of recent advancements in the use of pleural effusion- and ascites-derived organoids for drug resistance research, underscores their applications in personalized oncology, and explores future research directions.

PMID:41869438 | PMC:PMC13003542 | DOI:10.7150/jca.127511

  •  

Fluid-Derived Organoids from Pleural Effusion and Ascites: Emerging Models for Drug Resistance and Personalized Oncology

J Cancer. 2026 Mar 4;17(3):614-625. doi: 10.7150/jca.127511. eCollection 2026.

ABSTRACT

Malignant pleural effusion (MPE) and malignant ascites (MA) are common complications in advanced-stage cancers, often signifying disease progression and resistance to treatment. Compared to tissue biopsies or surgical specimens, materials derived from effusions offer advantages such as minimal invasiveness, ease of accessibility, and the feasibility of repeated collection during therapeutic interventions. Organoids generated from tumor cells in effusions, termed fluid-derived organoids (FDOs), have demonstrated the ability to maintain genetic heterogeneity and accurately replicate patient-specific tumor phenotypes. These characteristics position FDOs as promising models for investigating drug resistance mechanisms and informing personalized oncology strategies. In the context of lung cancer, organoids derived from pleural effusions have been employed to study acquired resistance to epidermal growth factor receptor (EGFR) tyrosine kinase inhibitors and immunotherapy. Similarly, in ovarian and gastrointestinal cancers, organoids derived from ascites have proven to be valuable platforms for examining chemotherapy resistance and conducting drug sensitivity testing. FDOs have shown significant potential for translational applications by effectively correlating ex vivo drug responses with clinical outcomes, thus facilitating real-time monitoring of resistance evolution. However, several challenges remain, such as achieving culture standardization, maintaining the integrity of tumor microenvironment components, and integrating with multi-omics approaches. This review provides a comprehensive overview of recent advancements in the use of pleural effusion- and ascites-derived organoids for drug resistance research, underscores their applications in personalized oncology, and explores future research directions.

PMID:41869438 | PMC:PMC13003542 | DOI:10.7150/jca.127511

  •  

Continual Learning in Large Language Models: Methods, Challenges, and Opportunities

arXiv:2603.12658v1 Announce Type: cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting-a critical limitation of the static pre-training paradigm inherent to modern LLMs. This survey presents a comprehensive overview of CL methodologies tailored for LLMs, structured around three core training stages: continual pre-training, continual fine-tuning, and continual alignment.Beyond the canonical taxonomy of rehearsal-, regularization-, and architecture-based methods, we further subdivide each category by its distinct forgetting mitigation mechanisms and conduct a rigorous comparative analysis of the adaptability and critical improvements of traditional CL methods for LLMs. In doing so, we explicitly highlight core distinctions between LLM CL and traditional machine learning, particularly with respect to scale, parameter efficiency, and emergent capabilities. Our analysis covers essential evaluation metrics, including forgetting rates and knowledge transfer efficiency, along with emerging benchmarks for assessing CL performance. This survey reveals that while current methods demonstrate promising results in specific domains, fundamental challenges persist in achieving seamless knowledge integration across diverse tasks and temporal scales. This systematic review contributes to the growing body of knowledge on LLM adaptation, providing researchers and practitioners with a structured framework for understanding current achievements and future opportunities in lifelong learning for language models.
  •  

Towards Effective and Efficient Graph Alignment without Supervision

arXiv:2603.08526v1 Announce Type: cross Abstract: Unsupervised graph alignment aims to find the node correspondence across different graphs without any anchor node pairs. Despite the recent efforts utilizing deep learning-based techniques, such as the embedding and optimal transport (OT)-based approaches, we observe their limitations in terms of model accuracy-efficiency tradeoff. By focusing on the exploitation of local and global graph information, we formalize them as the ``local representation, global alignment'' paradigm, and present a new ``global representation and alignment'' paradigm to resolve the mismatch between the two phases in the alignment process. We then propose \underline{Gl}obal representation and \underline{o}ptimal transport-\underline{b}ased \underline{Align}ment (\texttt{GlobAlign}), and its variant, \texttt{GlobAlign-E}, for better \underline{E}fficiency. Our methods are equipped with the global attention mechanism and a hierarchical cross-graph transport cost, able to capture long-range and implicit node dependencies beyond the local graph structure. Furthermore, \texttt{GlobAlign-E} successfully closes the time complexity gap between representative embedding and OT-based methods, reducing OT's cubic complexity to quadratic terms. Through extensive experiments, our methods demonstrate superior performance, with up to a 20\% accuracy improvement over the best competitor. Meanwhile, \texttt{GlobAlign-E} achieves the best efficiency, with an order of magnitude speedup against existing OT-based methods.
  •  

Process-Centric Analysis of Agentic Software Systems

arXiv:2512.02393v2 Announce Type: replace-cross Abstract: Agentic systems are modern software systems: they consist of orchestrated modules, expose interfaces, and are deployed in software pipelines. Unlike conventional programs, their execution, i.e., trajectories, is inherently stochastic and adaptive to the problems they solve. Evaluation of such systems is often outcome-centric. This narrow focus overlooks detailed insights, failing to explain how agents reason, plan, act, or change their strategies. Inspired by the structured representation of conventional software systems as graphs, we introduce Graphectory to systematically encode the temporal and semantic relations in such systems. Using Graphectory, we automatically analyze 4000 trajectories of two dominant agentic programming workflows, SWE-agent and OpenHands, with four backbone Large Language Models (LLMs), attempting to resolve SWE-bench issues. Our automated analyses (completed within four minutes) reveal that: (1) agents using richer prompts or stronger LLMs exhibit more complex Graphectory, reflecting deeper exploration, broader context gathering, and more thorough validation; (2) agents' strategies vary with problem difficulty and the underlying LLM - for resolved issues, strategies often follow coherent localization-patching-validation steps, while unresolved ones exhibit chaotic or backtracking behaviors; and (3) even successful agentic systems often display inefficient processes. We also implement a novel technique for real-time construction and analysis of Graphectory and Langutory during agent execution to flag trajectory issues. Upon detecting such issues, the technique notifies the agent with a diagnostic message and, when applicable, rolls back the trajectory. Experiments show that online monitoring and interventions improve resolution rates by 6.9%-23.5% across models for problematic instances, while significantly shortening trajectories with near-zero overhead.
  •  

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

arXiv:2603.03781v1 Announce Type: new Abstract: Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory benchmarks primarily target declarative memory, specifically semantic and episodic types, where all information is explicitly presented in dialogues. In contrast, real-world actions are also governed by non-declarative memory, including habitual and procedural types, and need to be inferred from diverse digital traces. To bridge this gap, we introduce Lifebench, which features densely connected, long-horizon event simulation. It pushes AI agents beyond simple recall, requiring the integration of declarative and non-declarative memory reasoning across diverse and temporally extended contexts. Building such a benchmark presents two key challenges: ensuring data quality and scalability. We maintain data quality by employing real-world priors, including anonymized social surveys, map APIs, and holiday-integrated calendars, thus enforcing fidelity, diversity and behavioral rationality within the dataset. Towards scalability, we draw inspiration from cognitive science and structure events according to their partonomic hierarchy; enabling efficient parallel generation while maintaining global coherence. Performance results show that top-tier, state-of-the-art memory systems reach just 55.2\% accuracy, highlighting the inherent difficulty of long-horizon retrieval and multi-source integration within our proposed benchmark. The dataset and data synthesis code are available at https://github.com/1754955896/LifeBench.
  •  
❌