❌

Reading view

BAF60A governs beta cell identity to control systemic glucose homeostasis

Diabetologia. 2026 Oct 3. doi: 10.1007/s00125-026-06884-2. Online ahead of print.

ABSTRACT

AIMS/HYPOTHESIS: Chromatin remodelling is critical for maintaining pancreatic beta cell identity and function, yet the key regulatory mechanisms remain incompletely defined. This study aimed to investigate the role of the switch/sucrose non-fermentable (SWI/SNF) complex subunit BAF60A in preserving beta cell fate and glucose homeostasis.

METHODS: Pdx1-Cre-mediated BAF60A-knockout (BaBKO) and BAF60A-overexpressing (BaBOE) mice, together with tamoxifen-inducible adult beta cell-specific Smarcd1 knockout (BaBKOTM) and Isl1 knockout (Isl1BKOTM) mice, were generated to evaluate the role of BAF60A in vivo. Glucose homeostasis was assessed through glucose tolerance tests, insulin tolerance tests and glucose-stimulated insulin secretion (GSIS) assays. Multiomic analyses, including RNA-seq, ATAC-seq, Cleavage Under Targets and Tagmentation (CUT&Tag) and single-cell RNA-seq, were performed to characterise chromatin accessibility and transcriptional changes. BAF60A-interacting proteins were identified with biotin identification (BioID) and GST pull-down assays. Beta cell lineage tracing was used to assess changes in cell identity. In addition, BAF60A and the dedifferentiation marker ALDH1A3 were examined in pancreatic islets from individuals with and without type 2 diabetes.

RESULTS: BaBKO mice exhibited significant glucose intolerance, impaired GSIS and pronounced loss of beta cell identity, accompanied by the acquisition of non-beta endocrine features. Inducible deletion of Smarcd1 in adult beta cells similarly impaired beta cell maturation and promoted dedifferentiation, as confirmed by lineage tracing. BAF60A deficiency reduced enhancer accessibility and downregulated beta cell identity genes. Mechanistically, BAF60A physically interacts with the transcription factor islet-1 (ISL1) to regulate transcription of target genes. Adult beta cell-specific Isl1 deletion recapitulated key features of BAF60A deficiency and abolished the beneficial effect of BAF60A overexpression on insulin secretion. Conversely, BaBOE mice exhibited improved glucose tolerance and enhanced GSIS under high-fat diet conditions. Adeno-associated virus-mediated BAF60A overexpression markedly reduced beta cell dedifferentiation in BKS-db/db mice. In human type 2 diabetes islets, BAF60A expression was significantly reduced and inversely correlated with ALDH1A3.

CONCLUSIONS/INTERPRETATION: This work establishes BAF60A-ISL1-dependent chromatin remodelling as a key mechanism that preserves beta cell identity and function under metabolic stress, providing mechanistic insight into beta cell failure in type 2 diabetes.

PMID:42829354 | DOI:10.1007/s00125-026-06884-2

  •  

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

arXiv:2608.26105v2 Announce Type: replace-cross Abstract: Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.
  •  

Differential Responses of COL17A1 Nonsense Mutations to Readthrough Drugs and NOG as a Novel Enhancer in Junctional Epidermolysis Bullosa

COL17A1 nonsense mutations generate premature termination codons, reducing collagen XVII through truncated protein production and/or nonsense-mediated decay (NMD), and causing junctional epidermolysis bullosa. Translational readthrough drug offers a potential approach to nonsense mutations suppression. Different COL17A1 nonsense mutations showed distinct responses to readthrough drugs and NMD inhibitors, supporting personalized therapeutic strategies.
  •  

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

arXiv:2609.04355v2 Announce Type: replace-cross Abstract: Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9$\times$ improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2$\times$ and 1.8$\times$ the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.
  •  

Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study

Background: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. codes are inconsistent and do not specify nonprescribed fentanyl use. Objective: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. Methods: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms (“fent,” “blues,” “M30s,” and “tranq”) in their EHR documentation. Physician annotators classified snippets into “nonprescribed fentanyl use,” “prescribed fentanyl use,” or “other,” with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and -scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. Results: Of 7389 snippets, 9.6% (n=709) were classified as “nonprescribed fentanyl use,” 40.3% (n=2981) were classified as “prescribed fentanyl use,” and 50% (n=3699) were classified as “other.” Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest -score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. Conclusions: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.
  •  

PREDICT-GBM: A multicenter platform advancing personalized glioblastoma radiotherapy planning

npj Digital Medicine, Published online: 09 September 2026; doi:10.1038/s41746-026-03194-0

PREDICT-GBM: A multicenter platform advancing personalized glioblastoma radiotherapy planning
  •  

Ensuring multiomics data reproducibility for artificial intelligence with reference materials as a common calibrator

Nature Biotechnology, Published online: 01 September 2026; doi:10.1038/s41587-026-03267-1

Reference materials should be adopted as a common calibrator for multiomics measurement and co‑profiled with study samples. Multiomics results should be reported as sample‑to‑reference ratios so that they are reproducible and suitable for artificial intelligence tools.
  •  

Your Embedding Model is SMARTer Than You Think

arXiv:2605.24938v1 Announce Type: cross Abstract: Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine-grained, local evidence critical for dense retrieval tasks. Multi-vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi-vector capabilities of standard single-vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late-interaction over these frozen hidden states during inference, SMART acts as a plug-and-play upgrade that consistently improves performance across diverse modalities, improving even the state-of-the-art models further on MMEB-V2. We also reveal SMART's superior performance, as simple lightweight post-training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single-vector model to outperform SoTA multi-vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval. We open source our code and weights at https://github.com/HanSolo9682/SMART.
  •  

NPSolver: Neural Poisson Solver with Iterative Physics Supervision

arXiv:2605.25786v1 Announce Type: cross Abstract: Efficiently solving Poisson equations on complex, irregular domains remains a fundamental challenge in scientific computing, as classical iterative solvers often suffer from prohibitive runtime due to ill-conditioned systems. While neural operators offer a fast alternative, they typically rely on large-scale labeled datasets or struggle with unstable training dynamics when using physics-informed residual losses. We propose \textsc{NPSolver}, a neural Poisson solver trained without solution labels via iterative physics supervision. Instead of relying on fully converged numerical solutions or raw PDE residuals, \textsc{NPSolver} utilizes a small number of preconditioned conjugate gradient (PCG) steps to refine its own predictions, providing a more stable and well-scaled training signal. Theoretical analysis confirms that this iterative supervision serves as a well-conditioned error proxy and that a stop-gradient design is essential for optimization stability. To better capture boundary-driven features under mixed boundary conditions, we further introduce the Boundary-Aware Transolver (\textsc{BA-Transolver}) architecture that explicitly separates interior and boundary tokenization. Extensive evaluations on 2D and 3D irregular geometries demonstrate that \textsc{NPSolver} outperforms both physics-informed and data-driven baselines. Furthermore, a downstream thermal control task highlights the model's capability for conducting efficient and reliable gradient-based boundary control. We will release our codes and data at https://github.com/intell-sci-comput/NPSolver.
  •  

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

arXiv:2605.25850v1 Announce Type: cross Abstract: This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large language models. This paper extends that idea by moving from a ternary reward to a Trajectory-Informed advantage reweighting, dynamically re-weights the abstention reward during Group Relative Policy Optimization (GRPO) training. The objective of this work focuses on abstention learning instead of improving truthfulness, serving as an exploration into hallucination reduction. The novelty of this paper lies in methodological innovation, advantage re-weighting, and benchmark selection. Leveraging GRPO's multiple trajectories as a natural abstention signal, this method uses a reward signal to explore knowledge boundaries and encourage consistency. By demonstrating that trajectories can be used as a confidence indicator of the policy relative to the query, they are then used to dynamically calculate the abstention advantage. AbstentionBench is used as the evaluation benchmark, as this work aims to contribute to the field of abstention learning. All datasets on the benchmark were tested against this method and various baselines. Empirical results demonstrate that TIAR achieves state-of-the-art abstention F1 scores across five of six evaluation categories, outperforming the static ternary baseline on 17 of 31 benchmark datasets while fully preserving baseline accuracy.
  •  
  •  

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

arXiv:2601.10611v4 Announce Type: replace-cross Abstract: Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from them, or do not disclose their training data or recipe. As a result, the open-source community lacks the foundations needed to improve on the state-of-the-art video (and image) language models. Crucially, many downstream applications require more than just high-level video understanding; they require grounding -- either by pointing or by tracking in pixels. Even proprietary models lack this capability. We present Molmo2, a new family of VLMs that are state-of-the-art among open-source models and demonstrate exceptional new capabilities in point-driven grounding in single image, multi-image, and video tasks. Our key contribution is a collection of 7 new video datasets and 2 multi-image datasets, including a dataset of highly detailed video captions for pre-training, a free-form video Q&A dataset for fine-tuning, a new object tracking dataset with complex queries, and an innovative new video pointing dataset, all collected without the use of closed VLMs. We also present a training recipe for this data utilizing an efficient packing and message-tree encoding scheme, and show bi-directional attention on vision tokens and a novel token-weight strategy improves performance. Our best-in-class 8B model outperforms others in the class of open weight and data models on short videos, counting, and captioning, and is competitive on long-videos. On video-grounding Molmo2 significantly outperforms existing open-weight models like Qwen3-VL (35.5 vs 29.6 accuracy on video counting) and surpasses proprietary models like Gemini 3 Pro on some tasks (38.4 vs 20.0 F1 on video pointing and 56.2 vs 41.1 J&F on video tracking).
  •  

FigAgent: Towards Automatic Method Illustration Figure Generation for AI Scientific Papers

arXiv:2603.29590v1 Announce Type: cross Abstract: Method illustration figures (MIFs) play a crucial role in conveying the core ideas of scientific papers, yet their generation remains a labor-intensive process. In this paper, we identify three key characteristics that substantially influence MIF generation quality, i.e., \emph{compositional complexity}, \emph{component similarity}, and \emph{design dynamics}. To handle these characteristics, we take inspiration from human authors' drawing practices and propose \textbf{FigAgent}, a novel multi-agent framework for automatically generating high-quality MIFs. Through multi-agent collaboration, our FigAgent distills drawing experiences across similar components of MIFs and encapsulates them into reusable tools that can be invoked during MIF generation, while evolving these tools to adapt to dynamic design requirements. Besides, a novel Explore-and-Select drawing strategy is introduced to mimic the human-like trial-and-error manner for gradually constructing MIFs with complex structures. Extensive experiments show the efficacy of our method. Project is available \href{https://zhuolingli.github.io/FigAgent-page-project/}{here}.
  •  

UniFluids: Unified Neural Operator Learning with Conditional Flow-matching

arXiv:2603.22309v1 Announce Type: cross Abstract: Partial differential equation (PDE) simulation holds extensive significance in scientific research. Currently, the integration of deep neural networks to learn solution operators of PDEs has introduced great potential. In this paper, we present UniFluids, a conditional flow-matching framework that harnesses the scalability of diffusion Transformer to unify learning of solution operators across diverse PDEs with varying dimensionality and physical variables. Unlike the autoregressive PDE foundation models, UniFluids adopts flow-matching to achieve parallel sequence generation, making it the first such approach for unified operator learning. Specifically, the introduction of a unified four-dimensional spatiotemporal representation for the heterogeneous PDE datasets enables joint training and conditional encoding. Furthermore, we find the effective dimension of the PDE dataset is much lower than its patch dimension. We thus employ $x$-prediction in the flow-matching operator learning, which is verified to significantly improve prediction accuracy. We conduct a large-scale evaluation of UniFluids on several PDE datasets covering spatial dimensions 1D, 2D and 3D. Experimental results show that UniFluids achieves strong prediction accuracy and demonstrates good scalability and cross-scenario generalization capability. The code will be released later.
  •  

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

arXiv:2603.21289v2 Announce Type: replace-cross Abstract: Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are costly and difficult to scale. To address this, we propose an unsupervised self-evolution training framework for multimodal reasoning that achieves stable performance improvements without using human-annotated answers or external reward models. For each input, we sample multiple reasoning trajectories and jointly model their within group structure. We use the Actor's self-consistency signal as a training prior, and introduce a bounded Judge based modulation to continuously reweight trajectories of different quality. We further model the modulated scores as a group level distribution and convert absolute scores into relative advantages within each group, enabling more robust policy updates. Trained with Group Relative Policy Optimization (GRPO) on unlabeled data, our method consistently improves reasoning performance and generalization on five mathematical reasoning benchmarks, offering a scalable path toward self-evolving multimodal models. The code are available at https://github.com/OPPO-Mente-Lab/LLM-Self-Judge.
  •  

Inactivating <i>SnRK1β1A</i> promotes broad-spectrum disease resistance in rice

Nature, Published online: 25 March 2026; doi:10.1038/s41586-026-10273-5

SnRK1β1A in rice promotes susceptibility to multiple fungal diseases, and disrupting this infection-inducible gene confers broad-spectrum resistance without compromising growth or yield under normal field conditions.
  •  

Towards unified brain-to-text decoding across speech production and perception

arXiv:2603.12628v1 Announce Type: new Abstract: Speech production and perception are the main ways humans communicate daily. Prior brain-to-text decoding studies have largely focused on a single modality and alphabetic languages. Here, we present a unified brain-to-sentence decoding framework for both speech production and perception in Mandarin Chinese. The framework exhibits strong generalization ability, enabling sentence-level decoding when trained only on single-character data and supporting characters and syllables unseen during training. In addition, it allows direct and controlled comparison of neural dynamics across modalities. Mandarin speech is decoded by first classifying syllable components in Hanyu Pinyin, namely initials and finals, from neural signals, followed by a post-trained large language model (LLM) that maps sequences of toneless Pinyin syllables to Chinese sentences. To enhance LLM decoding, we designed a three-stage post-training and two-stage inference framework based on a 7-billion-parameter LLM, achieving overall performance that exceeds larger commercial LLMs with hundreds of billions of parameters or more. In addition, several characteristics were observed in Mandarin speech production and perception: speech production involved neural responses across broader cortical regions than auditory perception; channels responsive to both modalities exhibited similar activity patterns, with speech perception showing a temporal delay relative to production; and decoding performance was broadly comparable across hemispheres. Our work not only establishes the feasibility of a unified decoding framework but also provides insights into the neural characteristics of Mandarin speech production and perception. These advances contribute to brain-to-text decoding in logosyllabic languages and pave the way toward neural language decoding systems supporting multiple modalities.
  •  

The Effects of Digital Health Interventions on Motor Symptoms, Nonmotor Symptoms, and Quality of Life in Patients With Parkinson Disease: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Parkinson disease (PD) is a progressive neurodegenerative disorder with increasing global prevalence, necessitating innovative management. Digital health interventions (DHIs) offer potential advantages for PD care; yet, a comprehensive systematic review and synthesis across all DHI types and core outcomes is still lacking. Objective: This review aimed to assess the effectiveness of DHIs for improving motor symptoms, nonmotor symptoms, and quality of life in patients with PD and to summarize the reach, uptake, and feasibility. Methods: We searched PubMed, Ovid Embase, Web of Science, CINAHL, Cochrane Central Register of Controlled Trials, and APA PsycINFO up to November 2025. Pooled standardized mean differences (SMDs) were calculated using random-effects models. We calculated 95% prediction intervals (PIs) to estimate the true effects. The revised Cochrane Risk of Bias 2 tool was used to assess risk of bias. Heterogeneity was assessed using I2, τ2, and 95% PI. Subgroup analyses, meta-regression, and sensitivity analyses were conducted to address heterogeneity and potential bias. The quality of evidence was assessed using GRADE (Grading of Recommendations Assessment, Development, and Evaluation). Results: The review included 112 randomized controlled trials involving 5594 participants. Significant postintervention improvements were identified in motor symptoms (SMD=–0.39, 95% CI –0.60 to –0.18, 95% PI –1.75 to 0.99; I2=80.3%) and overall nonmotor symptoms (SMD=–0.26, 95% CI –0.49 to –0.03, 95% PI –0.56 to 0.03; I2=13.8%), including cognitive function (SMD=0.47, 95% CI 0.22 to 0.72, 95% PI –0.41 to 1.35; I2=63.5%) and psychiatric symptoms (SMD=–0.42, 95% CI –0.74 to –0.09, 95% PI –1.82 to –0.99; I2=85.4%); however, there was no significant enhancement in quality of life (SMD=–0.19, 95% CI –0.47 to 0.09, 95% PI –1.50 to 1.12; I2=81.2%). The certainty of evidence was very low for quality of life, motor, and psychiatric symptoms and low for cognitive function and overall nonmotor symptoms. Improvements in motor symptoms and cognitive function remained stable at follow-up. Meta-regression analysis indicated that age, percentage of female participants, and supervision mode were possible sources of heterogeneity. Overall, 94 studies reported reach (median 37.5%), 38 reported fidelity (95.7%), and 105 reported dropout rates (9.1%). Conclusions: In contrast to previous reviews focused on single technologies or outcomes, this review provided the first comprehensive synthesis across all DHI types on multiple outcomes and indicated their potential as nonpharmacological interventions for PD management. However, current evidence is of low to very low certainty, and wide 95% PIs, together with high risk of bias and substantial heterogeneity, indicate considerable uncertainty regarding the true effect in future implementations. Therefore, findings should be interpreted with caution. These findings provide integrated evidence to guide the design and prioritization of future research. The results have important real-world implications, supporting cautious implementation while underscoring the need for more robust trials, particularly in resource-limited settings. Trial Registration: PROSPERO CRD42023492123; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023492123
  •  

Automated Reinforcement Learning: An Overview

arXiv:2201.05000v2 Announce Type: replace-cross Abstract: Reinforcement Learning and, recently, Deep Reinforcement Learning are popular methods for solving sequential decision-making problems modeled as Markov Decision Processes. RL modeling of a problem and selecting algorithms and hyper-parameters require careful consideration, as different configurations may entail completely different performances. These considerations are mainly the task of RL experts; however, RL is progressively becoming popular in other fields, such as combinatorial optimization, where researchers and system designers are not necessarily RL experts. Besides, many modeling decisions are typically made manually, such as defining state and action space, size of batches, batch update frequency, and time steps. For these reasons, automating different components of RL is of great importance, and it has attracted much attention in recent years. Automated RL provides a framework in which different components of RL, including MDP modeling, algorithm selection, and hyper-parameter optimization, are modeled and defined automatically. In this article, we present the literature on automated RL (AutoRL), including the recent large language model (LLM) based techniques. We also discuss the recent work on techniques that are not presently tailored for automated RL but hold promise for future integration into AutoRL. Furthermore, we discuss the challenges, open questions, and research directions in AutoRL.
  •  
❌