❌

Reading view

Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval

arXiv:2510.20486v2 Announce Type: replace-cross Abstract: Imbalanced labels can cause frequent samples to dominate AI-based quantitative remote sensing, degrading rare-event retrieval. In rain-rate retrieval based on satellite infrared brightness temperatures, this imbalance leads to systematic underestimation of rare high-intensity rainfall. In this study, Hurdle-Retrieval Model Imbalanced Learning (RMIL) is proposed. Following a divide-and-conquer strategy, Hurdle-RMIL separates zero inflation from the long-tailed distribution of positive rain. A hurdle model handles zero inflation, whereas RMIL exploits invariance under fixed observation conditions of the rainfall-to-satellite forward process to derive a Bayes-based transformation linking conditional distributions under naturally long-tailed and hypothetical balanced rainfall. This transformation enables the balanced-distribution model to be learned from natural samples without constructing a balanced dataset. Comparisons with conventional learning, classification-regression modeling, cost-sensitive learning, and generative learning using test data from multiple regions in China show that Hurdle-RMIL mitigates systematic underestimation and improves detection of rare high-intensity and extreme rainfall without markedly degrading lower-threshold accuracy. At 0.1-10 mm per hour, its root mean square error remains close to those of the best baselines, and it yields the highest equitable threat score (ETS) at most evaluated thresholds, with its advantage becoming more pronounced at high thresholds. At 30 mm per hour, its ETS is 0.051 versus 0.015 for the best baseline, and its mean error is -25.41 mm per hour versus -28.98 mm per hour. Case studies further show improved representations of rainfall intensity and spatial extent, demonstrating that Hurdle-RMIL effectively addresses rainfall-distribution imbalance and improves the retrieval of rare high-intensity rainfall.
  •  

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

arXiv:2605.20555v2 Announce Type: replace-cross Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO). In contrast to Reinforcement Learning with Verifiable Rewards (RLVR) methods, our proposal does not involve a Kullback Leibler (KL) regularization or critic; the trainable policy and the reference anchor are coupled through the logit averaging structure to leverage the reasoning expertise of the trainable policy while maintaining the formatting advantage of SFT. Our method is evaluated on MATH, cn-k12, and MMLU, and the results show a higher accuracy or at least comparable accuracy relative to the canonical KL-regularized GRPO.
  •  

Genetically encoded fluorescent reporters to visualize α-synuclein pathology in live brain

The development of genetically encoded fluorescent reporters, along with their corresponding knock-in mouse lines for labeling α-Syn inclusions, enables diverse applications in studying the propagation and pathological effects of α-Syn inclusions in the live brain.
  •  

Expansion of outer cortical CUX2 neurons requires adaptations for DNA repair

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10290-4

The transcription factor ATF4 is shown to regulate double-stranded DNA repair within vulnerable CUX2+ upper-layer 2/3 cortical neurons, enabling their survival during development.
  •  

Integrating Untargeted Metabolomics and Transcriptomics in Mice with Pulmonary Tuberculosis to Reveal Changes in Linoleic Acid and Its Metabolism in Lung Monocyte-Derived Macrophages

Pathogens. 2026 Feb 27;15(3):254. doi: 10.3390/pathogens15030254.

ABSTRACT

Pulmonary tuberculosis (TB) remains a major global health challenge. The molecular and metabolic responses of monocyte-derived macrophages (MDMs), which are critical for host defense against Mycobacterium tuberculosis (Mtb), are not fully characterized. A murine pulmonary TB model was established by intravenous injection of BALB/c mice with the attenuated Mtb strain H37Ra; controls received saline. After 8 weeks, lung MDMs were isolated for integrated transcriptomic and untargeted metabolomic profiling. Transcriptomic analysis identified 3970 differentially expressed genes (DEGs) in infected MDMs, including upregulated Ptpn1, Dgat2, and Alox5ap and downregulated Cyld, Zfp61, and Mapk11. Metabolomic profiling revealed 113 differentially accumulated metabolites (DAMs). Taurocholic acid and linoleic acid were identified as potential diagnostic biomarkers, both achieving an area under the curve (AUC) of 1.0 in ROC analysis. Integrated omics analysis showed a positive correlation between linoleic acid levels and the expression of Tbxas1, Acaa1b, and Acox1, implicating lipid metabolic pathways in the host response to TB. This multi-omics study delineates key molecular and metabolic alterations in lung MDMs during TB infection. The identified metabolites, taurocholic acid and linoleic acid, show promise as biomarkers, while dysregulated linoleic acid metabolism represents a potential target for novel diagnostic and therapeutic strategies against TB.

PMID:41901707 | PMC:PMC13029415 | DOI:10.3390/pathogens15030254

  •  

Evaluating Prompting Strategies for Chart Question Answering with Large Language Models

arXiv:2603.22288v1 Announce Type: cross Abstract: Prompting strategies affect LLM reasoning performance, but their role in chart-based QA remains underexplored. We present a systematic evaluation of four widely used prompting paradigms (Zero-Shot, Few-Shot, Zero-Shot Chain-of-Thought, and Few-Shot Chain-of-Thought) across GPT-3.5, GPT-4, and GPT-4o on the ChartQA dataset. Our framework operates exclusively on structured chart data, isolating prompt structure as the only experimental variable, and evaluates performance using two metrics: Accuracy and Exact Match. Results from 1,200 diverse ChartQA samples show that Few-Shot Chain-of-Thought prompting consistently yields the highest accuracy (up to 78.2\%), particularly on reasoning-intensive questions, while Few-Shot prompting improves format adherence. Zero-Shot performs well only with high-capacity models on simpler tasks. These findings provide actionable guidance for selecting prompting strategies in structured data reasoning tasks, with implications for both efficiency and accuracy in real-world applications.
  •  

Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs

arXiv:2603.20209v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) combine the linguistic strengths of LLMs with the ability to process multimodal data, enbaling them to address a broader range of visual tasks. Because MLLMs aim at more general, human-like competence than language-only models, we take inspiration from the Wechsler Intelligence Scales - an established battery for evaluating children by decomposing intelligence into interpretable, testable abilities. We introduce KidGym, a comprehensive 2D grid-based benchmark for assessing five essential capabilities of MLLMs: Execution, Perception Reasoning, Learning, Memory and Planning. The benchmark comprises 12 unique tasks, each targeting at least one core capability, specifically designed to guage MLLMs' adaptability and developmental potential, mirroring the stages of children's cognitive growth. Additionally, our tasks encompass diverse scenarios and objects with randomly generated layouts, ensuring a more accurate and robust evluation of MLLM capabilities. KidGym is designed to be fully user-customizable and extensible, allowing researchers to create new evaluation scenarios and adjust difficuly levels to accommodate the rapidly growing MLLM community. Through the evaluation of state-of-the-art MLLMs using KidGym, we identified significant insights into model capabilities and revealed several limitations of current models. We release our benchmark at: https://bobo-ye.github.io/KidGym/.
  •  
❌