❌

Normal view

Longitudinal Effects of a Smartphone Game (Tumaini) for HIV Prevention Among Kenyan Adolescents: 45-Month Trajectories of Condom Use–Related Proximal Outcomes From a Randomized Controlled Trial

Background: African adolescents and young adults account for a disproportionate number of new HIV infections. There is an urgent need to identify scalable and cost-effective behavioral HIV prevention strategies for this population. Using a condom at first sex is associated with a higher likelihood of consistent use later. Tumaini (“Hope for the Future” in Swahili; Emory University) is a choose-your-own-adventure smartphone game that has been shown to reduce the risk of unprotected first sex by end line in a 45-month randomized controlled trial in western Kenya. Objective: This study aimed to assess the impact of Tumaini on proximal outcomes related to condom use at first sex (specifically, behavioral intentions, self-efficacy, attitudes, and knowledge) longitudinally across mid-adolescence in the above trial. Methods: Adolescent participants (n=996, mean baseline age 14, SD 0.56 years) were randomized 1:1 to receive either a smartphone loaded with Tumaini or an attention-control math game for 5 to 7 weeks at 3 time points (mean age 14.0, SD 0.56; 15.3, SD 0.55; and 16.0, SD 0.56 years, respectively). They completed a behavioral survey at 13 time points, through mean age 17.7 (SD 0.56) years. Using generalized estimating equations and controlling for age at baseline, we modeled mean scores (overall and stratified by gender) on a range of condom-related survey items over time to assess mean differences at specific time points. We applied appropriate Bonferroni corrections to inferences about cross-arm differences in mean changes relative to baseline at 4 time points (after each intervention period and at end line; α=.05/4) and within-arm mean changes relative to baseline at each of the 12 post-baseline time points (α=.05/12). Analyses were conducted as intent-to-treat. Results: At end line, 97.8% (n=974) of the sample had been retained. Participants in both arms dedicated a mean total of >30 hours to their assigned game. There was significant improvement across all condom-related proximal outcomes in the intervention arm relative to the control arm immediately after initial intervention exposure. For almost all outcomes, a significant cross-arm difference was also present at end line and for most outcomes at the 2 intervening comparison time points. Some outcomes saw stronger intervention effects on female participants (eg, self-efficacy to refuse unprotected sex) or male participants (eg, knowledge that condoms are an effective way to prevent HIV). In each arm, intention to use a condom at first sex was consistently higher among male participants; however, female intervention-arm scores overtook male control-arm scores following initial intervention exposure. Conclusions: Tumaini significantly improved theory-based proximal outcomes related to condom use, with effects sustained 45 months post initial exposure and 16 months post most recent exposure. Adolescents benefited from even short-term exposure, though repeated exposure generally sustained and reinforced intervention effects. As access to smartphones increases, Tumaini has potential for high scalability and impact on condom-related outcomes. Trial Registration: ClinicalTrials.gov NCT04437667; https://clinicaltrials.gov/study/NCT04437667

Effects of Internet-Based Dementia Risk Reduction Education on Risk and Protective Factor Knowledge, Intentions, and Health Behaviors: Randomized Controlled Trial

Background: Dementia prevention through the reduction of modifiable risk factors is gaining attention as a public health strategy. However, public knowledge of dementia risk and protective factors remains low. Web-based education offers a potential solution to raise awareness and promote risk-reduction behaviors. Objective: This randomized controlled trial evaluated the effectiveness of DementiaRisk.ca, an internet-based multimedia educational intervention, in increasing knowledge of dementia risk factors, intentions to engage in risk reduction behaviors, and changes in health behaviors. Methods: A 2-arm randomized controlled trial was conducted with 510 participants (265 in the intervention group and 245 in the control group). Participants were randomized to receive either the e-learning about dementia risk and promoting brain health, which included a multimedia lesson and microlearning emails, or a control intervention focused on mild cognitive impairment. Outcomes included knowledge of dementia risk factors, intentions to engage in risk reduction, and health behaviors, measured at baseline (T1), 4 weeks (T2), and 2 months postintervention (T3). Outcomes were analyzed using linear mixed effects models with fixed effects for group, time, and their interaction, and a random intercept for participants. Results: Of the 510 randomized participants, 405 (79.4%) completed all intervention components. Participants were predominantly female (n=309, 60.6%) and aged 55 years or older (n=284, 55.7%). Baseline mean dementia knowledge scores were 17.0 (SD 5.5) in the intervention group and 17.4 (SD 6.0) in the control group. At T2, scores increased to 25.8 (SD 4.5) and 23.6 (SD 5.1), respectively, yielding a between-group difference of 2.2 points (95% CI 1.2‐3.2;

Breast Cancer Screening Knowledge and Sentiments in Singaporean Women: Mixed Methods Study Using Topic Modeling, Sentiment Analysis, and Structured Questionnaire Data

Background: Mammography screening uptake in Singapore remains below 40% despite campaigns and subsidies. Natural language processing (NLP) can extract nuanced attitudes from free text that fixed response options miss, revealing latent factors influencing breast cancer (BC) screening behavior. Objective: This study characterized women’s attitudes toward mammography using mixed methods data, examined associations between BC awareness and screening willingness, and identified barriers and facilitators through NLP of free-text responses. Methods: We conducted a cross-sectional study within the multicenter cohort in Singapore (October 2021-December 2023). In total, 4169 women aged 35‐59 years (median 48, IQR 43‐54) were recruited via convenience sampling (3 hospitals and 2 polyclinics). Participants completed online structured questionnaires on demographics and screening history, then a BC education quiz with feedback. Participants answering >80% correctly were classified as “BC-aware.” Posteducation, participants reported screening willingness (motivated or neutral) with optional free-text explanations. Logistic regression models (adjusted for study site, age, ethnicity, marital status, housing, and education) examined the associations with willingness. For 3819 English-language respondents, biterm topic modeling identified themes and sentiment analysis quantified emotional tone. Statistical significance: =.05. Results: Overall, 79% (3287/4169) were BC-aware, and 94% (3908/4169) reported increased motivation posteducation. BC-aware women had higher screening motivation than BC-unaware women (adjusted odds ratio [aOR] 2.88, 95% CI 2.19‐3.80;

Investigating the Effect of Hospital Infection Control Informatization on Optimizing Microbiological Specimen Submission Before Antibiotic Therapy: Failure Mode and Effects Analysis

Background: Antimicrobial resistance (AMR) poses a critical global health threat, with inappropriate antibiotic use being a major driver. Timely microbiological specimen submission before initiating antibiotic therapy is a cornerstone of antimicrobial stewardship (AMS), enabling pathogen-directed therapy and reducing unnecessary broad-spectrum exposure. However, suboptimal compliance remains common due to workflow interruptions, technological barriers, and behavioral factors. Failure Mode and Effects Analysis (FMEA), a proactive risk-assessment method widely used in health care quality improvement, provides a systematic framework to identify process vulnerabilities and prioritize corrective actions. Despite its increasing application, few studies have integrated FMEA with hospital informatization to optimize microbiological specimen submission workflows in routine AMS practice. Objective: This study aimed to systematically identify workflow risks affecting preantibiotic microbiological specimen submission and to design, implement, and evaluate informatization-enabled interventions using an FMEA-based framework. Methods: FMEA was conducted at a tertiary hospital in China. A multidisciplinary team identified potential failure modes across 4 domains: health information systems, personnel, administration, and external support. Risk Priority Numbers (RPNs) and Action Priority (AP) indices were calculated for each failure mode. Targeted interventions were implemented, including dual-verification barcode scanning, artificial intelligence-driven clinical decision support alerts, EHR-integrated training modules, and automated compliance dashboards. Pre- and postintervention specimen submission rates (January 2024-December 2024) were analyzed using the Mann-Kendall trend test. Results: The top 5 failure modes included PDA barcode scanning failures (RPN=175), inadequate clinical decision support (RPN=140), insufficient clinician awareness (RPN=56), suboptimal oversight mechanisms, and patient-related barriers. Postintervention, significant upward trends were observed in overall specimen submission rates (

eHealth Literacy and Type 2 Diabetes Prevention Among At-Risk Populations: Mechanistic Systematic Review Using Theory-Driven Thematic Analysis

Background: Type 2 diabetes (T2D) is emerging as a growing global public health crisis. Early and effective interventions can reduce T2D incidence among at-risk populations. Compared with traditional approaches, digital health technologies offer promising opportunities for prevention, with eHealth literacy (eHL) emerging as a critical determinant of digital prevention outcomes. Objective: This systematic review aims to synthesize and explain the pathways and mechanisms through which eHL supports T2D prevention among at-risk populations. Methods: We searched Scopus, Web of Science, and PubMed databases for English-language original research published between January 1, 2000, and August 14, 2025. Studies included were prevention research involving eHL engagement among populations at risk for T2D. Nonoriginal literature, such as editorials and abstracts, as well as research protocols, was excluded. The findings were synthesized using a thematic analysis approach, integrating the Theoretical Domains Framework with the eHL model. Two reviewers independently screened literature and extracted data, and discrepancies were resolved by a third reviewer. The Mixed Methods Appraisal Tool was used to assess risk of bias. Results: This review included 28 studies (n=13,100), mostly quantitative and published within the past decade, targeting people with prediabetes, prior gestational diabetes, and overweight/metabolic risk. Study quality was moderate to high (Mixed Methods Appraisal Tool 60%‐100%) with no high risk of bias. eHL supported prevention mainly through knowledge (28/28), behavioral regulation (16/28), social influences (15/28), environmental resources (12/28), and goals (11/28), while emotions, memory, attention, decision process, and beliefs about competence were rarely addressed. Health literacy (27/28), information literacy (20/28), and communicative eHL (20/28) were most common; critical eHL and media literacy were not addressed. Studies reported positive outcomes: high engagement, weight loss (≥5%), improved glycemic markers, and enhanced lifestyle behaviors. Conclusions: This is the first systematic exploration of eHL mechanism pathways in T2D prevention via theoretical mapping. We found interventions yield positive effects despite highly uneven mechanism application: extant research relies excessively on knowledge and behavioral pathways while underemphasizing emotional support, autonomy, and critical evaluation—factors linked to long-term adherence. We provide a mechanism-based framework and identify critical gaps, including the absence of focus on critical eHL and media literacy. This review is limited by substantial variation across studies that did not allow for meta-analysis and by the limited evidence base on eHL. Future interventions should explore and test emotional and autonomy support, information discernment training, and accessibility optimization in T2D prevention. These comprehensive, equity-focused intervention approaches will help ensure that eHL becomes a truly effective public health tool that benefits everyone, especially at-risk and vulnerable populations. Trial Registration: PROSPERO CRD42025630395; https://www.crd.york.ac.uk/PROSPERO/view/CRD42025630395

‘Pokémon Pokopia’ is a game about rehabilitating a broken world — and I love it

11 March 2026 at 03:48
The latest Pokémon game is a cozy life simulator like "Animal Crossing" or "Minecraft," but in a way that feels more grounded to our actual world.

Quality Challenges in Municipal Telecare Call Center Services: Qualitative Evaluation Using the Anchored, Realistic, Cocreated, Human, Integrated, and Evaluated (ARCHIE) Framework

Background: Telecare is seen as a promising technology aimed at enhancing the accessibility and efficiency of health care services. Although focus on quality has been highly prioritized within the health care services, there is a need to explore the quality of telecare services in general and municipal telecare call centers (CCs) in particular, as health and assistive technologies are increasingly being implemented in patients’ homes. Objective: The study sought to explore which factors influence the quality of telecare services provided by municipal telecare CCs in Norway, evaluated through the anchored, realistic, cocreated, human, integrated, and evaluated (ARCHIE) framework. Methods: The study had a multiple-case design. Interviews were the main source of data from 15 informants from 5 municipal telecare CCs across Norway. Observation and document studies were used for background and contextualization. To explore and evaluate quality, a combined deductive–inductive analysis was conducted. Results: Evaluated against the ARCHIE framework, none of the quality criteria were fully met. Due to the telecare service not being sufficiently anchored for all patients, it was challenging to provide realistic technologies. The collaborative work was difficult, with challenges in recruiting patients. The human principle was characterized by variation of knowledge and national guidelines. Municipal telecare CCs were not integrated into the health care services, and data must be used to a greater extent for evaluation and learning than is currently the case. Conclusions: The findings suggest that municipal telecare CC services have several shortcomings in providing high-quality health care. Relating the quality principles identified by the ARCHIE framework to normalization process theory constructs indicates that the CC service remains in a transitional phase of normalization. To improve the telecare CC services and enhance communication and integration, policymakers need to reduce fragmentation in the broader health care system. Further national standardization to professionalize the telecare CC services should be developed. The telecare CCs need to improve their service related to all indicators of the ARCHIE framework. Training for telecare operators should be prioritized.

Google rolls out new Gemini capabilities to Docs, Sheets, Slides, and Drive

10 March 2026 at 21:00
The idea behind the new features is to make the apps more personal and capable to help users get things done faster, right within the platforms themselves.

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use

arXiv:2603.08262v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm shift from passive information retrieval to dynamic, agentic interaction. While general-purpose tool learning has witnessed a surge in benchmarks, the financial sector, characterized by high stakes, strict compliance, and rapid data volatility, remains critically underserved. Existing financial evaluations predominantly focus on static textual analysis or document-based QA, ignoring the complex reality of tool execution. Conversely, general tool benchmarks lack the domain-specific rigor required for finance, often relying on toy environments or a negligible number of financial APIs. To bridge this gap, we introduce FinToolBench, the first real-world, runnable benchmark dedicated to evaluating financial tool learning agents. Unlike prior works limited to a handful of mock tools, FinToolBench establishes a realistic ecosystem coupling 760 executable financial tools with 295 rigorous, tool-required queries. We propose a novel evaluation framework that goes beyond binary execution success, assessing agents on finance-critical dimensions: timeliness, intent type, and regulatory domain alignment. Furthermore, we present FATR, a finance-aware tool retrieval and reasoning baseline that enhances stability and compliance. By providing the first testbed for auditable, agentic financial execution, FinToolBench sets a new standard for trustworthy AI in finance. The tool manifest, execution environment, and evaluation code will be open-sourced to facilitate future research.

Deconstructing Multimodal Mathematical Reasoning: Towards a Unified Perception-Alignment-Reasoning Paradigm

arXiv:2603.08291v1 Announce Type: new Abstract: Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems that involve both textual and visual modalities. However, current models still face significant challenges in real-world visual math tasks. They often misinterpret diagrams, fail to align mathematical symbols with visual evidence, and produce inconsistent reasoning steps. Moreover, existing evaluations mainly focus on checking final answers rather than verifying the correctness or executability of each intermediate step. To address these limitations, a growing body of recent research addresses these issues by integrating structured perception, explicit alignment, and verifiable reasoning within unified frameworks. To establish a clear roadmap for understanding and comparing different MMR approaches, we systematically study them around four fundamental questions: (1) What to extract from multimodal inputs, (2) How to represent and align textual and visual information, (3) How to perform the reasoning, and (4) How to evaluate the correctness of the overall reasoning process. Finally, we discuss open challenges and offer perspectives on promising directions for future research.

CORE-Acu: Structured Reasoning Traces and Knowledge Graph Safety Verification for Acupuncture Clinical Decision Support

arXiv:2603.08321v1 Announce Type: new Abstract: Large language models (LLMs) show significant potential for clinical decision support (CDS), yet their black-box nature -- characterized by untraceable reasoning and probabilistic hallucinations -- poses severe challenges in acupuncture, a field demanding rigorous interpretability and safety. To address this, we propose CORE-Acu, a neuro-symbolic framework for acupuncture clinical decision support that integrates Structured Chain-of-Thought (S-CoT) with knowledge graph (KG) safety verification. First, we construct the first acupuncture Structured Reasoning Trace dataset and a schema-constrained fine-tuning framework. By enforcing an explicit causal chain from pattern identification to treatment principles, treatment plans, and acupoint selection, we transform implicit Traditional Chinese Medicine (TCM) reasoning into interpretable generation constraints, mitigating the opacity of LLM-based CDS. Furthermore, we construct a TCM safety knowledge graph and establish a ``Generate--Verify--Revise'' closed-loop inference system based on a Symbolic Veto Mechanism, employing deterministic rules to intercept hallucinations and enforce hard safety boundaries. Finally, we introduce the Lexicon-Matched Entity-Reweighted Loss (LMERL), which corrects terminology drift caused by the frequency--importance mismatch in general optimization by adaptively amplifying gradient contributions of high-risk entities during fine-tuning. Experiments on 1,000 held-out cases demonstrate CORE-Acu's superior entity fidelity and reasoning quality. Crucially, CORE-Acu achieved 0/1,000 observed safety violations (95\% CI: 0--0.37\%), whereas GPT-4o exhibited an 8.5\% violation rate under identical rules. These results establish CORE-Acu as a robust neuro-symbolic framework for acupuncture clinical decision support, guaranteeing both reasoning auditability and strict safety compliance.

M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering

arXiv:2603.08369v1 Announce Type: new Abstract: Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplored bottleneck: inaccurate visual perception. Through systematic analysis, we find that the most failures originate from incorrect or incomplete visual evidence extraction rather than deficiencies in reasoning capability. Moreover, models tend to remain overly confident in their initial perceptions, making standard strategies such as prompt engineering, multi-round self-reflection, or posterior guidance insufficient to reliably correct errors. To address this limitation, we propose M3-ACE, a multi-agentic context engineering framework designed to rectify visual perception in multimodal math reasoning. Instead of directly aggregating final answers, our approach decouples perception and reasoning by dynamically maintaining a shared context centered on visual evidence lists. Multiple agents collaboratively contribute complementary observations, enabling the system to expose inconsistencies and recover missing perceptual information. To support stable multi-turn collaboration, we further introduce two lightweight tools: a Summary Tool that organizes evidence from different agents into consistent, complementary, and conflicting components, and a Refine Tool that filters unreliable samples and guides iterative correction. Extensive experiments demonstrate that M3-ACE substantially improves visual mathematical reasoning performance across multiple benchmarks. Our method establishes new state-of-the-art results 89.1 on the MathVision benchmark and achieves consistent improvements on other related datasets, including MathVista and MathVerse. These results highlight the importance of perception-centric multi-agent collaboration for advancing multimodal reasoning systems.

RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback

arXiv:2603.08561v1 Announce Type: new Abstract: Large language model (LLM)-based agents trained with reinforcement learning (RL) have shown strong potential on complex interactive tasks. However, standard RL paradigms favor static problem-solving over continuous adaptation: agents often converge to suboptimal strategies due to insufficient exploration, while learned knowledge remains implicit within parameters rather than explicitly retrievable, limiting effective experiential learning. To address these limitations, we introduce RetroAgent, an online RL framework that empowers agents to master complex interactive environments not just by solving, but by evolving. Concretely, RetroAgent features a hindsight self-reflection mechanism that produces dual intrinsic feedback: (1) intrinsic numerical feedback that that tracks incremental subtask completion relative to prior attempts, rewarding promising explorations, and (2) intrinsic language feedback that distills reusable lessons into a memory buffer, retrieved via our proposed Similarity & Utility-Aware Upper Confidence Bound (SimUtil-UCB) strategy balancing relevance, utility, and exploration to effectively leverage past experiences. Extensive experiments on two model families across four challenging agentic tasks demonstrate that RetroAgent significantly outperforms existing methods, achieving state-of-the-art results -- e.g., surpassing Group Relative Policy Optimization (GRPO)-trained agents by +18.3% on ALFWorld, +15.4% on WebShop, +27.1% on Sokoban, and +8.9% on MineSweeper -- while exhibiting strong test-time adaptation and generalization to out-of-distribution scenarios.

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

arXiv:2603.08652v1 Announce Type: new Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning. However, existing CoT-based T2I methods largely rely on abstract natural-language planning, which lacks the precision required for complex spatial layouts, structured visual elements, and dense textual content. In this work, we propose CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enabling explicit and verifiable intermediate planning for image generation. Given a text prompt, CoCo first generates executable code that specifies the structural layout of the scene, which is then executed in a sandboxed environment to render a deterministic draft image. The model subsequently refines this draft through fine-grained image editing to produce the final high-fidelity result. To support this training paradigm, we construct CoCo-10K, a curated dataset containing structured draft-final image pairs designed to teach both structured draft construction and corrective visual refinement. Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41.23% over direct generation, while also outperforming other generation methods empowered by CoT. These results demonstrate that executable code is an effective and reliable reasoning paradigm for precise, controllable, and structured text-to-image generation. The code is available at: https://github.com/micky-li-hd/CoCo

OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning

arXiv:2603.08655v1 Announce Type: new Abstract: We introduce OfficeQA Pro, a benchmark for evaluating AI agents on grounded, multi-document reasoning over a large and heterogeneous document corpus. The corpus consists of U.S. Treasury Bulletins spanning nearly 100 years, comprising 89,000 pages and over 26 million numerical values. OfficeQA Pro consists of 133 questions that require precise document parsing, retrieval, and analytical reasoning across both unstructured text and tabular data. Frontier LLMs including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro Preview achieve less than 5% accuracy on OfficeQA Pro when relying on parametric knowledge, and less than 12% with additional access to the web. When provided directly with the document corpus, frontier agents still struggle on over half of questions, scoring 34.1% on average. We find that providing agents with a structured document representation produced by Databricks' ai_parse_document yields a 16.1% average relative performance gain across agents. We conduct additional ablations to study the effects of model selection, table representation, retrieval strategy, and test-time scaling on performance. Despite these improvements, significant headroom remains before agents can be considered reliable at enterprise-grade grounded reasoning.

Agentic Critical Training

arXiv:2603.08706v1 Announce Type: new Abstract: Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contrast successful actions against suboptimal alternatives and thus lack awareness of action quality. Recent approaches attempt to address this by introducing self-reflection supervision derived from contrasts between expert and alternative actions. However, the training paradigm fundamentally remains imitation learning: the model imitates pre-constructed reflection text rather than learning to reason autonomously. We propose Agentic Critical Training (ACT), a reinforcement learning paradigm that trains agents to identify the better action among alternatives. By rewarding whether the model's judgment is correct, ACT drives the model to autonomously develop reasoning about action quality, producing genuine self-reflection rather than imitating it. Across three challenging agent benchmarks, ACT consistently improves agent performance when combined with different post-training methods. It achieves an average improvement of 5.07 points over imitation learning and 4.62 points over reinforcement learning. Compared to approaches that inject reflection capability through knowledge distillation, ACT also demonstrates clear advantages, yielding an average improvement of 2.42 points. Moreover, ACT enables strong out-of-distribution generalization on agentic benchmarks and improves performance on general reasoning benchmarks without any reasoning-specific training data, highlighting the value of our method. These results suggest that ACT is a promising path toward developing more reflective and capable LLM agents.
❌