❌

Normal view

Depictions of Depression in Generative AI Video Models: Mixed Methods Study of OpenAI’s Sora 2

Background: Generative AI video models are increasingly capable of producing complex depictions of mental health experiences, yet little is known about how these systems represent conditions such as depression. Because AI-generated content may reach people during vulnerable periods, understanding what visual narratives these models produce for sensitive concepts carries clinical relevance. Objective: This study aimed to characterize how OpenAI’s Sora 2 generative AI video model depicts depression and examine whether depictions differ between the consumer app and developer API access points, which differ in their product layer mediation. Methods: We generated 100 videos using the single-word prompt “Depression” across 2 access points: the consumer app (n=50, 50%) and developer API (n=50, 50%). Two trained coders independently coded narrative structure, visual environments, objects, figure demographics, and figure states. Interrater reliability was assessed using the Cohen κ, with dimensions showing insufficient agreement excluded from analysis. Computational features (visual aesthetics, audio, semantic content, and temporal dynamics) were extracted and compared between modalities using 2-tailed Welch tests with Benjamini-Hochberg false discovery rate correction. Results: App-generated videos exhibited a pronounced recovery bias: 78% (39/50) featured narrative arcs progressing from depressive states toward resolution compared with 14% (7/50) of API outputs. This divergence was reinforced across channels. App videos brightened over time (mean slope 2.90, SD 2.43 per second vs −0.18, SD 1.24 per second for the API; Cohen =1.59;

Design Guidelines for Online Health Forums: User-Centered Design Approach

Background: Online health forums are used widely, yet evidence of their effectiveness is inconsistent. Evidence-based forum design guidance grounded in theory and lived experience could improve the efficacy and outcomes of these forums for the many people using them worldwide. Objective: This study aimed to draw on the experience of online forum users and staff, and insights from existing research on technology design and self-determination theory, to generate a set of theoretically grounded guidelines for safe and well-being–supportive online forums. Methods: We conducted 54 semistructured interviews (36 forum users and 18 forum staff) and 4 design workshops with forum staff, combined with input from a multidisciplinary research team. Principles of qualitative framework analysis were used to adapt a preexisting framework for well-being–supportive technology to the context of online health forums. Results: The resulting design guidelines are framed around 4 overarching principles relating to the psychological needs for autonomy, competence, and relatedness as defined by self-determination theory, and the additional need for safety in online forums. Each principle is presented alongside pragmatic design heuristics and specific implementation strategies. Conclusions: User experiences of online health forums are mixed. We have drawn on self-determination theory to propose evidence-informed guidelines for service development and refinement, adaptable to specific user groups across diverse settings. International Registered Report Identifier (IRRID): RR2-https://doi.org/10.1136/bmjopen-2023-075142

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study

Background: The widespread adoption of electronic health records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for deidentifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in large language models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of health care data anonymization. Objective: This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes. Methods: We constructed a gold-standard corpus of 1000 Portuguese outpatient clinical notes, manually annotated by 5 trained researchers for 5 protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: 2 stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and 2 quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on a held-out test set of 500 notes using precision, recall, and -score metrics. Computational efficiency was quantified through end-to-end processing time. Results: The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-score of 0.855 (95% CI 0.823‐0.880), outperforming the stand-alone Llama-3.3-70B (0.726, 95% CI 0.704‐0.747), Dynex-QML-8B (0.733, 95% CI 0.709‐0.756), and Llama-3.1-8B (0.602, 95% CI 0.588‐0.615). Compared with Llama 3.3 70B, Dynex-QML (Llama 70B) improved macro-score by 0.128 (95% CI 0.091‐0.163; empirical 2-sided bootstrap

Digital Phenotyping of Lifestyle Profiles and Mental Well-Being in German Adults: Prospective Longitudinal Cohort Study

Background: Digital phenotyping uses passively collected smartphone-sensing data to characterize everyday behavior in naturalistic settings, and has become an important approach for studying mental well-being. Most previous studies have examined associations between individual sensing variables and mental health. However, mental well-being is likely reflected not by isolated behaviors but by combinations of co-occurring daily behaviors that together form lifestyles. Person-centered approaches capable of identifying these behavioral configurations may, therefore, provide more interpretable digital phenotypes; yet, such approaches have rarely been applied to passive smartphone-sensing data. Objective: This study aimed to examine whether smartphone-captured behavioral and environmental data could be used to derive interpretable day-level and person-level lifestyle profiles, and whether person-level profiles were associated with mental well-being. We also tested whether Big Five personality traits—extraversion, agreeableness, conscientiousness, openness, and negative emotionality—moderated these associations. Methods: The study used a 2-week prospective longitudinal cohort design with a sample of 553 German adults (mean age 42.12, SD 12.89 years; 44.65% female) drawn from an initial sample recruited according to quotas designed to reflect the German population. Ten smartphone-sensing indicators captured 5 domains, including communication and social media app use, mobility, physical activity, environmental context, and phone-use intensity. Mental well-being was assessed using the Warwick–Edinburgh Mental Well-Being Scale, and personality was assessed using the 15-item Big Five Inventory–2 Extra-Short Form. We used multilevel latent profile analysis to identify day-level profiles nested within person-level profiles. Associations between profiles and mental well-being were tested using classification-error–adjusted mean comparisons and omnibus Wald tests. Moderation was examined using hierarchical regressions comparing models with and without profile-by-personality interactions. Results: Eight day-level profiles and 7 person-level profiles were identified. Day-level profiles reflected distinct combinations of smartphone-sensing indicators. Person-level profiles represented different distributions of these daily patterns. Profiles differed significantly only in positive functioning (Wald ²=13.39; =.04), not in overall mental well-being, positive affect, or satisfying interpersonal relationships. The physically active and unplugged profile had higher positive functioning than the mobile and always-on social profile (mean 3.94, SD 0.63 vs mean 3.61, SD 0.74; Cohen =0.47; 95% CI 0.21‐0.73). No other pairwise differences were significant. Sensitivity analyses excluding the smallest profile produced comparable results, supporting the robustness of the findings. Personality-by-profile interactions did not significantly improve prediction for any well-being outcome. Conclusions: The findings extend the field by showing that transparent, person-centered digital phenotypes can distinguish variation in positive functioning, although causal conclusions cannot be drawn. In real-world settings, such interpretable profiles could support understandable monitoring tools and, following prospective replication and validation, inform personalized multibehavior interventions that target combinations of behaviors rather than single behaviors in isolation.

Java News Roundup: New OpenJDK JEPs, CDI 5.0, Spring, Open Liberty, RefactorFirst, ADK for Kotlin

15 September 2026 at 04:15

This week's Java roundup for September 7th, 2026, features news highlighting: new JEPs for ahead-of-time compilation and structured concurrency; GA releases of Jakarta CDI 5.0 and ADK for Kotlin 1.0; the September 2026 edition of Open Liberty; point releases of TornadoVM and RefactorFirst; a maintenance release of Micronaut; and first releases candidates of Groovy 6.0 and Gradle 9.8.

By Michael Redlich

Evaluation of the Square Eyes Model as a Screening Tool for Identifying Digital Technologies in Wearable Camera Images Among Children: Laboratory Study

Background: Accurate measurements of children’s digital technology use are essential for understanding its potential implications on health and well-being. Wearable cameras can provide such measurements, but image coding is a high burden for researchers. Machine learning–based object-recognition models have the potential to reduce this burden by identifying images containing technology. Objective: This study aims to evaluate the performance of an object recognition model, the Square Eyes model, as a screening tool for identifying technologies in wearable camera images among children for further human review, as well as to examine the potential influence of face-blurring methods on the model’s performance. Methods: This study used data collected on 48 children (aged 3‐14 y) during an approximately 1-hour laboratory session. The children performed various technology-related tasks while wearing a camera. A total of 221,226 images were coded by humans and processed through the Square Eyes model. The performance of the Square Eyes model as a screening tool was evaluated by (1) assessing agreement between the model and human coding; (2) evaluating the N-back algorithm, an algorithm embedded in the model aimed to flag images requiring human review; and (3) examining the potential influence of facial-blurring on model performance. Results: Humans detected technology in 92,745 (41.9%) images, and the Square Eyes model detected technologies with an overall accuracy of 78.0%. When considering specific technologies, agreement between the model and human coders was the highest for (n=19,148, 54.3%) and (n=8492, 44.5%) and lowest for smaller devices such as (n=2600, 31.3%) and (n=3685, 25.1%). The model’s N-back algorithm effectively flagged images that required further human review, with only 7144 (3.2%) images that were not flagged for screening containing a human-coded technology. An explorative analysis indicated that using a square face-blurring with border could have reduced the model’s ability to accurately detect technologies. Conclusions: The Square Eyes model demonstrated overall satisfying accuracy in detecting technologies and successfully flagged images that required further review by humans. These findings suggest that the model could be used as an effective screening tool for reducing the burden of human coding. However, the model could be improved to more accurately detect smaller devices, and the form of facial blurring in images should be considered.
  • ✇MIT Technology Review
  • The AI industry has taken a doomer turn. What now? Will Douglas Heaven
    This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassa
     

The AI industry has taken a doomer turn. What now?

15 September 2026 at 01:54

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X.

Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology.

Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)

Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.

It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both.

And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.

Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call.

But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.

(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)

But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.

What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. The implication is that OpenAI has built a model so good it’s dangerous.  

But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.

The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.

OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.  

That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.  

Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.   

To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

Digitally Adapting LGBTQ-Affirmative Cognitive Behavioral Therapy for Chinese Men Who Have Sex With Men Living With HIV: User-Centered Design Approach

Background: Chinese men who have sex with men living with HIV (MSMLWH) experience substantial psychological distress driven by minority stress and HIV-related challenges. However, culturally tailored digital mental health interventions that address HIV-specific maladaptive cognitive schemas and culturally specific psychosocial stressors remain scarce in China. Objective: This study aimed to systematically adapt an evidence-based cognitive behavioral therapy (CBT) intervention Effective Skills to Empower Effective Men (ESTEEM) into a WeChat (Tencent) Mini-Program–based intervention (iESTEEM) specifically for Chinese MSMLWH and to evaluate its preliminary feasibility and usability. Methods: We used a three-phase user-centered design approach guided by the Assessment, Decision, Adaptation, Production, Topical Experts, Integration, Training, and Testing (ADAPT-ITT) framework. The study proceeded in three phases: (1) a qualitative needs assessment using semistructured interviews with 20 MSMLWH (mean age 23.25, SD 3.08 years); (2) systematic intervention adaptation and platform development, including theater testing (n=5); and (3) a 2-week pilot study involving 10 MSMLWH and five counselors to evaluate feasibility, usability, and acceptability through focus groups and objective platform analytics. Results: Phase 1 identified 3 major themes of psychological distress: persistent health anxiety fueled by catastrophizing, intersectional stigma internalization, the disclosure dilemma, and intimacy barriers rooted in defectiveness and shame schemas. Participants also prioritized anonymity and bite-sized learning. Guided by these findings, iESTEEM was developed as a counselor-assisted, privacy-preserving WeChat Mini-Program incorporating HIV-specific scenarios, multimodal learning modules, and a back-end risk-alert system. During the 2-week pilot, participants logged into the platform 14.1 (SD 6.7) times per person and completed 134.3 (SD 103.1) minutes of learning activities; all participants accessed module 1, and 90% (9/10) accessed modules 2‐5. Anxiety scores decreased from 8.9 (SD 2.3) to 7.2 (SD 3.0), whereas depression scores remained stable. All participants expressed a willingness to continue using the program and to recommend it to peers. Participants and counselors endorsed its contextual relevance, privacy protections, and clinical utility. Conclusions: This study provides a theory- and evidence-informed model for culturally adapting digital mental health interventions for highly stigmatized populations. By integrating lesbian, gay, bisexual, transgender, and queer (LGBTQ)-affirmative CBT principles, HIV-specific adaptations, and a privacy-preserving, counselor-assisted WeChat Mini-Program, iESTEEM demonstrated promising preliminary feasibility, acceptability, and engagement among Chinese MSMLWH. These findings support the potential of culturally tailored digital interventions to expand access to psychological support for this stigmatized population in resource-constrained settings. Ongoing randomized controlled trials will further evaluate its efficacy, implementation outcomes, and mechanism of action. Trial Registration: Chinese Clinical Trial Registry ChiCTR2400080263; https://www.chictr.org.cn/showproj.html?proj=216926

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

15 September 2026 at 00:27
The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.
❌