❌

Normal view

DuckDuckGo installs are up 30% as users reject being ‘force-fed’ Google’s AI Search

27 May 2026 at 06:32
Google overhauled Search at I/O 2026, replacing blue links with AI agents. The backlash has been swift. DuckDuckGo app installs spiked 30% as users seek a way out.

Self-Reported Health Outcomes in Metabolic Health YouTube Comments: Cross-Sectional Study and Rule-Based Natural Language Processing Framework Development and Validation

Background: YouTube is increasingly used for healthcasting, the sharing of evidence-based dietary and lifestyle interventions by domain experts. In the metabolic health domain, channels focused on therapeutic carbohydrate restriction have accumulated audiences of millions. A distinctive feature is the comment section, where viewers share first-person accounts of health changes, constituting a unique source of real-world outcome data at scale. However, extracting structured health information from unstructured comments presents computational challenges. Objective: This observational, cross-sectional study aims to develop and validate a precision-optimized computational framework for extracting self-reported health outcomes from healthcasting YouTube comments and to characterize the prevalence, distribution across health aspects, and channel-level variation of reported outcomes across a large-scale metabolic health corpus. Methods: This study analyzed 43,111 unique YouTube comments from 110 videos across 11 therapeutic carbohydrate restriction-focused healthcasting channels (37,458 unique authors; data span November 2013 to January 2026; collected via YouTube data application programming interface version 3). The methodology comprised 3 construction phases and 5 validation studies. The construction phases were (1) exploratory corpus characterization, (2) iterative development of a 35-aspect hierarchical health outcome ontology, and (3) precision-optimized rule-based classification, validated through precision validation (stratified sample of n=500), recall estimation (n=510), external validation on 5 held-out channels (n=12,653 comments), large language model–assisted interrater reliability assessment, and transformer baseline comparison against Bidirectional Encoder Representations from Transformers (BERT) and Robustly Optimized BERT Pretraining Approach (ROBERTa) classifiers. A supplementary aspect–based sentiment analysis contextualized the positive-only design. Results: The framework identified 1790 positive health outcome reports (1790/43,111, 4.15% prevalence), achieving 97.6% (488/500) precision (95% CI 95.7%-98.6%) and estimated 56.2% recall (95% CI 43.4%-67.9%). The reports described 6674 positive outcomes, distributed across 35 health aspects and 18 named disease conditions extending beyond weight loss: pain and inflammation reduction (1137/6674, 17%), type 2 diabetes improvement (977/6674, 14.6%), skin health (784/6674, 11.8%), and psychological well-being (731/6674, 11%). Over half (3355/6674, 50.3%) spanned multiple research objectives. Significant channel-level variation was observed (χ²10=927.5; P<.001), with positive outcome rates ranging from 1.32% to 10.40% (odds ratio 8.68, 95% CI 7.10-10.61). Transformer baselines achieved higher recall but lower precision, confirming their advantage for high-confidence corpus generation. A supplementary aspect-based sentiment analysis indicated a positive-to-negative ratio of approximately 4.6:1 (n=1003), with negative experiences (59/495, 11.9%) predominantly involving gastrointestinal and cardiovascular concerns. Conclusions: This study presents, to our knowledge, the first validated, rule-based framework for extracting self-reported metabolic health outcomes from healthcasting YouTube comments at corpus scale. Unlike existing recall-oriented social media health classifiers, the precision-optimized design achieves the confidence threshold required for outcomes research without manual review. These findings demonstrate that expert-led health content comment sections constitute a scalable, complementary data source for monitoring real-world engagement with dietary interventions, with implications for public health surveillance, platform design, and health communication research.

Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care: Preclinical Algorithm Validation

Background: The application of generative artificial intelligence to simplify medication use instructions has the potential to enhance people’s health by improving treatment adherence. Objective: We evaluated the performance of large language models (LLMs) in generating medication usage instructions to complement prescriptions in primary health care. Methods: This randomized, blinded experimental preclinical study used prescription-inducing scenarios, assigned to 62 health care professionals, to validate instructions generated by LLMs during electronic prescriptions. The instructions were generated by ChatGPT-4.0 (OpenAI), Llama3.1-8B (Meta), and Llama3.1-8B-RAG (Meta) using retrieval-augmented generation based on patient information leaflets. Performance metrics assessed adequacy, completeness, clarity, language simplification, usefulness, and errors in the generated instructions, with scores to analyze overall and individual metrics. Results: The 3 models yielded high overall scores for producing qualified instructions (ChatGPT-4.0: median 88.4, IQR 22.8; Llama3.1-8B: median 66.5, IQR 50.9; Llama3.1-8B-RAG: median 79.9, IQR 34.4; Kruskal-Wallis test P=.003). Llama3.1-8B-RAG received evaluations with similar overall scores to ChatGPT-4.0 (post hoc test, P=.05) and similar to Llama3.1-8B (post hoc test, P=.44). ChatGPT-4.0 outperformed Llama3.1-8B (Bonferroni test, P<.001). Regarding specific domains, Llama3.1-8B-RAG received scores equivalent to those of ChatGPT-4.0 for adequacy (mean 6.24, SD 2.3 vs mean 6.82, SD 2.1; post hoc test, P=.54); completeness (mean 5.94, SD 2.2 vs 6.55, SD 1.9; post hoc test P=.38), clarity (mean 5.77, SD 2.4 vs mean 6.68, SD 1.9; post hoc test P=.09), and usefulness (mean 5.42, SD 2.4 vs mean 5.96, SD 2.2; post hoc test P=.63). ChatGPT-4.0 received higher scores in the language simplification criterion than Llama3.1-8B-RAG (mean 7.05, SD 1.5 vs mean 5.44, SD 2.6; post hoc test P<.001). Interrater variability in assigning scores ranged from 4.2% (n=3) to 85.8% (n=6) among primary health care professionals. Instructions leading to incorrect use of the medication had similar frequency among the models(ChatGPT-4.0: n=15, 22.7%; Llama3.1-8B: n=19, 22.8%; Llama3.1-8B-RAG: n=19, 22.8%; chi-square test P=.71). The frequencies of hallucination were similar (ChatGPT-4.0: n=7, 10.6%; Llama3.1-8B: n=9, 13.6%; Llama3.1-8B-RAG: n=6, 9.1%; chi-square test P=.67). Conclusions: The open-source LLM enhanced with external information presented similar performance to the closed-source model, except for ChatGPT4.0, which was superior in language simplification of messages. LLM generation demonstrated potential for instructing patients on medication use. Nonetheless, the introduction of this innovation into the electronic prescribing workflow demands prescriber validation for human oversight of the technology and requires a strategy for LLM performance governance.

Safety of Telemedicine Versus In-Person Care for Patients With Tracheal Devices: Propensity Score–Matched Cohort Study

Background: Patients with tracheal diseases often require long-term follow-up after tracheal device placement, with a risk of adverse events that may lead to emergency care and unplanned interventions. Telemedicine has been proposed as an alternative to in-person follow-up to improve access and continuity of care. Objective: The primary objective of this study was to compare the need for emergency department (ED) visits between telemedicine and in-person groups. Secondary objectives included comparing hospital readmissions, 30-day hospital readmissions, and unplanned interventions between groups. Methods: This retrospective, single-institution study included adult patients with tracheal devices who underwent telemedicine and in-person outpatient clinic visits between 2020 and 2024. To balance the groups, we used 1:1 propensity score matching. We collected demographic and clinical data and evaluated the need for ED visits, hospital readmissions, 30-day hospital readmissions, and unplanned interventions. Kaplan-Meier estimation of time to first ED visit was performed to assess outcomes after outpatient visits. Results: A total of 483 patients (n=277, 57% telemedicine and n=206, 43% in-person) underwent 2487 visits (1258 telemedicine and 1229 in-person). After propensity score matching, 336 patients remained (168 in each group). There were no significant differences in the need for ED visits, hospital readmissions, or unplanned interventions. The telemedicine group had significantly fewer 30-day hospital readmissions (odds ratio 0.38, 95% CI 0.16-0.87; =.02). Kaplan-Meier analysis indicated no statistically significant difference in ED-free visits. Conclusions: Telemedicine follow-up was associated with outcomes comparable to those of in-person follow-up in this cohort of adult patients with tracheal devices, with no evidence of an increased need for ED visits. In the matched analysis, telemedicine was associated with lower odds of 30-day hospital readmission.
  • ✇MIT Technology Review
  • The Download: puncturing the AI jobs panic Thomas Macaulay
    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. A reality check on the AI jobs hysteria Despite the growing hysteria over AI’s threat to white-collar jobs, there’s still scant evidence that the technology has had a large-scale impact on the labor market. Analysis of US labor data shows that unemployment in occupations most exposed to AI is actually lower than in less-exposed jobs. There are also no
     

The Download: puncturing the AI jobs panic

26 May 2026 at 20:10

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

A reality check on the AI jobs hysteria

Despite the growing hysteria over AI’s threat to white-collar jobs, there’s still scant evidence that the technology has had a large-scale impact on the labor market.

Analysis of US labor data shows that unemployment in occupations most exposed to AI is actually lower than in less-exposed jobs. There are also no signs that large numbers of workers are shifting from AI-threatened professions into supposedly safer manual-labor jobs.

It’s true that things aren’t great in the job market—but the question is why. Here’s what the data really says about AI and jobs.

—David Rotman

Opinion: It’s time to address the looming crisis in entry-level work

—Georgios Petropoulos, an assistant professor at the USC Marshall School of Business

AI has not yet produced mass unemployment. But it may be quietly weakening the first rung of the career ladder.

A recent Stanford study found that young workers in AI-exposed occupations suffered a sharp decline in employment after the spread of generative AI. The same pattern didn’t appear in low-exposure jobs, suggesting AI is replacing junior tasks that once gave young workers their first foothold.

It’s time to rethink how we train, prepare, and support young people entering the workforce. Read this op-ed on how job seekers, businesses, and society can adapt.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 The Pope has called for governments to regulate AI 
In his first major teaching document, Pope Leo said AI must be “disarmed.” (BBC)
+ He warned that AI fuels war and misinformation. (CNN)
+ But could also “open up a horizon extending in all directions.” (Engadget)
+ Anthropic cofounder Chris Olah also spoke at the event. (Reuters $)

2 SpaceX has launched its biggest and most powerful rocket
The Starship V3 made its test flight debut two days after Elon Musk announced SpaceX’s IPO.(Guardian)+ SpaceX pulled off the launch, but not the landing. (Ars Technica)
+ The rocket could be key to SpaceX’s valuation. (Fortune $)
+ But rivals to the company are rising. (MIT Technology Review)

3 Huawei says it can make industry-leading chips within five years
The Chinese tech giant announced a breakthrough in chip design. (Reuters $)
+ Its progress underscores Beijing’s push to neutralize US sanctions. (NBC)
+ Chinese chip stocks rallied after the announcement. (Bloomberg $)

4 A new vaccine may protect against the Ebola strain behind the current crisis
Tests have shown promising results for the mRNA vaccine. (New Scientist)
+ Another Ebola vaccine that could be ready for trials in months. (BBC)
+ But vaccines face a new problem: their name. (MIT Technology Review)

5 A swimmer broke a world record at the ‘Steroid Olympics’
Athletes at the Enhance Games were encouraged to take dope. (Wired $)
+ Silicon Valley elites have backed the competition. (WP $)
+ Which fits right into 2026’s longevity vibes. (MIT Technology Review)

6 The EU plans to fine Google a massive antitrust penalty
For allegedly favoring its own services in search results. (CNBC)
+ It would be the largest penalty for breaching the Digital Markets Act. (Reuters $) 

7 US quantum computing subsidies may not be legal
Congressional critics say the funding has been misused. (Ars Technica)

8 AI is minting new billionaires—and workers want their share
The Samsung labor showdown reflects global concerns. (Rest of World)

9 China has launched artificial human embryos into orbit
To find out whether we can reproduce beyond Earth. (Gizmodo)

10 Jony Ives has designed Ferrari’s first fully-electric car
The legendary Apple designer has created a polarizing aesthetic. (FT $) 


Quote of the day

“Technology is never neutral, because it takes on the characteristics of those who devise, finance, regulate, and use it.” 

—Pope Leo issues a warning about AI in his first encyclical letter, entitled ‘Magnifica humanitas: On Safeguarding the Human Person in the Time of Artificial Intelligence.”

One More Thing

portrait of Monica Sanders
ALYSSA SCHUKAR


How climate vulnerability and the digital divide are linked

In Anacostia, a historic African-American section of Washington, DC, Monica Sanders is measuring Wi-Fi speeds. It’s below the FCC’s minimum to qualify as a broadband service. She then checks the temperature: 46.9 °F.

Sanders, an adjunct professor of law at Georgetown University, frequently records this combination of weak internet access and environmental conditions. Her work shows how underinvestment in infrastructure can leave underserved communities more exposed to climate risks like extreme heat and flooding.

Discover how the digital divide is shaping climate vulnerability in the US.

—Colleen Hagerty

We can still have nice things

A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ Here’s a joyful way to settle sibling squabbles: a mandatory dance-off.
+ Build the metropolis of your dreams in this browser-based city simulation game.
+ Watch this hypnotic tiny train move in a perfect, endless loop on a rotating turntable.
+ Take a nostalgic look at early computing history with this curated gallery of vintage punch cards.

A Non-Canonical Role of SMAD4 in Regulating 3D Genome Architecture to Inhibit Lung Squamous Cell Carcinoma Development

Adv Sci (Weinh). 2026 May 26:e75839. doi: 10.1002/advs.75839. Online ahead of print.

ABSTRACT

Lung squamous cell carcinoma (LUSC) lacks clearly defined key drivers and effective targeted therapies, reflecting an incomplete understanding of its molecular pathogenesis. Here, we identify SMAD4 as a critical regulator of three-dimensional (3D) genome organization in LUSC and uncover a mechanistic link between tumor suppressor loss and oncogenic transcriptional activation. By integrating clinical datasets, genetically engineered mouse models, human and murine LUSC cell lines, and multi-omics analyses, we demonstrate that SMAD4 deficiency promotes LUSC progression by unleashing EP300-mediated enhancer-promoter looping at the SOX2 locus. Mechanistically, SMAD4 does not directly bind SOX2 regulatory elements but instead constrains chromatin looping by sequestering EP300 away from loop anchor regions. Loss of SMAD4 leads to enhanced H3K27ac deposition, aberrant SOX2 activation, and increased LUSC tumor cell proliferation. Together, these findings reveal a non-canonical role for a transcription factor (e.g., SMAD4) in regulating dysregulated 3D genome architecture to inhibit tumor development.

PMID:42189071 | DOI:10.1002/advs.75839

A Non-Canonical Role of SMAD4 in Regulating 3D Genome Architecture to Inhibit Lung Squamous Cell Carcinoma Development

Adv Sci (Weinh). 2026 May 26:e75839. doi: 10.1002/advs.75839. Online ahead of print.

ABSTRACT

Lung squamous cell carcinoma (LUSC) lacks clearly defined key drivers and effective targeted therapies, reflecting an incomplete understanding of its molecular pathogenesis. Here, we identify SMAD4 as a critical regulator of three-dimensional (3D) genome organization in LUSC and uncover a mechanistic link between tumor suppressor loss and oncogenic transcriptional activation. By integrating clinical datasets, genetically engineered mouse models, human and murine LUSC cell lines, and multi-omics analyses, we demonstrate that SMAD4 deficiency promotes LUSC progression by unleashing EP300-mediated enhancer-promoter looping at the SOX2 locus. Mechanistically, SMAD4 does not directly bind SOX2 regulatory elements but instead constrains chromatin looping by sequestering EP300 away from loop anchor regions. Loss of SMAD4 leads to enhanced H3K27ac deposition, aberrant SOX2 activation, and increased LUSC tumor cell proliferation. Together, these findings reveal a non-canonical role for a transcription factor (e.g., SMAD4) in regulating dysregulated 3D genome architecture to inhibit tumor development.

PMID:42189071 | DOI:10.1002/advs.75839

A Dynamical Framework for Cognitive Processes Based on Transformations and Semantic Equivalence

arXiv:2605.23942v1 Announce Type: new Abstract: This paper proposes a structural and dynamical framework for modeling cognitive processes within a cybernetic perspective. Cognitive states are represented as elements of a state space evolving through an iterative update rule of the form \[ X_{t+1} = \pi\big(F(f(X_t))\big), \] where $f$ describes internal transformations, $F$ represents interpretative mappings, and $\pi$ enforces semantic equivalence. The model is interpreted as a feedback system integrating transformation, observation, and stabilization. A categorical formulation is introduced to capture compositional structure, while the associated dynamics are analyzed through fixed-point arguments and contraction conditions ensuring stability. To demonstrate the operational character of the framework, a computational illustration is provided, together with a qualitative analysis of the induced dynamics. A concrete linguistic application shows how context-dependent interpretation can be modeled as a trajectory toward a stable semantic class. The proposed approach connects dynamical systems, category theory, and cognitive modeling, and provides a unified representation of cognition as a feedback-driven process evolving toward invariant interpretations.

Inference Time Context Sparsity: Illusion or Opportunity?

arXiv:2605.24168v1 Announce Type: new Abstract: Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position is that these constraints are artificial and unnecessary, and that the future of LLM inference lies in extreme but principled sparsity along the context dimension. This position is supported by several strands of empirical and theoretical evidence. First, we find the insistence on dense attention unreasonable, since in a long context a query effectively projects O(N) attention information into a hidden space of dimension d

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

arXiv:2605.24219v2 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate only the final output, missing failures that originate in intermediate Thought-Action-Observation steps. We present Trajel, a dataset and evaluation framework for auditing trajectory-level hallucinations in multi-agent industrial workflows. Trajel introduces a five-type hallucination taxonomy (factual, referential, logical, procedural, and scope-based) over expert-annotated agent traces from AssetOpsBench. We benchmark supervised detection models at the subtask, trajectory, and long-context levels. Our results show that the most common failure modes are missed by existing benchmarks, that nearly half of hallucinated trajectories involve multiple types at once, and that automated detectors with high binary accuracy still misclassify the subtlest types. Trajectory-aware detection significantly outperforms standard post-hoc verification, making taxonomy-grounded evaluation necessary for safer agentic deployment.

ConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

arXiv:2605.24399v1 Announce Type: new Abstract: Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational pathology, for complex tumor subtypes where morphology alone can be challenging to distinguish, pathology reports and molecular measurements may provide additional diagnostic evidence alongside whole-slide images, yet existing models often fail to clarify how diverse signals assemble into recognizable diagnostic concepts. We propose ConceptM$^3$oE (Concept Multimodal MoE), which embeds concept formation directly within interaction-aware mixture-of-experts (MoE) pathways. The architecture decomposes evidence into modality-specific, redundant, and synergistic experts, which are then projected into structured concept bottlenecks mapping latent features to a hierarchy of morphology and biomarker concepts. To prevent the information loss typical of interpretable bottlenecks, we utilize residual pathways within each expert to allow task-relevant signals to flow both through the concepts and directly to the final task prediction, so that high performance is maintained alongside interpretability. Across an institutional pediatric brain tumor cohort and a public glioma cohort, the framework delivers competitive performance to unconstrained models while producing reasoning traces validated by an independent neuropathologist. In data-limited regimes, ConceptM$^3$oE improves limited-data performance, increasing macro-F1 from 56.41% to 66.70% at small training sizes compared to non-concept-informed baselines, while also showing faster training convergence consistent with the regularizing effect of concept learning. This work offers a scalable path toward high-performance medical AI that is inherently verifiable and better aligned with the complex decision-making of clinical practice.

Advancing Graph Few-Shot Learning via In-Context Learning

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning. However, existing methods often face two key limitations. First, the predominant graph few-shot learning paradigm relies on supervised tasks, failing to leverage the vast number of unlabeled nodes in the graph. Second, many approaches require complex task adaptation or fine-tuning during inference, limiting their efficiency and applicability. Inspired by the powerful in-context learning capabilities of large language models, we propose a novel model named VISION for adVancIng graph few-Shot learning via In-cOntext LearNing to address these challenges. Our model reframes graph few-shot learning as a fine-tuning-free sequence reasoning problem. At its core is a context-aware network that initializes nodes with role embeddings and employs a dual-context fusion module to synergistically integrate local topological structures and global task-level dependencies. This allows our model to dynamically generate class-aware representations for the query set conditioned on the support set context in a single forward pass. To effectively train our model, we introduce an unsupervised task generator that creates structure-adaptive features and constructs diverse pseudo-tasks from abundant unlabeled data. Our method unifies unsupervised meta-learning with graph in-context learning, achieving efficient inference. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our model. Our public code can be found

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

arXiv:2605.24484v1 Announce Type: new Abstract: Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing solvers are typically limited to symmetric settings or degrade in performance when switching to asymmetric settings due to input inconsistencies or inherent structural differences, substantially limiting their practicality in real-world scenarios that encompass both scenarios. To address this limitation, we define the spatial position of each node based on the relative distances to a specific set of pivots and further propose a Spatial Pivot-Aligned Coordinate-free Embedding (SPACE) framework that unifies node representation and solution generation across symmetric and asymmetric VRPs. Specifically, we construct a bidirectional Frechet representation using a novel furthest pivot sampling strategy to enable invariant node representations across distinct problem settings. Furthermore, we introduce a weight-decomposed adaptive decoding mechanism that decouples geometric perception from problem representations, mitigating the overfitting of constraint decisions to a specific geometry setting. Extensive experiments on 110 VRP variants, comprising 55 symmetric problems and their asymmetric counterparts, demonstrate that SPACE achieves promising zero-shot generalization in both symmetric and asymmetric VRPs.

TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

arXiv:2605.24489v1 Announce Type: new Abstract: Enzyme-reaction retrieval is a fundamental problem in computational biology, underpinning enzyme characterization, reaction mechanism elucidation, and the rational design of metabolic pathways and biocatalysts. As a bidirectional task, it entails both enzyme-to-reaction and reaction-to-enzyme mapping. However, existing approaches suffer from poor generalization across tasks and distributions, with performance highly sensitive to dataset splits and substantial asymmetry between retrieval directions. To address these challenges, we present TIGER, a Text-Informed Generalized Enzyme-Reaction Retrieval framework that leverages protein-to-text generation models to distill textual semantic knowledge from enzyme sequences, providing a generalized representation that bridges enzymes and biochemical reactions. To ensure the quality and reliability of textual semantics, we design a Dynamic Gating Network that adaptively fuses text-derived knowledge with sequence features, enabling more consistent and informative enzyme representations, while a Structure-Shared Feature Projector aligns enzyme and reaction representations within a unified latent space. Extensive experiments demonstrate that, under bidirectional retrieval supervision, TIGER significantly outperforms state-of-the-art baselines across diverse distributions and exhibits strong robustness and transferability across tasks.

Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems

arXiv:2605.24490v1 Announce Type: new Abstract: Multi-agent LLM decision systems for portfolio management still lack a principled way to assign credit across specialist agents, remain vulnerable to cold-start dominance under regime shifts, and offer limited transparency into how final allocations are formed. We propose Market Regime Council (MRC), a cooperative multi-agent decision system that computes exact Shapley credits across all single, pairwise, and Grand-coalition outputs for online agent weighting. Instantiated with N=3 specialist agents, at each trading period, MRC recomputes coalition-based Shapley weights from exponentially weighted performance histories, uses a Bayesian adaptive mixture to stabilize early periods, applies regime-dependent multipliers to adjust agent authority, and records each rebalance through a five-layer causal trace. Over 1,037 trading days across 13 crypto assets and five seeds, MRC achieves a Sharpe ratio of 1.51 and a cumulative return of 440.1%, ranking first on CR, SR, and IR among active baselines and attaining the lowest MDD among active methods. Ablation results show that the gains come from Shapley-weighted integration across coalition outputs rather than from any single stage in isolation. Code and demo data are included in the supplementary material.

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

arXiv:2605.24497v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new security risks, making them particularly vulnerable to jailbreak attacks. Existing approaches often rely on static CoT templates to elicit harmful outputs, but such fixed designs suffer from limited diversity, adaptability, and effectiveness. To overcome these limitations, we propose an adaptive evolutionary CoT jailbreak framework, called AE-CoT. Specifically, the method first rewrites harmful goals into mild prompts with teacher role-play and decomposes them into semantically coherent reasoning fragments to construct a pool of CoT jailbreak candidates. Then, within a structured representation space, we perform multi-generation evolutionary search, where candidate diversity is expanded through fragment-level crossover and a mutation strategy with an adaptive mutation-rate control mechanism. An independent scoring model provides graded harmfulness evaluations, and high-scoring candidates are further enhanced with a harmful CoT template to induce more destructive generations. Extensive experiments across multiple models and datasets demonstrate the effectiveness of the proposed AE-CoT, consistently outperforming state-of-the-art jailbreak methods.

Hypothesis Generation and Inductive Inference in Children and Language Models

arXiv:2605.24528v1 Announce Type: new Abstract: Real world decision-making requires constructing mental models under uncertainty over evidence, over the underlying causal rules, and over the state of the world itself. Which computational principles underpin human inference under such conditions, and do LLM-based agents exhibit similar behavior given matching constraints? We address these questions using an inductive inference Box Task in which participants, human children and LLM-based agents, infer a latent cause through sequential interaction with an uncertain environment. We formalize this task as program induction with Bayesian particle-based inference, admitting two complementary interpretations: (1) as a constraint satisfaction process over hypotheses, and (2) as a program synthesis problem in which hypotheses are executable programs evaluated against evidence. Using the constraint-based formulation, we show that children's behavior is best explained by a combination of subjective evidence reliability and online hypothesis generation, accounting for both their evidence-seeking patterns and their dissociation between task completion and rule generalization. Using the program synthesis formulation, we treat LLM-based agents as model organisms: controllable systems that allow systematic manipulation of task conditions. Across backends, LLM-based agents replicate children's responses to changes in evidence reliability and observability, including discounting unreliable evidence, seeking to resolve partial information, and dissociating between task completion and causal generalization. At the same time, LLM-based agents tend to over-observe and over-comply with instructions relative to children. These results suggest that while children and LLM-based agents adapt similarly to environmental structure, their information-seeking behavior exhibits distinct underlying costs and inductive biases.

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on constructed benchmarks to assess safety from predefined perspectives, or employ dynamic red-teaming to probe potential vulnerabilities. While effective, these approaches face challenges, as they depend heavily on expert domain knowledge, offer limited systematic guarantees, and are vulnerable to rapid obsolescence. To address these limitations, we introduce a novel framework POLARIS that brings the rigor of specification-based software testing to AI safety. POLARIS first compiles unstructured natural-language policies into First-Order Logic (FOL) representations, establishing a traceable link between high-level rules and concrete test cases. This formalization enables the construction of a Semantic Policy Graph, where complex policy violation scenarios are encoded as traversable paths. By systematically exploring this graph, POLARIS uncovers compositional violation patterns, which are then instantiated into executable natural-language test queries, enabling coverage-driven and reproducible safety testing. Experiments demonstrate that POLARIS achieves higher policy coverage and attack success counts compared to established baselines. Crucially, by bridging formal methods and AI safety, POLARIS provides a principled, automated approach to ensuring LLMs adhere to safety-critical policies with verifiable traceability. We release our code at https://github.com/huac-lxy/POLARIS.
  • ✇cs.AI, q-bio.NC updates on arXiv.org
  • Energy Shields for Fairness Filip Cano · Thomas A. Henzinger · Konstantin Kueffner
    arXiv:2605.24926v1 Announce Type: new Abstract: Runtime fairness is not a one-time constraint but a dynamic property evaluated over a sequence of decisions. To ensure fairness at runtime, it is necessary to account for past decisions, information neglected by conventional, static classifiers. Traditional fairness shields enforce runtime fairness abruptly, by intervening \emph{deterministically} whenever a sequence of decisions violates the target for a running fairness measure. This motiv
     

Energy Shields for Fairness

arXiv:2605.24926v1 Announce Type: new Abstract: Runtime fairness is not a one-time constraint but a dynamic property evaluated over a sequence of decisions. To ensure fairness at runtime, it is necessary to account for past decisions, information neglected by conventional, static classifiers. Traditional fairness shields enforce runtime fairness abruptly, by intervening \emph{deterministically} whenever a sequence of decisions violates the target for a running fairness measure. This motivates our \emph{main conceptual contribution: \textbf{energy shields}.} An energy shield is a novel, lightweight, adaptive controller that monitors a sequence of decisions and intervenes \emph{probabilistically} to ensure runtime fairness smoothly, by utilizing physics-inspired energy functions to nudge the sequence toward fairness: the more unfair the decisions, the stronger the nudging force becomes. This makes energy shields the \emph{\textbf{first}} fairness shields to provide both \emph{short-term safety and long-term liveness guarantees}. Safety ensures that the running fairness measure stays within a running target interval with high probability, and liveness ensures that the limit of the fairness measure lies within the limit target interval. Intuitively, the short-term specifies the tolerated fairness values and the long-term specifies the desired fairness values. We also provide a synthesis procedure for constructing the least intrusive energy shield for a given target specification, and demonstrate its efficiency experimentally. We evaluate our energy shields against existing fairness shields through the lens of short- and long-term fairness.
❌