❌

Reading view

Exclusive eBook: How AGI Became a Consequential Conspiracy Theory

In this exclusive subscriber-only eBook, you’ll learn about how the idea that machines will be as smart as—or smarter than—humans has hijacked an entire industry.

by Will Douglas Heaven October 30, 2025

Table of Contents:

  • How Silicon Valley got AGI-pilled
  • The great AGI conspiracy
  • How AGI hijacked an industry
  • The great AGI conspiracy, concluded

Related Stories:

Access all subscriber-only eBooks:

  •  

Human-AI Co-design for Clinical Prediction Models

arXiv:2601.09072v1 Announce Type: new Abstract: Developing safe, effective, and practically useful clinical prediction models (CPMs) traditionally requires iterative collaboration between clinical experts, data scientists, and informaticists. This process refines the often small but critical details of the model building process, such as which features/patients to include and how clinical categories should be defined. However, this traditional collaboration process is extremely time- and resource-intensive, resulting in only a small fraction of CPMs reaching clinical practice. This challenge intensifies when teams attempt to incorporate unstructured clinical notes, which can contain an enormous number of concepts. To address this challenge, we introduce HACHI, an iterative human-in-the-loop framework that uses AI agents to accelerate the development of fully interpretable CPMs by enabling the exploration of concepts in clinical notes. HACHI alternates between (i) an AI agent rapidly exploring and evaluating candidate concepts in clinical notes and (ii) clinical and domain experts providing feedback to improve the CPM learning process. HACHI defines concepts as simple yes-no questions that are used in linear models, allowing the clinical AI team to transparently review, refine, and validate the CPM learned in each round. In two real-world prediction tasks (acute kidney injury and traumatic brain injury), HACHI outperforms existing approaches, surfaces new clinically relevant concepts not included in commonly-used CPMs, and improves model generalizability across clinical sites and time periods. Furthermore, HACHI reveals the critical role of the clinical AI team, such as directing the AI agent to explore concepts that it had not previously considered, adjusting the granularity of concepts it considers, changing the objective function to better align with the clinical objectives, and identifying issues of data bias and leakage.
  •  

Companion Agents: A Table-Information Mining Paradigm for Text-to-SQL

arXiv:2601.08838v1 Announce Type: cross Abstract: Large-scale Text-to-SQL benchmarks such as BIRD typically assume complete and accurate database annotations as well as readily available external knowledge, which fails to reflect common industrial settings where annotations are missing, incomplete, or erroneous. This mismatch substantially limits the real-world applicability of state-of-the-art (SOTA) Text-to-SQL systems. To bridge this gap, we explore a database-centric approach that leverages intrinsic, fine-grained information residing in relational databases to construct missing evidence and improve Text-to-SQL accuracy under annotation-scarce conditions. Our key hypothesis is that when a query requires multi-step reasoning over extensive table information, existing methods often struggle to reliably identify and utilize the truly relevant knowledge. We therefore propose to "cache" query-relevant knowledge on the database side in advance, so that it can be selectively activated at inference time. Based on this idea, we introduce Companion Agents (CA), a new Text-to-SQL paradigm that incorporates a group of agents accompanying database schemas to proactively mine and consolidate hidden inter-table relations, value-domain distributions, statistical regularities, and latent semantic cues before query generation. Experiments on BIRD under the fully missing evidence setting show that CA recovers +4.49 / +4.37 / +14.13 execution accuracy points on RSL-SQL / CHESS / DAIL-SQL, respectively, with larger gains on the Challenging subset +9.65 / +7.58 / +16.71. These improvements stem from CA's automatic database-side mining and evidence construction, suggesting a practical path toward industrial-grade Text-to-SQL deployment without reliance on human-curated evidence.
  •  

Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents

arXiv:2601.08841v1 Announce Type: cross Abstract: The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this study, we explore how structured knowledge, specifically, subject-predicate-object triples, can enhance the clustering and classification of scientific papers. We propose a modular pipeline that combines unsupervised clustering and supervised classification over multiple document representations: raw abstracts, extracted triples, and hybrid formats that integrate both. Using a filtered arXiv corpus, we extract relational triples from abstracts and construct four text representations, which we embed using four state-of-the-art transformer models: MiniLM, MPNet, SciBERT, and SPECTER. We evaluate the resulting embeddings with KMeans, GMM, and HDBSCAN for unsupervised clustering, and fine-tune classification models for arXiv subject prediction. Our results show that full abstract text yields the most coherent clusters, but that hybrid representations incorporating triples consistently improve classification performance, reaching up to 92.6% accuracy and 0.925 macro-F1. We also find that lightweight sentence encoders (MiniLM, MPNet) outperform domain-specific models (SciBERT, SPECTER) in clustering, while SciBERT excels in structured-input classification. These findings highlight the complementary benefits of combining unstructured text with structured knowledge, offering new insights into knowledge-infused representations for semantic organization of scientific documents.
  •  

A Marketplace for AI-Generated Adult Content and Deepfakes

arXiv:2601.09117v1 Announce Type: cross Abstract: Generative AI systems increasingly enable the production of highly realistic synthetic media. Civitai, a popular community-driven platform for AI-generated content, operates a monetized feature called Bounties, which allows users to commission the generation of content in exchange for payment. To examine how this mechanism is used and what content it incentivizes, we conduct a longitudinal analysis of all publicly available bounty requests collected over a 14-month period following the platform's launch. We find that the bounty marketplace is dominated by tools that let users steer AI models toward content they were not trained to generate. At the same time, requests for content that is "Not Safe For Work" are widespread and have increased steadily over time, now comprising a majority of all bounties. Participation in bounty creation is uneven, with 20% of requesters accounting for roughly half of requests. Requests for "deepfake" - media depicting identifiable real individuals - exhibit a higher concentration than other types of bounties. A nontrivial subset of these requests involves explicit deepfakes despite platform policies prohibiting such content. These bounties disproportionately target female celebrities, revealing a pronounced gender asymmetry in social harm. Together, these findings show how monetized, community-driven generative AI platforms can produce gendered harms, raising questions about consent, governance, and enforcement.
  •  

Global Benchmark Database

arXiv:2405.10045v3 Announce Type: replace-cross Abstract: This paper presents Global Benchmark Database (GBD), a comprehensive suite of tools for provisioning and sustainably maintaining benchmark instances and their metadata. The availability of benchmark metadata is essential for many tasks in empirical research, e.g., for the data-driven compilation of benchmarks, the domain-specific analysis of runtime experiments, or the instance-specific selection of solvers. In this paper, we introduce the data model of GBD as well as its interfaces and provide examples of how to interact with them. We also demonstrate the integration of custom data sources and explain how to extend GBD with additional problem domains, instance formats and feature extractors.
  •  

Exploring the Secondary Risks of Large Language Models

arXiv:2506.12382v4 Announce Type: replace-cross Abstract: Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
  •  

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p
  •  

Circulating metabolites, genetics and lifestyle factors in relation to future risk of type 2 diabetes

Nat Med. 2026 Jan 14. doi: 10.1038/s41591-025-04105-8. Online ahead of print.

ABSTRACT

The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.

PMID:41535386 | DOI:10.1038/s41591-025-04105-8

  •  

Multi-omics to study chronic respiratory diseases and viral infections

Eur Respir Rev. 2026 Jan 14;35(179):240286. doi: 10.1183/16000617.0286-2024. Print 2026 Jan.

ABSTRACT

Despite recent advances, the underlying mechanisms of the development and progression of many chronic respiratory diseases remain to be elucidated. Factors such as heterogeneity and complexity of human diseases and difficulty interpreting large datasets hinder research into chronic respiratory diseases. Omics assesses the changes in specific biological entities, such as mRNA expression, epigenetics/epigenomics, genomics, proteomics, metagenomics and metabolomics, and provides valuable insights into the roles of these processes in chronic respiratory diseases. High-throughput omics at bulk, single-cell and spatial levels empower the exploration of disease-related changes through untargeted data-driven statistical methods. Multi-omics is the exploration and integration of multiple biological processes, which compared to a single-omics, can provide a substantially greater and more holistic overview of the pathogenic mechanisms that underpin complex diseases. Multi-omics analysis can comprehensively characterise the mechanisms that drive chronic respiratory diseases, capturing unique biological signatures and cellular interactions at different omics levels. Use of these methods has begun to identify key factors and biomarkers in chronic respiratory diseases. Here, we review current omics approaches and highlight recent advances in respiratory research achieved using multi-omics and integrative methods. Our review provides a valuable resource for researchers and clinicians in this area.

PMID:41534886 | DOI:10.1183/16000617.0286-2024

  •  

Complement-secreting CAFs are associated with better prognosis in pancreatic cancer: single-cell multiomics

Gut. 2026 Jan 13:gutjnl-2025-335683. doi: 10.1136/gutjnl-2025-335683. Online ahead of print.

ABSTRACT

BACKGROUND: Accumulating evidence has demonstrated that distinct tumour-promoting and tumour-restraining cancer-associated fibroblast (CAF) subtypes coexist in pancreatic ductal adenocarcinoma.

OBJECTIVE: To develop targeted CAF therapeutic strategies by reprogramming tumour-promoting CAF subtypes.

DESIGN: We leveraged multiomics technologies to systematically identify and characterise CAF subtypes transcriptionally, epigenetically and spatially and correlate them with clinicopathological features.

RESULTS: We found that complement-secreting CAFs (csCAFs), initially identified by our group and inflammatory CAFs (iCAFs) share significant overlap in their transcriptional profiles and chromatin accessibility. iCAFs specifically express transcription factors from the heme and oxidative homeostasis pathway and the activator protein 1 family, which are both involved in cellular response to oxidative stress. Notably, the composition of csCAFs among all CAFs declined during pancreatic carcinogenesis, while trajectory analysis showed that csCAFs could potentially differentiate into iCAFs. Spatially resolved analysis indicated that tumour regions with a higher csCAF composition were associated with lower levels of TGF-β ligands, fewer M2 tumour-associated macrophages and increased levels of lipid mediators. Additionally, we identified a spatially defined CXCL12-CXCR4 ligand-receptor interaction between csCAFs and T cells, but in distinct patterns between different metastatic organs. Patients with a higher composition of csCAFs have significantly longer overall survival and recurrence-free survival through multiplex immunohistochemistry and bulk RNA-seq deconvolution.

CONCLUSION: Our study demonstrates that csCAFs may represent an early-stage iCAF subtype and suggests a promising strategy for reprogramming iCAFs into csCAFs.

PMID:41534892 | DOI:10.1136/gutjnl-2025-335683

  •  

Evidence for Digital Health Tools Designed to Support the Triage of Musculoskeletal Conditions in Primary, Urgent, and Emergency Care Settings: Scoping Review

Background: The digital health research field is growing rapidly, and a summary of the available digital tools for triaging musculoskeletal conditions is needed. Effective and safe digital triage tools for musculoskeletal conditions could support patients in making informed care decisions, aid clinicians and patients in navigating care, and may contribute to reducing ED overcrowding and healthcare costs. Objective: To identify and describe digital health tools for use by adults to triage musculoskeletal conditions across primary, urgent, or emergency care settings. Methods: Our scoping review was conducted following the Johanna Briggs Institute recommendations for scoping reviews and Arksey & O’Malley’s framework. Systematic searches in MEDLINE (OVID), CINAHL (EBSCO), PsycINFO (EBSCO), Embase (OVID), Cochrane Library, Web of Science, OpenGrey, GoogleScholar, arXiv.org, medRxiv.org, and an extensive grey literature search were conducted with a librarian scientist from inception to Sept 18, 2025. Studies had to recruit adults (18+ years) with musculoskeletal conditions that identified a digital health tool designed to triage or diagnose in primary, urgent, or emergency care settings and report primary data to be included. Two reviewer pairs independently screened abstracts and full-text articles Relevant data were extracted in duplicate, and results were summarized descriptively. Results: The search yielded 5695 records, and we screened 189 full-text articles. Thirty-four studies (n=37,509 patients) met the inclusion criteria. The most common musculoskeletal conditions reported were rheumatoid/inflammatory arthritis (n=13, 38%). Nineteen (59%) studies reported on symptom checkers, 13 (44%) studies on triage/diagnosis tools, and 2 (6%) were studies of diagnostic predictor tools. There were 16 unique digital health tools. Two tools were built for triaging musculoskeletal conditions, and were not publicly available outside the UK National Health Service. Most tools were generic tools designed to screen for general health problems, including musculoskeletal conditions. The most common approach to evaluating performance (eg, accuracy) of the tools was to compare the concordance of the tool to a clinician diagnosis or triage recommendation. Sensitivity and specificity ranged from 39%-91% and 23%-80%, respectively. Reported accuracy of included tools ranged from 33% to 98%. Conclusions: Musculoskeletal conditions remain a blind spot for people designing, implementing, and evaluating digital health for triage: few tools were specifically designed for musculoskeletal conditions, and most existing tools performed poorly when applied to musculoskeletal populations. The evidence base supporting accuracy of digital health for triaging and diagnosing musculoskeletal conditions is weak, and tool performance was inconsistent and lacking transparency. We recommend health systems and clinicians use a multi-modal approach, integrating both digital health tools and clinical decision-making to safely triage and diagnose until a more robust tool for musculoskeletal conditions is available. Future tool developers need to use transparent, standardized processes that prioritize tool safety, clinical value, and trustworthiness when designing for clinicians and patients.
  •  

From Agents to Governance: Essential AI Skills for Clinicians in the Large Language Model Era

Large language models are rapidly transitioning from pilot schemes to routine clinical practice. This creates an urgent need for clinicians to develop the necessary skills to strike the right balance between seizing opportunities and taking accountability. We propose a 3-tier competency framework to support clinicians’ evolution from cautious users to responsible stewards of artificial intelligence (AI). Tier 1 (foundational skills) defines the minimum competencies for safe use, including prompt engineering, human–AI agent interaction, security and privacy awareness, and the clinician-patient interface (transparency and consent). Tier 2 (intermediate skills) emphasizes evaluative expertise, including bias detection and mitigation, interpretation of explainability outputs, and the effective clinical integration of AI-generated workflows. Tier 3 (advanced skills) establishes leadership capabilities, mandating competencies in ethical governance (delineating accountability and liability boundaries), regulatory strategy, and model life cycle management—specifically, the ability to govern algorithmic adaptation and change protocols. Integrating this framework into continuing medical education programs and role-specific job descriptions could enhance clinicians’ ability to use AI safely and responsibly. This could standardize deployment and support safer clinical practice, with the potential to improve patient outcomes.
  •  

Africa’s Digital Health Revolution: The Digital Fit-Viability Model to Move From Innovation to Scaled Implementation

Digital innovations hold immense potential to transform health care delivery, particularly in sub-Saharan Africa, where financial, geographical, and infrastructural constraints continue to hinder progress toward universal health care delivery. Although a growing health tech sector offers creative solutions, few digital health interventions reach scaled implementation. In this paper, we present the digital fit/viability model—an adapted determinant framework to describe facilitators and barriers to moving from digital tools to integrated digital health implementation. We then use this model to describe the specific challenges and recommended solutions when developing digital health tools for health systems in sub-Saharan Africa.
  •  

The Download: next-gen nuclear, and the data center backlash

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

How next-generation nuclear reactors break out of the 20th-century blueprint  

The popularity of commercial nuclear reactors has surged in recent years as worries about climate change and energy independence drowned out concerns about meltdowns and radioactive waste. The problem is, building nuclear power plants is expensive and slow.  

A new generation of nuclear power technology could reinvent what a reactor looks like—and how it works. Advocates hope that new tech can refresh the industry and help replace fossil fuels without emitting greenhouse gases.  Here’s what that might look like.

—Casey Crownhart

Next-gen nuclear is one of our 10 Breakthrough Technologies this year. If you want to learn more about why it made the list, sign up to receive The Spark, our weekly newsletter all about energy and climate change, tomorrow. You can also check out the rest of the technologies on the list here.

Data centers are amazing. Everyone hates them.

The hyperscale datacenter is a marvel of our age. A masterstroke of engineering across multiple disciplines. They are nothing short of a technological wonder. People hate them.  

People hate them in Virginia, which leads the nation in their construction. They hate them in Nevada, where they slurp up the state’s precious water. They hate them in Michigan, and Arizona, and South Dakota. They hate them all around the world, it’s true. But they really hate them in Georgia. Read our story about why they’re provoking so much fury. 

—Mat Honan

This story first featured in The Debrief with Mat Honan, a weekly newsletter about the biggest stories in tech from our editor in chief. Sign up here to get the next one in your inbox on Friday.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Iran is systematically crippling Starlink
The satellite internet service is meant to be impossible to jam—but the Iranian authorities are doing just that. (Rest of World)  
+ Messages getting around Iran’s internet block suggest that thousands of people have been killed. (NYT $)
+ On the ground in Ukraine’s largest Starlink repair shop. (MIT Technology Review)

2 Studies claiming microplastics harm us are being called into question
Some scientists say the discoveries are probably the result of contamination and false positives. (The Guardian) 

3 Trump is trying to temper the data center backlash 
He hopes cajoling tech companies to pay more and thus reduce people’s energy bills will do the trick. (WP $) 
+ Microsoft has just become the first tech company to promise it will do just that. (NYT $)
+ We know AI is power hungry. But just how big is the scale of the problem? (MIT Technology Review) 

4 US emissions jumped last year
Thanks to a combination of rising electricity demand, and more coal being burned to meet it. (NYT $)
+ But it’s not all bad news: coal power generation in India and China finally started to decline. (The Guardian)
+ Four bright spots in climate news in 2025. (MIT Technology Review)

5 Elon Musk needs to face consequences for his actions
If we tolerate him unleashing a flood of harassment of women and children, what will come next? (The Atlantic $) 
+ The US Senate has passed a bill that could give non-consensual deepfake victims a new way to fight back. (The Verge $)

6 Why the US is set to lose the race back to the moon 🚀🌔
Cuts to NASA aren’t helping, but they’re not the only problem. (Wired $)

7 Google’s Veo AI model can now turn portrait images into vertical videos
Really slick ones, too. (The Verge $)
+ AI-generated influencers are sharing fake images of them in bed with celebrities on Instagram. (404 Media $)

8 Former NYC mayor Eric Adams has been accused of a crypto ‘pump and dump’ 
He promoted a token that saw its market cap briefly soar to $580 million before plummeting. (Coindesk)

9 Are you a middle manager? Here’s some good news for you
Your skills are not being replaced by AI any time soon. (Quartz) 

10 Even miniscule lifestyle tweaks can extend your lifespan
A study of 60,000 adults found just a little bit more sleep and exercise makes a huge difference. (New Scientist $)
+ Aging hits us in our 40s and 60s. But well-being doesn’t have to fall off a cliff. (MIT Technology Review)

Quote of the day

“What I’m hopeful for in ’26 is for more people speaking up. Speaking truth to power is the point of freedom of speech, is the point of American society.”

—LinkedIn cofounder Reid Hoffman tells Wired he wants more people in Silicon Valley to start pushing back against the Trump administration this year. 

One more thing

two women collaborating on their laptops in a lecture hall
DEEP LEARNING INDABA 2024

What Africa needs to do to become a major AI player

Africa is still early in the process of adopting AI technologies. But researchers say the continent is uniquely hospitable to it for several reasons, including a relatively young and increasingly well-educated population, a rapidly growing ecosystem of AI startups, and lots of potential consumers.  

However, ambitious efforts to develop AI tools that answer the needs of Africans face numerous hurdles. Taken together, researchers worry, they could hold Africa’s AI sector back and hamper its efforts to pave its own pathway in the global AI race. Read the full story.

—Abdullahi Tsanni

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Still keen to do a bit of reflecting on the year behind and the one ahead? This free guide might help!
+ Turns out British comedian Rik Mayall had some pretty solid life advice.
+ I want to stay in this house in São Paolo.  
+ If you want to stop doomscrolling, it’s worth looking at your sleep habits. ($)

  •  

A nowhere-to-hide mechanism ensures complete piRNA-directed DNA methylation

Nature, Published online: 14 January 2026; doi:10.1038/s41586-025-09940-w

In mice, a SPOCD1–TPR-dependent ‘nowhere-to-hide’ mechanism is required for complete non-stochastic piRNA-directed LINE1 DNA methylation by preventing transposons from escaping surveillance within heterochromatin.
  •  

Enhancing telesurgical safety with predictive digital twin synchronization: a framework for latency compensation in robotic surgery

npj Digital Medicine, Published online: 13 January 2026; doi:10.1038/s41746-025-02283-w

Enhancing telesurgical safety with predictive digital twin synchronization: a framework for latency compensation in robotic surgery
  •  
❌