❌

Reading view

The Download: the US digital rights crackdown, and AI companionship

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

What it’s like to be banned from the US for fighting online hate  

Just before Christmas the Trump administration dramatically escalated its war on digital rights by banning five people from entering the US. One of them, Josephine Ballon, is a director of HateAid, a small German nonprofit founded to support the victims of online harassment and violence. The organization is a strong advocate of EU tech regulations, and so finds itself attacked in campaigns from right-wing politicians and provocateurs who claim that it engages in censorship. 

EU officials, freedom of speech experts, and the five people targeted all flatly reject these accusations. Ballon told us that their work is fundamentally about making people feel safer online. But their experiences over the past few weeks show just how politicized and besieged their work in online safety has become. Read the full story. 

—Eileen Guo

TR10: AI companions

Chatbots are skilled at crafting sophisticated dialogue and mimicking empathetic behavior. They never get tired of chatting. It’s no wonder, then, that so many people now use them for companionship—forging friendships or even romantic relationships. 

72% of US teenagers have used AI for companionship, according to a study from the nonprofit Common Sense Media. But while chatbots can provide much-needed emotional support and guidance for some people, they can exacerbate underlying problems in others—especially vulnerable people or those with mental health issues. 

Although some early attempts to regulate this space are underway, AI companionship is going nowhere. Read why we made it one of our 10 Breakthrough Technologies this year, and check out the rest of the list.

And, if you want to learn more about what we predict for AI this year, sign up to join me for our free LinkedIn Live event tomorrow at 12.30pm ET.

Why inventing new emotions feels so good  

Have you ever felt “velvetmist”?  

It’s a “complex and subtle emotion that elicits feelings of comfort, serenity, and a gentle sense of floating.” It’s peaceful, but more ephemeral and intangible than contentment. It might be evoked by the sight of a sunset or a moody, low-key album.  

If you haven’t ever felt this sensation—or even heard of it—that’s not surprising. A Reddit user generated it with ChatGPT, along with advice on how to evoke the feeling. Don’t scoff: Researchers say more and more terms for these “neo-­emotions” are showing up online, describing new dimensions and aspects of feeling. Read our story to learn more about why. 

—Anya Kamenetz

This story is from the latest print issue of MIT Technology Review. If you haven’t already, subscribe now to receive the next edition as soon as it lands (and benefit from some hefty seasonal discounts too!)

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Ads are coming to ChatGPT 
For American users initially, with plans to expand soon. (CNN)
+ Here’s how they’ll work. (Wired $)

2 What will we be able to salvage after the AI bubble bursts? 
It will be ugly, but there are plenty of good uses for AI that we’ll want to keep. (The Guardian) 
+ What even is the AI bubble? (MIT Technology Review)

3 It’s almost impossible to mine Greenland’s natural resources 
It has vast supplies of rare earth elements, but its harsh climate and environment make them very hard to access. (The Week)

4 Iran is now 10 days into its internet shutdown
It’s one of the longest and most extreme we’ve ever witnessed. (BBC)
+  Starlink isn’t proving as helpful as hoped as the regime finds ways to jam it. (Reuters $)
+ Battles are raging online about what’s really going on inside Iran. (NYT $)

5 America is heading for a polymarket disaster 
Prediction markets are getting out of control, and some people are losing a lot of money. (The Atlantic $)
+ They were first embraced by political junkies, but now they’re everywhere. (NYT $)

6 How to fireproof a city 
Californians are starting to fight fires before they can even start. (The Verge $)
+ How AI can help spot wildfires. (MIT Technology Review)

7 Stoking ‘deep state’ conspiracy theories can be dangerous 
Especially if you’re then given the task of helping run one of those state institutions, as Dan Bongino is now learning. (WP $)
+ Why everything is a conspiracy now. (MIT Technology Review)

8 Why we’re suddenly all having a ‘Very Chinese Time’ 🇨🇳
It’s a fun, flippant trend—but it also shows how China’s soft power is growing around the globe. (Wired $) 

9 Why there’s no one best way to store information
Each one involves trade-offs between space and time. (Quanta $)

10 Meat may play a surprising role in helping people reach 100
Perhaps because it can assist with building stronger muscles and bones. (New Scientist $)

Quote of the day

“That’s the level of anxiety now – people watching the skies and the seas themselves because they don’t know what else to do.”

—A Greenlander tells The Guardian just how seriously she and her fellow compatriots are taking Trump’s threat to invade their country. 

One more thing

three silhouetted people in a boat crossing the water in the dark toward a beam of light
KATHERINE LAM

Inside a romance scam compound—and how people get tricked into being there

Gavesh’s journey started, seemingly innocently, with a job ad on Facebook promising work he desperately needed.

Instead, he found himself trafficked into a business commonly known as “pig butchering”—a form of fraud in which scammers form close relationships with targets online and extract money from them. The Chinese crime syndicates behind the scams have netted billions of dollars, and they have used violence and coercion to force their workers, many of them trafficked like Gavesh, to carry out the frauds from large compounds, several of which operate openly in the quasi-lawless borderlands of Myanmar.

Big Tech may hold the key to breaking up the scam syndicates—if these companies can be persuaded or compelled to act. Read the full story.

—Peter Guest & Emily Fishbein

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Blue Monday isn’t real (but it is an absolute banger of a track.) 
+ Some great advice here about how to be productive during the working day.
+ Twelfth Night is one of Shakespeare’s most fun plays—as these top actors can attest. 
+ If the cold and dark gets to you, try making yourself a delicious bowl of soup. 

  •  

AnyECG: Evolved ECG Foundation Model for Holistic Health Profiling

arXiv:2601.10748v1 Announce Type: cross Abstract: Background: Artificial intelligence enabled electrocardiography (AI-ECG) has demonstrated the ability to detect diverse pathologies, but most existing models focus on single disease identification, neglecting comorbidities and future risk prediction. Although ECGFounder expanded cardiac disease coverage, a holistic health profiling model remains needed. Methods: We constructed a large multicenter dataset comprising 13.3 million ECGs from 2.98 million patients. Using transfer learning, ECGFounder was fine-tuned to develop AnyECG, a foundation model for holistic health profiling. Performance was evaluated using external validation cohorts and a 10-year longitudinal cohort for current diagnosis, future risk prediction, and comorbidity identification. Results: AnyECG demonstrated systemic predictive capability across 1172 conditions, achieving an AUROC greater than 0.7 for 306 diseases. The model revealed novel disease associations, robust comorbidity patterns, and future disease risks. Representative examples included high diagnostic performance for hyperparathyroidism (AUROC 0.941), type 2 diabetes (0.803), Crohn disease (0.817), lymphoid leukemia (0.856), and chronic obstructive pulmonary disease (0.773). Conclusion: The AnyECG foundation model provides substantial evidence that AI-ECG can serve as a systemic tool for concurrent disease detection and long-term risk prediction.
  •  

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

arXiv:2601.11505v1 Announce Type: cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D management datasets. Current datasets differ substantially in structure and are time-consuming to access and process, which impedes data integration and reduces the comparability and generalizability of algorithmic developments. This work aims to establish a unified and accessible data resource for T1D algorithm development. Multiple publicly available T1D datasets were consolidated into a unified resource, termed the MetaboNet dataset. Inclusion required the availability of both continuous glucose monitoring (CGM) data and corresponding insulin pump dosing records. Additionally, auxiliary information such as reported carbohydrate intake and physical activity was retained when present. The MetaboNet dataset comprises 3135 subjects and 1228 patient-years of overlapping CGM and insulin data, making it substantially larger than existing standalone benchmark datasets. The resource is distributed as a fully public subset available for immediate download at https://metabo-net.org/ , and with a Data Use Agreement (DUA)-restricted subset accessible through their respective application processes. For the datasets in the latter subset, processing pipelines are provided to automatically convert the data into the standardized MetaboNet format. A consolidated public dataset for T1D research is presented, and the access pathways for both its unrestricted and DUA-governed components are described. The resulting dataset covers a broad range of glycemic profiles and demographics and thus can yield more generalizable algorithmic performance than individual datasets.
  •  

The AI healthcare gold rush is here

AI companies are clustering around healthcare and fast.  In just the past week, OpenAI bought health startup Torch, Anthropic launched Claude for healthcare, and Sam Altman-backed MergeLabs closed a $250 million seed round at an $850 million valuation. The money and products are pouring into health and voice AI, but so are concerns about hallucination risks, inaccurate medical information, and […]
  •  

Cellular neighborhoods in cancer

Nat Cancer. 2026 Jan 16. doi: 10.1038/s43018-025-01107-w. Online ahead of print.

ABSTRACT

The concept of cellular neighborhoods, defined as recurring structures within the tissue with characteristic cell compositions and interactions, has transformed our understanding of the complexity and dynamics of tumor ecosystems. Recent advances in spatial omics and computational modeling have enabled high-resolution mapping of these neighborhoods, providing unprecedented insights into their roles in shaping tumor heterogeneity, evolution and therapeutic responses. Despite these advances, a unified framework for interpreting cellular neighborhoods remains lacking. This Perspective synthesizes emerging concepts and insights, focusing on the definition and classification of cellular neighborhoods in cancer, computational methods for identifying and comparing them, and their clinical relevance.

PMID:41545713 | DOI:10.1038/s43018-025-01107-w

  •  

Contaminating plasmid sequences and disrupted vector genomes in the liver following adeno-associated virus gene therapy

Nature Medicine, Published online: 16 January 2026; doi:10.1038/s41591-025-04073-z

Analyses of liver biopsies from a child with spinal muscular atrophy treated with adeno-associated virus gene therapy who developed hepatitis reveal contaminating manufacturing plasmids and disrupted vector genomes, possibly resulting from recombination events.
  •  

Exclusive eBook: How AGI Became a Consequential Conspiracy Theory

In this exclusive subscriber-only eBook, you’ll learn about how the idea that machines will be as smart as—or smarter than—humans has hijacked an entire industry.

by Will Douglas Heaven October 30, 2025

Table of Contents:

  • How Silicon Valley got AGI-pilled
  • The great AGI conspiracy
  • How AGI hijacked an industry
  • The great AGI conspiracy, concluded

Related Stories:

Access all subscriber-only eBooks:

  •  

Human-AI Co-design for Clinical Prediction Models

arXiv:2601.09072v1 Announce Type: new Abstract: Developing safe, effective, and practically useful clinical prediction models (CPMs) traditionally requires iterative collaboration between clinical experts, data scientists, and informaticists. This process refines the often small but critical details of the model building process, such as which features/patients to include and how clinical categories should be defined. However, this traditional collaboration process is extremely time- and resource-intensive, resulting in only a small fraction of CPMs reaching clinical practice. This challenge intensifies when teams attempt to incorporate unstructured clinical notes, which can contain an enormous number of concepts. To address this challenge, we introduce HACHI, an iterative human-in-the-loop framework that uses AI agents to accelerate the development of fully interpretable CPMs by enabling the exploration of concepts in clinical notes. HACHI alternates between (i) an AI agent rapidly exploring and evaluating candidate concepts in clinical notes and (ii) clinical and domain experts providing feedback to improve the CPM learning process. HACHI defines concepts as simple yes-no questions that are used in linear models, allowing the clinical AI team to transparently review, refine, and validate the CPM learned in each round. In two real-world prediction tasks (acute kidney injury and traumatic brain injury), HACHI outperforms existing approaches, surfaces new clinically relevant concepts not included in commonly-used CPMs, and improves model generalizability across clinical sites and time periods. Furthermore, HACHI reveals the critical role of the clinical AI team, such as directing the AI agent to explore concepts that it had not previously considered, adjusting the granularity of concepts it considers, changing the objective function to better align with the clinical objectives, and identifying issues of data bias and leakage.
  •  

Companion Agents: A Table-Information Mining Paradigm for Text-to-SQL

arXiv:2601.08838v1 Announce Type: cross Abstract: Large-scale Text-to-SQL benchmarks such as BIRD typically assume complete and accurate database annotations as well as readily available external knowledge, which fails to reflect common industrial settings where annotations are missing, incomplete, or erroneous. This mismatch substantially limits the real-world applicability of state-of-the-art (SOTA) Text-to-SQL systems. To bridge this gap, we explore a database-centric approach that leverages intrinsic, fine-grained information residing in relational databases to construct missing evidence and improve Text-to-SQL accuracy under annotation-scarce conditions. Our key hypothesis is that when a query requires multi-step reasoning over extensive table information, existing methods often struggle to reliably identify and utilize the truly relevant knowledge. We therefore propose to "cache" query-relevant knowledge on the database side in advance, so that it can be selectively activated at inference time. Based on this idea, we introduce Companion Agents (CA), a new Text-to-SQL paradigm that incorporates a group of agents accompanying database schemas to proactively mine and consolidate hidden inter-table relations, value-domain distributions, statistical regularities, and latent semantic cues before query generation. Experiments on BIRD under the fully missing evidence setting show that CA recovers +4.49 / +4.37 / +14.13 execution accuracy points on RSL-SQL / CHESS / DAIL-SQL, respectively, with larger gains on the Challenging subset +9.65 / +7.58 / +16.71. These improvements stem from CA's automatic database-side mining and evidence construction, suggesting a practical path toward industrial-grade Text-to-SQL deployment without reliance on human-curated evidence.
  •  

Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents

arXiv:2601.08841v1 Announce Type: cross Abstract: The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this study, we explore how structured knowledge, specifically, subject-predicate-object triples, can enhance the clustering and classification of scientific papers. We propose a modular pipeline that combines unsupervised clustering and supervised classification over multiple document representations: raw abstracts, extracted triples, and hybrid formats that integrate both. Using a filtered arXiv corpus, we extract relational triples from abstracts and construct four text representations, which we embed using four state-of-the-art transformer models: MiniLM, MPNet, SciBERT, and SPECTER. We evaluate the resulting embeddings with KMeans, GMM, and HDBSCAN for unsupervised clustering, and fine-tune classification models for arXiv subject prediction. Our results show that full abstract text yields the most coherent clusters, but that hybrid representations incorporating triples consistently improve classification performance, reaching up to 92.6% accuracy and 0.925 macro-F1. We also find that lightweight sentence encoders (MiniLM, MPNet) outperform domain-specific models (SciBERT, SPECTER) in clustering, while SciBERT excels in structured-input classification. These findings highlight the complementary benefits of combining unstructured text with structured knowledge, offering new insights into knowledge-infused representations for semantic organization of scientific documents.
  •  

A Marketplace for AI-Generated Adult Content and Deepfakes

arXiv:2601.09117v1 Announce Type: cross Abstract: Generative AI systems increasingly enable the production of highly realistic synthetic media. Civitai, a popular community-driven platform for AI-generated content, operates a monetized feature called Bounties, which allows users to commission the generation of content in exchange for payment. To examine how this mechanism is used and what content it incentivizes, we conduct a longitudinal analysis of all publicly available bounty requests collected over a 14-month period following the platform's launch. We find that the bounty marketplace is dominated by tools that let users steer AI models toward content they were not trained to generate. At the same time, requests for content that is "Not Safe For Work" are widespread and have increased steadily over time, now comprising a majority of all bounties. Participation in bounty creation is uneven, with 20% of requesters accounting for roughly half of requests. Requests for "deepfake" - media depicting identifiable real individuals - exhibit a higher concentration than other types of bounties. A nontrivial subset of these requests involves explicit deepfakes despite platform policies prohibiting such content. These bounties disproportionately target female celebrities, revealing a pronounced gender asymmetry in social harm. Together, these findings show how monetized, community-driven generative AI platforms can produce gendered harms, raising questions about consent, governance, and enforcement.
  •  

Global Benchmark Database

arXiv:2405.10045v3 Announce Type: replace-cross Abstract: This paper presents Global Benchmark Database (GBD), a comprehensive suite of tools for provisioning and sustainably maintaining benchmark instances and their metadata. The availability of benchmark metadata is essential for many tasks in empirical research, e.g., for the data-driven compilation of benchmarks, the domain-specific analysis of runtime experiments, or the instance-specific selection of solvers. In this paper, we introduce the data model of GBD as well as its interfaces and provide examples of how to interact with them. We also demonstrate the integration of custom data sources and explain how to extend GBD with additional problem domains, instance formats and feature extractors.
  •  

Exploring the Secondary Risks of Large Language Models

arXiv:2506.12382v4 Announce Type: replace-cross Abstract: Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior research has primarily focused on jailbreak attacks, less attention has been given to non-adversarial failures that subtly emerge during benign interactions. We introduce secondary risks a novel class of failure modes marked by harmful or misleading behaviors during benign prompts. Unlike adversarial attacks, these risks stem from imperfect generalization and often evade standard safety mechanisms. To enable systematic evaluation, we introduce two risk primitives verbose response and speculative advice that capture the core failure patterns. Building on these definitions, we propose SecLens, a black-box, multi-objective search framework that efficiently elicits secondary risk behaviors by optimizing task relevance, risk activation, and linguistic plausibility. To support reproducible evaluation, we release SecRiskBench, a benchmark dataset of 650 prompts covering eight diverse real-world risk categories. Experimental results from extensive evaluations on 16 popular models demonstrate that secondary risks are widespread, transferable across models, and modality independent, emphasizing the urgent need for enhanced safety mechanisms to address benign yet harmful LLM behaviors in real-world deployments.
  •  

GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards

arXiv:2601.08183v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p
  •  

Circulating metabolites, genetics and lifestyle factors in relation to future risk of type 2 diabetes

Nat Med. 2026 Jan 14. doi: 10.1038/s41591-025-04105-8. Online ahead of print.

ABSTRACT

The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.

PMID:41535386 | DOI:10.1038/s41591-025-04105-8

  •  
❌