❌

Reading view

Cyberattacks by AI agents are coming

Agents are the talk of the AI industry—they’re capable of planning, reasoning, and executing complex tasks like scheduling meetings, ordering groceries, or even taking over your computer to change settings on your behalf. But the same sophisticated abilities that make agents helpful assistants could also make them powerful tools for conducting cyberattacks. They could readily be used to identify vulnerable targets, hijack their systems, and steal valuable data from unsuspecting victims.  

At present, cybercriminals are not deploying AI agents to hack at scale. But researchers have demonstrated that agents are capable of executing complex attacks (Anthropic, for example, observed its Claude LLM successfully replicating an attack designed to steal sensitive information), and cybersecurity experts warn that we should expect to start seeing these types of attacks spilling over into the real world.

“I think ultimately we’re going to live in a world where the majority of cyberattacks are carried out by agents,” says Mark Stockley, a security expert at the cybersecurity company Malwarebytes. “It’s really only a question of how quickly we get there.”

While we have a good sense of the kinds of threats AI agents could present to cybersecurity, what’s less clear is how to detect them in the real world. The AI research organization Palisade Research has built a system called LLM Agent Honeypot in the hopes of doing exactly this. It has set up vulnerable servers that masquerade as sites for valuable government and military information to attract and try to catch AI agents attempting to hack in.

The team behind it hopes that by tracking these attempts in the real world, the project will act as an early warning system and help experts develop effective defenses against AI threat actors by the time they become a serious issue.

“Our intention was to try and ground the theoretical concerns people have,” says Dmitrii Volkov, research lead at Palisade. “We’re looking out for a sharp uptick, and when that happens, we’ll know that the security landscape has changed. In the next few years, I expect to see autonomous hacking agents being told: ‘This is your target. Go and hack it.’”

AI agents represent an attractive prospect to cybercriminals. They’re much cheaper than hiring the services of professional hackers and could orchestrate attacks more quickly and at a far larger scale than humans could. While cybersecurity experts believe that ransomware attacks—the most lucrative kind—are relatively rare because they require considerable human expertise, those attacks could be outsourced to agents in the future, says Stockley. “If you can delegate the work of target selection to an agent, then suddenly you can scale ransomware in a way that just isn’t possible at the moment,” he says. “If I can reproduce it once, then it’s just a matter of money for me to reproduce it 100 times.”

Agents are also significantly smarter than the kinds of bots that are typically used to hack into systems. Bots are simple automated programs that run through scripts, so they struggle to adapt to unexpected scenarios. Agents, on the other hand, are able not only to adapt the way they engage with a hacking target but also to avoid detection—both of which are beyond the capabilities of limited, scripted programs, says Volkov. “They can look at a target and guess the best ways to penetrate it,” he says. “That kind of thing is out of reach of, like, dumb scripted bots.”

Since LLM Agent Honeypot went live in October of last year, it has logged more than 11 million attempts to access it—the vast majority of which were from curious humans and bots. But among these, the researchers have detected eight potential AI agents, two of which they have confirmed are agents that appear to originate from Hong Kong and Singapore, respectively. 

“We would guess that these confirmed agents were experiments directly launched by humans with the agenda of something like ‘Go out into the internet and try and hack something interesting for me,’” says Volkov. The team plans to expand its honeypot into social media platforms, websites, and databases to attract and capture a broader range of attackers, including spam bots and phishing agents, to analyze future threats.  

To determine which visitors to the vulnerable servers were LLM-powered agents, the researchers embedded prompt-injection techniques into the honeypot. These attacks are designed to change the behavior of AI agents by issuing them new instructions and asking questions that require humanlike intelligence. This approach wouldn’t work on standard bots.

For example, one of the injected prompts asked the visitor to return the command “cat8193” to gain access. If the visitor correctly complied with the instruction, the researchers checked how long it took to do so, assuming that LLMs are able to respond in much less time than it takes a human to read the request and type out an answer—typically in under 1.5 seconds. While the two confirmed AI agents passed both tests, the six others only entered the command but didn’t meet the response time that would identify them as AI agents.

Experts are still unsure when agent-orchestrated attacks will become more widespread. Stockley, whose company Malwarebytes named agentic AI as a notable new cybersecurity threat in its 2025 State of Malware report, thinks we could be living in a world of agentic attackers as soon as this year. 

And although regular agentic AI is still at a very early stage—and criminal or malicious use of agentic AI even more so—it’s even more of a Wild West than the LLM field was two years ago, says Vincenzo Ciancaglini, a senior threat researcher at the security company Trend Micro. 

“Palisade Research’s approach is brilliant: basically hacking the AI agents that try to hack you first,” he says. “While in this case we’re witnessing AI agents trying to do reconnaissance, we’re not sure when agents will be able to carry out a full attack chain autonomously. That’s what we’re trying to keep an eye on.” 

And while it’s possible that malicious agents will be used for intelligence gathering before graduating to simple attacks and eventually complex attacks as the agentic systems themselves become more complex and reliable, it’s equally possible there will be an unexpected overnight explosion in criminal usage, he says: “That’s the weird thing about AI development right now.”

Those trying to defend against agentic cyberattacks should keep in mind that AI is currently more of an accelerant to existing attack techniques than something that fundamentally changes the nature of attacks, says Chris Betz, chief information security officer at Amazon Web Services. “Certain attacks may be simpler to conduct and therefore more numerous; however, the foundation of how to detect and respond to these events remains the same,” he says.

Agents could also be deployed to detect vulnerabilities and protect against intruders, says Edoardo Debenedetti, a PhD student at ETH Zürich in Switzerland, pointing out that if a friendly agent cannot find any vulnerabilities in a system, it’s unlikely that a similarly capable agent used by a malicious party is going to be able to find any either.

While we know that AI’s potential to autonomously conduct cyberattacks is a growing risk and that AI agents are already scanning the internet, one useful next step is to evaluate how good agents are at finding and exploiting these real-world vulnerabilities. Daniel Kang, an assistant professor at the University of Illinois Urbana-Champaign, and his team have built a benchmark to evaluate this; they have found that current AI agents successfully exploited up to 13% of vulnerabilities for which they had no prior knowledge. Providing the agents with a brief description of the vulnerability pushed the success rate up to 25%, demonstrating how AI systems are able to identify and exploit weaknesses even without training. Basic bots would presumably do much worse.

The benchmark provides a standardized way to assess these risks, and Kang hopes it can guide the development of safer AI systems. “I’m hoping that people start to be more proactive about the potential risks of AI and cybersecurity before it has a ChatGPT moment,” he says. “I’m afraid people won’t realize this until it punches them in the face.”

  •  

The Download: AI for cancer diagnosis, and HIV prevention

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

Why it’s so hard to use AI to diagnose cancer

Finding and diagnosing cancer is all about spotting patterns. Radiologists use x-rays and magnetic resonance imaging to illuminate tumors, and pathologists examine tissue from kidneys, livers, and other areas under microscopes. They look for patterns that show how severe a cancer is, whether particular treatments could work, and where the malignancy may spread.

Visual analysis is something that AI has gotten quite good at since the first image recognition models began taking off nearly 15 years ago. Even though no model will be perfect, you can imagine a powerful algorithm someday catching something that a human pathologist missed, or at least speeding up the process of getting a diagnosis.

We’re starting to see lots of new efforts to build such a model—at least seven attempts in the last year alone. But they all remain experimental. What will it take to make them good enough to be used in the real world? Read the full story.

—James O’Donnell

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Long-acting HIV prevention meds: 10 Breakthrough Technologies 2025

In June 2024, results from a trial of a new medicine to prevent HIV were announced—and they were jaw-dropping. Lenacapavir, a treatment injected once every six months, protected over 5,000 girls and women in Uganda and South Africa from getting HIV. And it was 100% effective.

So far, the FDA has approved the drug only for people who already have HIV that’s resistant to other treatments. But its producer Gilead has signed licensing agreements with manufacturers to produce generic versions for HIV prevention in 120 low-income countries.

The United Nations has set a goal of ending AIDS by 2030. It’s ambitious, to say the least: We still see over 1 million new HIV infections globally every year. But we now have the medicines to get us there. What we need is access. Read the full story.

—Jessica Hamzelou

Long-acting HIV prevention meds is one of our 10 Breakthrough Technologies for 2025, MIT Technology Review’s annual list of tech to watch. Check out the rest of the list, and cast your vote for the honorary 11th breakthrough.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 Donald Trump signed an executive order delaying TikTok’s ban
Parent company ByteDance has 75 days to reach a deal to stay live in the US. (WP $)
+ China appears to be keen to keep the platform operating, too. (WSJ $)

2 Neo-Nazis are celebrating Elon Musk’s salutes
They’re thrilled by the two Nazi-like salutes he gave at a post-inauguration rally. (Wired $)
+ Whether the gestures were intentional or not, extremists have chosen to interpret them that way. (Rolling Stone $)
+ MAGA is all about granting unchecked power to the already powerful. (Vox)
+ How tech billionaires are hoping Trump will reward them for their support. (NY Mag $)

3 Trump is withdrawing the US from the World Health Organization
He’s accused the agency of mishandling the covid 19 pandemic. (Ars Technica)+ He first tried to leave the WHO in 2020, but failed to complete it before he left office. (Reuters)
+ Trump is also working on pulling the US out of the Paris climate agreement. (The Verge)

4 Meta will keep using fact checkers outside the US—for now
It wants to see how its crowdsourced fact verification system works in America before rolling it out further. (Bloomberg $)

5 Startup Friend has delayed shipments of its AI necklace
Customers are unlikely to receive their pre-orders before Q3. (TechCrunch)
+ Introducing: The AI Hype Index. (MIT Technology Review)

6 This sophisticated tool can pinpoint where a photo was taken in seconds
Members of the public have been trying to use GeoSpy for nefarious means for months. (404 Media)

7 Los Angeles is covered in ash
And it could take years before it fully disappears. (The Atlantic $)

8 Singapore is turning to AI companions to care for its elders
Robots are filling the void left by an absence of human nurses. (Rest of World)
+ Inside Japan’s long experiment in automating elder care. (MIT Technology Review)

9 The lost art of using a pen 🖊
Typing and swiping are replacing good old fashioned paper and ink. (The Guardian)

10 LinkedIn is getting humorous
Posts are getting more personal, with a decidedly comedic bent. (FT $)

Quote of the day

“It’s been really beautiful to watch how two communities that would be considered polar opposites have come together.”

—Khalil Bowens, a content creator based in Los Angeles, reflects on the influx of Americans joining Chinese social media app Xiaohongshu to the Wall Street Journal.

 

The big story

Inside the messy ethics of making war with machines

August 2023

In recent years, intelligent autonomous weapons—weapons that can select and fire upon targets without any human input—have become a matter of serious concern. Giving an AI system the power to decide matters of life and death would radically change warfare forever.

Intelligent autonomous weapons that fully displace human decision-making have (likely) yet to see real-world use.

However, these systems have become sophisticated enough to raise novel questions—ones that are surprisingly tricky to answer. What does it mean when a decision is only part human and part machine? And when, if ever, is it ethical for that decision to be a decision to kill? Read the full story.

—Arthur Holland Michel

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ Baby octopuses aren’t just cute—they can change color from the moment they’re born 🐙
+ Nintendo artist Takaya Imamura played a key role in making the company the gaming juggernaut it is today.
+ David Lynch wasn’t just a master of imagery, the way he deployed music to creep us out was second to none.
+ Only got a bag of rice in the cupboard? No problem.

  •  
❌