Cassie Shum discusses why knowledge graphs serve as a critical foundation for agentic systems. Moving beyond basic RAG, she explains 4 practical architectural patterns: context bundling, decision provenance, code as truth, and agent visibility. She demonstrates an engineering harness built on a knowledge graph to streamline feedback loops, optimize token usage, and maintain system reliability. By Cassie Shum
Cassie Shum discusses why knowledge graphs serve as a critical foundation for agentic systems. Moving beyond basic RAG, she explains 4 practical architectural patterns: context bundling, decision provenance, code as truth, and agent visibility. She demonstrates an engineering harness built on a knowledge graph to streamline feedback loops, optimize token usage, and maintain system reliability.
This is todayβs edition of The Download, our weekday newsletter that provides a daily dose of whatβs going on in the world of technology.
Meet the under-35s shaping the future of biotech
Every year, MIT Technology Review puts together our 35 Innovators Under 35, a list of some of the brightest and best young minds working across science and technology. This yearβs honorees include nine people transforming biotech, whose work spans everything from lifesaving innovations to groundbreaking lo
This is todayβs edition of The Download, our weekday newsletter that provides a daily dose of whatβs going on in the world of technology.
Meet the under-35s shaping the future of biotech
Every year, MIT Technology Review puts together our 35 Innovators Under 35, a list of some of the brightest and best young minds working across science and technology. This yearβs honorees include nine people transforming biotech, whose work spans everything from lifesaving innovations to groundbreaking longevity tech.
Their innovations include a βreprogrammingβ therapy that reverses vision loss, tiny brain electrodes inspired by Japanese art, and a personalized gene-editing treatment for a baby with a rare genetic disorder. There are even efforts to design new viruses with generative AI, which (hopefully) will produce new drugs or soak up pollution.
This story is from The Checkup, our weekly biotech newsletter. Sign up to receive it in your inbox every Thursday.
Biotechnology is one of four categories in our 35 Innovators Under 35 list for 2026, featuring young people worldwide doing groundbreaking work in science and technology. Meet the rest of them here, or explore the full list across the AI, computing and robotics, biotechnology, and climate and energy categories.
This founder is making cheaper, cleaner steel
The steel industry isnβt exactly known for innovation. Very little has changed about purifying iron ore since the process was invented and commercialized in the 1850s. But Laureen Meroueh, founder of Hertha Metals, has an idea that could change that.
Meroueh may have found a way to clean up steelmaking without driving up the price. Her new furnace turns iron ore into refined liquid steel in a single step and swaps coal for natural gas. Together, those changes slash emissions by at least half, she says, and cut costs by 25% compared with steelmaking as usual.
Iβve combed the internet to find you todayβs most fun/important/scary/fascinating stories about technology.
1 Anthropic says it has blocked potential plots to build biological weapons The company identified five such cases. (NYT $) + And six cases of using AI to build software for conventional weapons (BBC) + Governments are also using Claude for surveillance.(Axios) + While Russia-linked hackers used it to automate attacks on Ukraine. (Quartz) + The threats were revealed in a new Anthropic report. (Guardian) + Bill Gates says AI needs new guardrails. (MIT Technology Review)
2 California has banned addictive social media features for under-16s The law prohibits infinite scroll and autoplay. (Guardian) + It also introduces new rules for AI and companion chatbots. (Reuters $) + Itβs the first law of its kind in the US. (NYT $) + Social media encourages the worst AI boosterism. (MIT Technology Review)
3 Two AI researchers have left Anthropic and Google over safety risks They left a day after Jacob Coxonβs viral departure from Anthropic. (NBC News) + Elon Musk called their concerns a βsetupβ and a βpsyop.β (Guardian) + AI fears are pushing Congress toward tougher regulation. (WSJ $)
4 Sam Altman is pitching OpenAIβs cyber defenses to power companies The meetings followed reports of AI attacks on critical systems. (Politico $) + Altman also told staff that OpenAI is open to slowing down AI. Bloomberg $)
5 After years of fighting AI, music labels are starting to embrace it Universal is partnering with ElevenLabs on an AI remix platform.(Gizmodo) + AI is complicating definitions of creativity. (MIT Technology Review)
6 Chinese drugmakers are challenging US dominance in weight-loss drugs Theyβre developing hundreds of GLP-1 treatments for global markets. (WSJ $)
7 Electric air taxis have begun official test flights in Texas Theyβre the first flights under the White Houseβs new pilot program. (Verge)
8 Chinese drones are helping to rescue survivors of Nepalβs floods Theyβre delivering food and airlifting bodies from flood-hit areas. (Ars Technica)
9 NASA and IBM have built an AI model to map the moon It could help locate ice and identify safer landing sites. (Register)
10 One man is on a quest to digitally preserve Americaβs public restrooms His Restroom Archive is a museum-style repository of 3D scans. (404 Media)
Quote of the day
βI didnβt ask Facebook to build a profile of my familyβI posted a video of me singing in the car with my kids.βΒ
βKalie Roberts, a travel content creator, says in an Instagram reel that Meta AI used years of Facebook posts to piece together her childrenβs identities and pinpoint where her family lives.
One more thing
Chinese tech workers are starting to train their AI doublesβand pushing back
In April, a GitHub project called Colleague Skill struck a nerve by claiming to βdistillβ a workerβs skills and personalityβand replicate them with an AI agent. Though the project was a spoof, it prompted a wave of soul-searching among otherwise enthusiastic early adopters.
A number of tech workers told MIT Technology Review that their bosses are already encouraging them to document their workflows for automation via tools like OpenClaw. Many now fear that they are being flattened into code and losing their professional identity.
A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)
+ Worried about Flock cameras? These guys designed a car to fool them. + Webbβs Near-Infrared Camera has captured a galactic mergerβs dazzling final phase. + An exquisitely preserved 66-million-year-old bird feather was found in a fossilised dinosaur dropping. + A plucky preservationist travelled 1,700 miles and made 52 calls from a rare phone box to keep it in service.
Every year, MIT Technology Review puts together a list of some of the brightest and best young minds working across science and technology. Our 35 Innovators Under 35 are the ones to watchβpeople whose research and technical work stands to shape the future of their fields.
This year, the list includes nine people who are transforming biotech. And this week, Iβm going to give you a taste of some of the very cool stuff five of them are working on, which includes lifesaving innovations and gr
Every year, MIT Technology Review puts together a list of some of the brightest and best young minds working across science and technology. Our 35 Innovators Under 35 are the ones to watchβpeople whose research and technical work stands to shape the future of their fields.
This year, the list includes nine people who are transforming biotech. And this week, Iβm going to give you a taste of some of the very cool stuff five of them are working on, which includes lifesaving innovations and groundbreaking βage reversalβ tech.Β Β
1.Β Preventing maternal deaths
Letβs start with Paschal Kija, a 28-year-old who has developed a device to treat postpartum hemorrhageβa dangerous birth complication that contributes to around 29% of maternal deaths in his home country, Tanzania. The Mkanda Salama (βSafe Wrapβ in Swahili) is easy to use and costs just $70. A study found that it stopped postpartum bleeding in 73% of women within 20 minutes.
2.Β Making brain electrodes inspired by Japanese art
For decades, scientists have been developing, testing, and implanting brain electrodes. These devices are literally inserted into peopleβs brains, so while they can help us understand brain activity and treat various neurological disorders, itβs not totally surprising that they can also cause a bit of damage. Xiao Yang, 34, is working on ultra-small electrodes, which she hopes will have less of an impact on surrounding brain tissue. Her electrodes are flexible, tooβin fact, they look a lot like actual neurons.
Yang is also creating sheets of electrodes to study brain cells in the lab. Inspired by kirigamiβthe traditional Japanese art of cutting paper to form three-dimensional shapesβsheβs created a sheet of electrodes with a honeycombed structure shaped like a spiral basket. And sheβs already using it to study brain cells.
3.Β Developing an all-new treatment for baby KJ
In 2024, Kyle βKJβ Muldoon Jr. was born with a rare and potentially fatal genetic disorder.Β Sarah Grandinette was a member of a team that developed an entirely new, personalized treatment for himβa gene-editing therapy essentially designed to correct a genetic misspelling.
Grandinette, who is now 26, created cells with KJβs genetic variant and used them to screen gene-editing approaches; then she tested potential medicines in mice and monkeys. KJ ultimately got his first dose of the resulting treatment when he was about seven months old. He responded well and was eventually discharged from hospital. Heβs βdoing pretty great,β she says.
4.Β Reversing the aging process to treat eye disease
The buzziest tech in longevity right now centers on reprogrammingβattempts to rewind the age of cells by resetting them to a more embryonic-like state. In a study published in 2020, Yuancheng (Ryan) Lu (now 34) and his colleagues showed that a reprogramming therapy reversed vision loss in aged, blind mice. Now an almost identical version of that therapy is being tested in people with eye disease. Life Biosciences, the company developing the drug, dosed its first volunteer in June.
5.Β Using AI to design new viruses
Last year, Samuel King used a generative AI model to come up with new genetic blueprints for bacteriophagesβteeny viruses that can infect bacteria. Once he had those blueprints, he printed them out as strands of DNA. In experiments, he found that those AI-designed viruses could create new copies of themselves, burst out of bacterial cells, and infect other nearby bacteria. Viruses arenβt alive, but King, 27, hopes that AI-designed life forms might one day be used to make drugs or soak up pollution.
You can read more about these innovators, and the others on the biotech list, here.
This article first appeared in The Checkup,Β MIT Technology ReviewβsΒ weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first,Β sign up here.
This is todayβs edition of The Download, our weekday newsletter that provides a daily dose of whatβs going on in the world of technology.
God told them to sell crypto. Their investors lost everything.
When Eli Regalado first heard God speak to him, he wondered whether he was hallucinating. According to Eli and his wife, Kaitlyn, He told them to get married, buy a house, and start having kids. Then in 2021, divine guidance steered them in an unexpected new direction: crypto.
That October
This is todayβs edition of The Download, our weekday newsletter that provides a daily dose of whatβs going on in the world of technology.
God told them to sell crypto. Their investors lost everything.
When Eli Regalado first heard God speak to him, he wondered whether he was hallucinating. According to Eli and his wife, Kaitlyn, He told them to get married, buy a house, and start having kids. Then in 2021, divine guidance steered them in an unexpected new direction: crypto.
That October, the Regalados later testified in court, they received holdings in a little-known digital coin. βTake this to my people for a wealth transfer,β Eli heard God say. Over time, they came to believe that He wanted them to launch their own coin.
The Regalados created INDXcoin, which they promoted through family, friends, and contacts in evangelical Christian circles. In all, more than 500 people handed over more than $3 million. But within a year, the project collapsed. Investors lost it all, leaving many to wonder where the funds went and whether they had fallen victim to an elaborate fraud.
This article is part of the Big Story series, the home of MIT Technology Reviewβs most important and ambitious reporting. You can read the rest of the series here.Β
The story was produced in partnership with Type Investigations and with support from the Fund for Investigative Journalism.
This road map could help us decide whether to deploy solar geoengineering
Scientists have spent half a century exploring whether we could counteract climate change by releasing reflective particles into the stratosphere, mimicking the cooling effects of volcanic eruptions. But even after hundreds of studies, we still donβt know how well it would work or what else it might doβand thereβs no systematic plan for clearing up that uncertainty.
Reflective, a research organization, has now attempted to fill that gap. The San Francisco nonprofit has published a detailed road map of the experiments, studies, and infrastructure that it says would be needed to make informed decisions about the use of solar geoengineering, MIT Technology Review can reveal.
This founder is teaching chips how to recycle (their energy)
Throughout the history of the computer chip, engineers have treated waste heat as an inevitable cost of a calculation. Hannah Earley, however, thinks itβs a design choice.
Earley, 31, is cofounder and CTO of Vaire Computing, which builds chips that recycle energy usually thrown away as heat, a strategy known as reversible computing. The approach could make data centers (and our laptops and phones) much more energy efficient.
Last year, Vaire announced a key breakthrough: a chip with a resonator that recovered more energy than it lost, even after the energy needed to power the component was taken into account.
Hannah Earley is one of the computing and robotics honorees on our 35 Innovators Under 35 list for 2026. Meet the rest of them here, or explore the full list across the biotechnology, AI, computing and robotics, and climate and energy categories.
Can the US battery market untangle from China?
βCasey Crownhart
The US energy storage market is growing at a record pace, which could shore up the grid and cut emissions. Crucially, this is all happening with the help of cheap Chinese batteries, which the Trump administration is trying to phase out.
Reducing reliance on any single source of crucial energy technology makes sense. But the tension raises a broader question for me: how much should countries take advantage of cheap, available tech, and how much should they cut themselves off from foreign sources to develop their own, even if it costs more?
This story is from The Spark, our weekly climate tech newsletter. Sign up to receive it in your inbox every Wednesday.
The must-reads
Iβve combed the internet to find you todayβs most fun/important/scary/fascinating stories about technology.
1 OpenAIβs agents used at least 10 websites for unauthorized communications Researchers found they bypassed restrictions on posting online.(Reuters $) + The company faces a Senate probe into the Hugging Face breach. (Axios) + Its hacking issues may indicate cultural problems. (MIT Technology Review)
2Β Another Anthropic model hacked a real system during testing A misconfigured environment gave it internet access. (CBS News) + The January incident went undetected until last month. (Reuters $) + AI agents are not your βcoworkers.β (MIT Technology Review)
3 Apple has entered the foldable phone market with the $1,999 iPhone Duo It opens into a 7.6-inch display and launches October 23. (NPR) + Apple is betting its design and privacy will give it an edge. (Reuters $) + And that foldables can solve the smartphoneβs sameness problem. (NPR $) + Samsung responded with a campaign touting its foldable lead. (CNBC) + In China, Apple enters a crowded market dominated by Huawei. (SCMP)
4 US prosecutors have called Huawei a criminal enterprise at trial They accuse the company of stealing American technology. (Reuters $) + And helping Iran snoop on its citizens. (AP News) + The trial could impact Trumpβs upcoming meeting with Xi. (WSJ $)
5 California is warming to nuclear power after decades of opposition The state may extend Diablo Canyon and lift its ban on new reactors. (NYT $) + China is betting on big nuclear reactors. (MIT Technology Review)
6 Chinese professionals are becoming gig workers training AI Lawyers and engineers are training models for extra income. (Rest of World) + Gig workers are training humanoids at home. (MIT Technology Review)
7 The new Apple Watch can listen to conversations happening nearby Apple says users must opt in, but others cannot. (Wired $)
8 Pink noise during sleep could help the brain clear away waste Timed bursts boosted brain fluid flow in a small study. (New Scientist $)
9 A lost supercontinent may have triggered the explosion of life Gondwanaβs formation fueled volcanic activity and warmed the planet. (404 Media)
10 GTA VI has sparked a debate over whether virtual romance is cheating Players can date, have sex with, and shower gifts on virtual partners. (Guardian)
Quote of the day
βWe must work to crush any dissent to Doomβs vision of public safety.βΒ
βA Seattle policy adviser dressed as Doctor Doom protests the cityβs expanding network of Flock and Axon surveillance systems at a Public Safety Committee meeting, 404 Media reports.
One more thing
Digging for clues about the North Poleβs past
In the past, getting to the North Pole involved a treacherous trip through ice many meters thick. But last year, a research vessel encountered open water and thin ice, which created an easy passage. It provided a reminder of how quickly the Arctic is changing.Β
Now scientists are digging deep below the seabed to find out if the Arctic Ocean was ever ice-freeβand what that could mean for the future of Earthβs northernmost waters.Β
A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line.)
+ Dutch kids have been declared the worldβs happiest (again). Hereβs why. + Travel through music history by picking a country and decade on Radiooooo. + These 16 majestic aerial photos reveal wildlife from perspectives you rarely see. + A Toronto cafe is pushing croissant engineering to new heights with its egg-shaped, custard-filled βCrogg.β
On July 22, 2026, a transmission line fault in Ashburn, Virginiaβthe heart of the worldβs largest data center clusterβknocked more than 3 gigawatts of load off the grid in seconds. And it wasnβt the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time.
The AI power debate is mostly about generation: more turbines
On July 22, 2026, a transmission line fault in Ashburn, Virginiaβthe heart of the worldβs largest data center clusterβknocked more than 3 gigawatts of load off the grid in seconds. And it wasnβt the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time.
The AI power debate is mostly about generation: more turbines, more solar, more transmission. The grid needs more electrons. But the outages in Virginia werenβt supply failures; they were architecture failures. And a giant wave of interconnections is arriving on that same architecture, putting grid reliability at risk. Itβs a problem nobody wants to own.
Asking more from the grid
The grid was built around predictable loads: steel mills, refineries, and houses at dinnertime. Different load sizes, same processβdrawing power smoothly, misbehaving occasionally, and recovering gracefully.
But AI data centers donβt behave that way.
An AI campus can swing 70% of its load in milliseconds during a training run, then trip offline just as fast at the first sign of trouble upstream to protect billions in compute. Each is rational alone. Together, at gigawatt scale, theyβre a problem the grid has never solvedβand the next wave of data center campuses is planned at exactly that scale.
Where the old stack breaks
The standard data center power stack hasnβt changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage uninterruptible power supply (UPS) units condition it, and it reaches the racks. Push that design to AI scale, and it cracks in three places.
First, the UPS sits deep inside the building, close to the racks. But its batteries are an undersized spare tire, designed to handle an outage for a few minutes, not to absorb load swings this fast and volatile around the clock.
Second, the UPS spends most of its life in bypass. Legacy converters waste enough power that operators run in eco-mode: A static switch feeds the racks directly from the grid and nothing filters in either direction. The computeβs swings go out raw, and grid transientsβsub-millisecond events that can damage or take down equipmentβcome in too fast for any switch to catch.
This isnβt sloppy engineering. Itβs careful engineering the load has outgrown.
Moving into the path
The fix is three moves, made together.
Move it upβfrom 480 volts to medium voltage (13.8 kilovolts and higher), the voltage large sites draw from the grid.
Move it outβfrom the data hall to modular enclosures near the substation so the building holds only compute and the cooling that keeps it alive.
Move it into the pathβinstead of a battery that watches and reacts, a system every electron runs through, all the time. Thereβs nothing to detect and nothing to switch because nothing was ever routed around it.
On paper, three straightforward upgrades. In practice, they rewrite every line item downstream.
Making the change
When thousands of GPUs spin up together, the system absorbs the swing and hands the grid a flat load profile. When a disturbance hits, the equipment behind it never notices. A difficult neighbor becomes a predictable one. And when the utility needs help, it becomes a useful one.
Interconnection changes, too. The utility certifies one medium-voltage box instead of untangling every transformer, UPS, chiller, pump, and switchgear lineup behind it. Engineers swap chip generations without a fresh interconnection study. Months come off the permitting timeline.
Inside the fence, UPS rooms become compute or cooling space. Density per construction dollar climbs.
And the economics flip. Equipment that runs at medium voltage, sits outside, and stores its own energy can qualify for tax credits, and earn revenue in grid programs like peak shaving and demand response. Backup power stops being insurance and starts paying for itself.
The architecture test
In early 2026, we tested a full-scale system at the National Laboratory of the Rockies, a U.S. Department of Energy facility and the only place in the Western Hemisphere that can replicate real grid faults and AI-scale load swings concurrently in the same loop.
We hit it from both directions: real AI load profiles hit the compute side at full medium voltage. Grid faults hit the utility side, including a full zero-voltage event. The compute side didnβt flinch. Neither did the grid side. It cleared the large-load voltage ride-through requirements from the Electric Reliability Council of Texas (ERCOT), the grid operator, with room to spare.
Those rules exist because operators no longer take facilities this size on faith, and more are coming. Most of the industry treats them as hurdles. A medium-voltage, inline system clears them out of the box. Compliance isnβt an added feature. Itβs what the architecture does.
The new layer
Much of what looks like a grid problem in the AI buildout sits inside the fence, in equipment sized for a load that no longer exists. Move the right pieces up, out, and into the path, and a grid liability becomes a grid asset. Density goes up. Permitting time comes down. Backup power earns its keep.
The engineering worksβand the next wave of AI factories is being built on it. The industry hasnβt named this layer yet. We call it the medium-voltage AI UPS. The name matters less than the choice: those factories can arrive as a strain on the grid or as strength for it. We already know how to build the second kind.Β Β Β Β
This content was produced by ON.energy. It was not written by MIT Technology Reviewβs editorial staff.
The US is hitting records for the rapid growth of its energy storage market. Thatβll go a long way to shoring up the grid, increasing reliability and also cutting emissions, since batteries can help store energy from intermittent renewables like wind and solar.
Crucially, this is all happening with the help of cheap Chinese batteries, though thereβs been a concerted effort to reduce the USβs reliance on them. Most recently, in an executive order in late August, the Trump administration dec
The US is hitting records for the rapid growth of its energy storage market. Thatβll go a long way to shoring up the grid, increasing reliability and also cutting emissions, since batteries can help store energy from intermittent renewables like wind and solar.
Crucially, this is all happening with the help of cheap Chinese batteries, though thereβs been a concerted effort to reduce the USβs reliance on them. Most recently, in an executive order in late August, the Trump administration declared a national emergency that essentially bans Chinese batteries from being used in grid-scale energy storage systems.
Thereβs an argument to be made about reducing reliance on any single source of a crucial energy technology. But all this tension raises a broader question for me: How much should countries take advantage of cheap, available tech, versus cutting off major sources to force development of their own factories even if that comes at a higher cost?
This is hardly Americaβs first push to move away from Chinese influence in the battery supply chain. One of the major policy tools used in recent years is restricting the tax credits designed to incentivize use of the new technologies. Limiting the types of projects that are eligible can help reduce the cost of local technologies so theyβre more competitive with otherwise cheaper imported options.
Back in 2022, the US government designed the tax credits that were part of the Inflation Reduction Act to restrict where a batteryβs minerals could be mined, processed, or recycled, as well as where a battery and its components were assembled.
Those tax credits underwent a makeover in 2025, but the Trump administration has taken a similar tack. New legislation requires that starting in 2026, 55% of the cost of materials used for new energy storage projects must come from outside China and other restricted countries or the projects wonβt qualify for tax credits.Β
And we canβt forget about tariffs. Import taxes for batteries increased to 25% in January, up from 7.5%.
But the new executive order is a more drastic move. It bans the installation of βany foreign-produced bulk-power system electric equipmentβ that poses a national security risk. The order specifically calls out battery energy storage systems, as well as inverters and transformers.
βAn outright ban was a bit of a surprise, and it does create a bit of concern for domestic players in the US,β says Shan Tomouk, energy storage and energy lead for Benchmark Mineral Intelligence, an energy industry analyst.
The move is likely to slow deployment of grid-connected energy storage projects in the near term, according to analysis from BloombergNEF, an energy consultancy. Projects could face delays as developers wait for clarity on the rules.
Depending on the detailed guidance from the Department of Energy, which is expected by the end of the year, some projects may need to find alternative sources for their cells, whether theyβre domestically produced or imported from other countries. These will likely be more expensive than Chinese imports, says Isshu Kikuma, an energy storage analyst at BloombergNEF. βWorst case, those projects could get canceled,β he says.
Technically, the order applies even to existing energy storage plants, though itβs unlikely that theyβll be taken offline because of their batteriesβ origin. Since most of these plants currently use Chinese batteries, enforcing the order to the letter would essentially mean removing most installed battery energy storage from the US grid, Kikuma says.
In the longer term, the US will eventually be able to meet its own demand for batteries. The country could have enough capacity by about 2030, though some factories may not ramp up or run at their full capability, meaning domestic supply wonβt actually meet demand until later in the 2030s.Β
New factories from LG Energy Solutions, Samsung SDI, Ford, and SK On are set to come online or ramp up by next year. In an ironic twist, a slowing EV market is helping, as some factories originally designed for vehicle batteries are retooling to build cells for grid storage instead.Β
But it will come at a cost. Today, batteries produced in the US are still significantly more expensive than those made in China. Even switching to imports from other countries like South Korea would likely be more expensive.
This is a crucial issue that goes beyond the US and even beyond batteries. China is miles ahead of much of the rest of the world on technologies like solar panels and batteries. Through years of government support and experience with research and manufacturing, the nation is an energy powerhouse.
Thereβs a delicate political balance to maintain as the world figures out how to navigate this situation. Thereβs cheap technology on offer, which can help drastically reduce emissions and energy costs. But there can be risks associated with relying too much on any one player for crucial technologies.
This article is from The Spark, MIT Technology Reviewβs weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here.Β
arXiv:2609.09625v1 Announce Type: new
Abstract: As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focus on specific enabling techniques, such as learning modules, knowledge graphs, and large language models, while providing limited insight into how cognition can be systematically i
arXiv:2609.09625v1 Announce Type: new
Abstract: As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focus on specific enabling techniques, such as learning modules, knowledge graphs, and large language models, while providing limited insight into how cognition can be systematically integrated into DT architectures. To address this issue, this paper proposes a four-layer CDT architecture consisting of the physical layer, digital-twin layer, cognitive layer, and task layer. The proposed architecture establishes a self-evolving closed operational loop spanning these four layers, in which physical states are synchronized into digital representations, cognition constructs task-specific cognitive models through knowledge, memory, and attention, and task-level decisions are generated under practical constraints. Operational feedback further refines cognitive experience and updates relationships and annotations in the digital representation, enabling subsequent task interpretation, initiation, and reasoning to evolve with system operation. Based on this framework, two representative operation modes are characterized: user-request-driven cognition and self-driven cognition. We further discuss key enabling mechanisms and deployment challenges associated with semantic communication, knowledge querying, task orchestration, and closed-loop synchronization. A lightweight simulation study illustrates reliable closed-loop task feasibility under limited semantic information and improved operational efficiency through accumulated task experience. The proposed framework provides a structured foundation for the design and development of future CDT systems.
arXiv:2609.09754v1 Announce Type: new
Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucina
arXiv:2609.09754v1 Announce Type: new
Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.
arXiv:2609.09774v1 Announce Type: new
Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code
arXiv:2609.09774v1 Announce Type: new
Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code was adapted while the separately stored high-level procedure was not reported to change; this phase does not constitute an autonomous memory-agent evaluation. In the controlled phase, an early pilot produced one task on which two memory conditions selected a more expensive item while the no-memory condition selected the reference minimum. Follow-up probes did not establish a recurring row-order or identity-binding pattern. We then tested four forms of mismatch: changed quantities, a different evidence representation, a conflict between local and global optimization, and distributed promotion evidence, across 32 formal cells. Each cell used one temperature-0 generation with the same local qwen3:8b configuration and no adaptive retry. Across these pairs, none of the predefined diagnostic interference signatures appeared on the tasks for which they were defined when current-task evidence was explicit and sufficient. The result identifies a tested region of non-interference: a procedural memory can be mismatched without becoming behaviorally disruptive. It does not establish general safety or a mechanism. The remaining question is which additional conditions turn applicability mismatch into observable, memory-caused error.
arXiv:2609.10055v1 Announce Type: new
Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective
arXiv:2609.10055v1 Announce Type: new
Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective hierarchy-guided refinement. We also construct PhenoNormBench, a unified benchmark comprising 13,390 samples from seven Human Phenotype Ontology datasets. OntologyAligner achieved state-of-the-art performance on HPO normalization, with 88.78% Macro Top-1 Accuracy and 86.75% Micro Top-1 Accuracy, exceeding the strongest baseline by 4.85 and 5.07 percentage points, respectively. Ablation analyses showed complementary contributions from all three stages, and sensitivity analyses demonstrated stability across candidate-set sizes and model backbones. Applications to MONDO, MEDIC, and NCBITaxon further established portability to other ontologies. OntologyAligner offers a generalizable framework for accurate mapping of biomedical text to structured ontology concepts. PhenoNormBench and the code are publicly available at https://github.com/zhelishisongjie/OntologyAligner.
arXiv:2609.10335v1 Announce Type: new
Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when equipped with specialized modules, can rival state-of-the-art LMMs on complex geometry problems. Our framework integrates a Geomet
arXiv:2609.10335v1 Announce Type: new
Abstract: Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when equipped with specialized modules, can rival state-of-the-art LMMs on complex geometry problems. Our framework integrates a Geometric Vision Parser, which translates diagrams into symbolic form, with a Symbolic Solver that performs formal deductions, thereby mitigating hallucinations and promoting interpretable reasoning. To enable rigorous evaluation, we curate a benchmark of challenging problems from the 2025 Chinese Zhongkao examinations, ensuring data novelty and testing deeper deductive skills. Experiments demonstrate that our approach achieves performance comparable to Gemini 2.5 Pro while delivering clearer, human-like solutions.
arXiv:2607.15957v3 Announce Type: cross
Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior. However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains
arXiv:2607.15957v3 Announce Type: cross
Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior. However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness.Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.
arXiv:2609.09348v1 Announce Type: cross
Abstract: Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet the architectural relationship between the two is seldom made explicit. We build on Wang et al.'s taxonomy of Continuum Orchestration Systems employing
arXiv:2609.09348v1 Announce Type: cross
Abstract: Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet the architectural relationship between the two is seldom made explicit. We build on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extend it with two further dimensions. The AI Augmentation Paradigm measures how LLMs are exploited, while the Feedback channel captures whether and through which system path the execution feedback returns to the LLM in order to close the MAPE control loop at the LLM Orchestration layer. We apply this taxonomy to six recent system architectures and find a common gap, as none combines full LLM orchestration with full agent-layer feedback in a Cloud Continuum setting. We relate this gap to a missing cross-tier feedback abstraction, bridging the incommensurable per-tier signals and the LLM Orchestrator.
arXiv:2609.09438v1 Announce Type: cross
Abstract: Critical systems sit near boundaries between qualitatively distinct behaviors. When inferring models of neural activity, this proximity to criticality is thought to require the precise tuning of parameters. Here, we show that as the number of neurons increases, criticality can emerge naturally without fine-tuning. When computing observable statistics from parameters (the forward problem), some small regions in parameter space map to large region
arXiv:2609.09438v1 Announce Type: cross
Abstract: Critical systems sit near boundaries between qualitatively distinct behaviors. When inferring models of neural activity, this proximity to criticality is thought to require the precise tuning of parameters. Here, we show that as the number of neurons increases, criticality can emerge naturally without fine-tuning. When computing observable statistics from parameters (the forward problem), some small regions in parameter space map to large regions in statistics space. These special parameters are precisely those near criticality. Thus, when inferring parameters from experimental measurements (the inverse problem), models concentrate near critical points, and this concentration becomes stronger as the system grows. We illustrate this flow toward criticality across many large-scale recordings in the mouse brain. In the Curie-Weiss model of Ising spins, we find that all of the recordings collapse to a first-order phase transition, despite substantial differences in the underlying systems. Together, these results suggest a resolution to the tension between criticality and fine-tuning in models of neural activity.
arXiv:2609.09606v1 Announce Type: cross
Abstract: Neural radiance fields (NeRFs) and 3D Gaussian Splatting (3DGS) encode a scene with complementary inductive biases, but existing cross-representation distillation typically fixes one representation as teacher for the entire scene. A globally fixed teacher can propagate local reconstruction errors. We present RouteBridge, a bidirectional framework that selects the teaching direction for each ray. Its reliability estimator combines photometric res
arXiv:2609.09606v1 Announce Type: cross
Abstract: Neural radiance fields (NeRFs) and 3D Gaussian Splatting (3DGS) encode a scene with complementary inductive biases, but existing cross-representation distillation typically fixes one representation as teacher for the entire scene. A globally fixed teacher can propagate local reconstruction errors. We present RouteBridge, a bidirectional framework that selects the teaching direction for each ray. Its reliability estimator combines photometric residuals with representation-specific geometric evidence and routes supervision from NeRF to 3DGS, from 3DGS to NeRF, or abstains. A renderer-independent interface transfers color, opacity, and normalized depth without shared features or point correspondence. On mip-NeRF 360, the NeRF and 3DGS exports reach 28.56 and 28.77 dB, respectively. The 3DGS export improves over 3DGS by 1.56 dB and over NeRF-GS by 0.45 dB while reducing LPIPS to 0.207. On static three-view DTU, RouteBridge obtains 21.12 dB. Ablations show that both adaptive routing and geometric ray targets contribute to the improvement.
arXiv:2609.09692v1 Announce Type: cross
Abstract: Chain-of-Thought (CoT) prompting enables LLMs to perform explicit, step-by-step reasoning, creating opportunities for sophisticated autonomous robots. However, recent research reveals that reasoning models verbalize their actual decision processes only 25-39% of the time, with faithfulness degrading 44% on complex tasks. This paper presents CT-SAFR (Chain-of-Thought Safety and Faithfulness for Robotics), a multi-layered verification framework ac
arXiv:2609.09692v1 Announce Type: cross
Abstract: Chain-of-Thought (CoT) prompting enables LLMs to perform explicit, step-by-step reasoning, creating opportunities for sophisticated autonomous robots. However, recent research reveals that reasoning models verbalize their actual decision processes only 25-39% of the time, with faithfulness degrading 44% on complex tasks. This paper presents CT-SAFR (Chain-of-Thought Safety and Faithfulness for Robotics), a multi-layered verification framework achieving 94.2% hallucination detection (n = 500, 95% CI: 91.8-95.9%) with sub-500ms latency. Through a warehouse robot case study, this work demonstrates 87% reduction in unsafe reasoning outputs (p
arXiv:2609.09884v1 Announce Type: cross
Abstract: Recent advances in Intrinsic Image Decomposition (IID) have increasingly relied on generative models. However, progress remains limited by three key challenges: (a) insufficient physical consistency, (b) high computational cost at inference time, and (c) limited generalization capabilities. In this work, we show that latent bridge matching (LBM) effectively addresses these limitations for albedo estimation. We introduce a novel LBM-based archite
arXiv:2609.09884v1 Announce Type: cross
Abstract: Recent advances in Intrinsic Image Decomposition (IID) have increasingly relied on generative models. However, progress remains limited by three key challenges: (a) insufficient physical consistency, (b) high computational cost at inference time, and (c) limited generalization capabilities. In this work, we show that latent bridge matching (LBM) effectively addresses these limitations for albedo estimation. We introduce a novel LBM-based architecture that enforces physical consistency through a pixel reconstruction loss, benefits from the inherent efficiency of LBM low-cost inference, and improves generalization across diverse datasets by incorporating a shading conditioning. In this extended version, we additionally show that conditioning the shading estimator itself on the predicted albedo further improves reconstruction fidelity, and we benchmark our best model against stateof-the-art IID methods across five real and synthetic datasets.
arXiv:2609.10125v1 Announce Type: cross
Abstract: Trochlear dysplasia (TD) is an abnormality of the femoral trochlea associated with anterior knee pain and patellar instability. The sulcus angle (SA) is used to assess trochlear morphology, but it is typically measured on a single axial MR slice with no clear guidance on which to select, making it sensitive to slice selection and landmark placement. We propose an automatic framework for continuous SA profiling from super-resolved MR volumes. Cli
arXiv:2609.10125v1 Announce Type: cross
Abstract: Trochlear dysplasia (TD) is an abnormality of the femoral trochlea associated with anterior knee pain and patellar instability. The sulcus angle (SA) is used to assess trochlear morphology, but it is typically measured on a single axial MR slice with no clear guidance on which to select, making it sensitive to slice selection and landmark placement. We propose an automatic framework for continuous SA profiling from super-resolved MR volumes. Clinically acquired axial, coronal, and sagittal MR scans are combined using implicit neural representations to reconstruct a high-resolution volume. SA measurements are computed across the trochlear region using two landmark detection U-Net models. The approach was evaluated on the public fastMRI dataset and a small in-house cohort of patients with TD. Compared with conventional manual single-slice SA measurements, the proposed automated method yielded a mean absolute error of 11.6$^\circ$ while providing continuous characterization of trochlear morphology. Population-level analysis demonstrated distinct mean SA profiles between the public cohort and the in-house TD cohort, highlighting the potential of profile-based assessment to characterize TD. By reducing reliance on a single manually selected axial slice, the proposed framework extends conventional SA assessment to a continuous profile-based description of trochlear morphology without additional imaging, while remaining conceptually linked to current clinical assessment. Further validation is required. The code is available: https://github.com/wehrlimi/SA_Profile.
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning an
arXiv:2609.10464v1 Announce Type: cross
Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.
arXiv:2609.10522v1 Announce Type: cross
Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically g
arXiv:2609.10522v1 Announce Type: cross
Abstract: Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.