Aging Dis. 2025 Dec 15. doi: 10.14336/AD.2025.1444. Online ahead of print.
ABSTRACT
Chronic gastritis (CG) is a highly prevalent, age-associated inflammatory disorder of gastric mucosa and a key precursor of gastric cancer in older adults. Beyond Helicobacter pylori infection and environmental insults, accumulating evidence indicates that chronic, low-grade inflammation coupled with aging biology, "gastric inflammaging", plays a central role in driving mucosal degeneration, atrophy, and malignant transformation. Here, we synthesize current mechanistic and multi-omics evidence to conceptualize CG as a tractable model of organ-specific inflammaging. We first summarize how hallmarks of aging-including cellular senescence and the senescence-associated secretory phenotype (SASP), mitochondrial dysfunction, impaired autophagy, immune exhaustion, and microbiome dysbiosis-converge to create a self-perpetuating inflammatory microenvironment in the stomach. We then review emerging single-cell and spatial multi-omics studies that delineate senescence-inflammation niches and reveal how these molecular neighborhoods relate to disease stage and cancer risk. Finally, we discuss therapeutic implications, highlighting geroscience-guided interventions such as senolytics/senomorphics, inflammasome and cGAS-STING pathway modulators, microbiota- and metabolite-targeted strategies, lifestyle interventions, and natural products, and propose a precision framework linking inflammaging biomarkers to patient stratification and clinical endpoints. Reframing CG as a gastric inflammaging model may provide a prototype for organ-specific healthy aging strategies and near-term gerotherapeutic trials aimed at extending healthspan.
Chin Med J (Engl). 2025 Nov 28;138(24):3332-50. doi: 10.1097/CM9.0000000000003922. Online ahead of print.
ABSTRACT
Personalized medicine for gastric cancer continues to face numerous challenges, primarily due to the complexity of clinical decision making and the difficulty of integrating multimodal data. Artificial intelligence (AI), with its powerful capabilities in feature learning and pattern recognition, is emerging as a key technology to overcome these barriers. It provides critical support in areas such as early screening, histological subtyping, prediction of treatment response, and prognostic risk stratification. This review examines the application of AI in diagnosing and treating gastric cancer, with particular attention to the current mainstream AI methodologies, including feature engineering and deep learning and the rapidly evolving pretrained foundation models and multimodal large models. With the integration of medical images, digital pathology, multiomics data, and structured clinical information, AI systems are increasingly effective at capturing tumor heterogeneity and supporting complex clinical decisions in real time. On the one hand, task-specific models have demonstrated excellent performance in subtyping, staging, and prognosis assessment. On the other hand, the rise of foundation models and general-purpose large models is redefining the limits of AI in cross-task transfer, complex reasoning, and human-machine interaction. These technologies hold promise in addressing key obstacles such as data scarcity, modality heterogeneity, and fragmented clinical workflows, offering a feasible path toward a unified and efficient AI-driven diagnostic and therapeutic system for gastric cancer. As technological maturity progresses alongside the development of robust safety and ethical frameworks, AI is expected to evolve from a static auxiliary interpretation tool into an intelligent decision-making platform capable of semantic understanding, dynamic feedback, and multidisciplinary collaboration-therefore playing a pivotal role across the full spectrum of precision medicine in gastric cancer.
You’re reading the web edition of STAT’s Health Tech newsletter, our guide to how technology is transforming the life sciences. Sign up to get it delivered in your inbox every Tuesday and Thursday.
A tumultuous year at the Food and Drug Administration will be capped off at the agency’s devices center with the departure of two key leaders, just as regulators are sorting through challenges related to artificial intelligence and launching new initiatives on software as a medical device regulation.
Sources tell us Jessica Paulsen, a 15-year veteran of FDA and acting deputy director of its Digital Health Center of Excellence is leaving the agency. She’s been leading the center since last summer when the last acting head, SonjaFulmer, left FDA for Mayo Clinic. Fulmer took over for Troy Tazbaz who left in January to return to Oracle. The center’s work includes communicating with industry and developing guidances relevant to digital health. (FDA did not respond to a request for comment.)
Neuralink, Elon Musk’s frothy brain-computer interface company, poached David McMullen, director of FDA’s office of neurological and physical medicine devices, which is in charge of regulating Neuralink. McMullen spent three years atop the office and previously worked at the National Institute for Mental Health.
Both Paulsen and McMullen were at the forefront of important conversations about the future of regulation. I grabbed the screenshot above of the two leaders from a video of last month’s Digital Health Advisory Committee meeting on generative AI-enabled mental health devices. Separately, McMullen’s office will have oversight of behavioral health devices under the FDA’s new TEMPO pilot.
New to me: As part of the funding package that reopened the government last month, lawmakers passed full-year 2026 funding for FDA. Buried within the Senate report accompanying the legislation, lawmakers direct FDA to, within 90 days, (February) report on its authorities to regulate AI medical devices, and within 180 days, (May) report on “the status of the FDA’s efforts regarding engagement on AI in drug development.”
The Government Accountability Office last week released a report on medical device recalls which found, among other things, that “insufficient staff limit FDA’s ability to conduct oversight activities.” In other words, the FDA already does not have enough staff to oversee medical devices and is losing key leadership at a time when new technology and initiatives may require additional horsepower.
The future of the mammogram
Applying AI to mammograms to help radiologists spot signs of breast cancer is increasingly common but researchers and AI companies want to apply new analyses to the routine screening tests to trigger more proactive care to prevent future cancers, heart attacks, and strokes. In one important breakthrough, the startup Clairity received FDA authorization for AI that offer a prediction of somone’s five-year breast cancer risk based on a mammogram alone.
This essay is part of a First Opinion series on the future of the National Institutes of Health and American science.
Around the world, nations with robust research and development infrastructure race to create therapeutics that meet the needs of their residents. Simply put, they dictate research priorities based on need. During the Covid pandemic, the United States was one of the first countries to gain access to vaccines to protect its citizens.
In an exceptional year for biotech and pharma, there were so many outstanding CEOs, I couldn’t single out just one for my 2025 Best Biopharma CEO honor.
Here are this year’s winners:
The dealmakers
2025 has been the best year for biotechs being acquired by pharma companies since 2019, with nearly $240 billion worth of deals announced or closed through November, according to Stifel.
Background: To date, there is no comprehensive paper that systematically synthesizes the effect of generative AI chatbot’s impact on mental health. Can generative AI chatbots help reduce our psychological distress? Objective: To comprehensively assess existing evidence, a systematic review and meta-analysis is essential to evaluate the overall effectiveness, identify gaps, and guide future research in this evolving field. This paper aims to: 1) synthesize current evidence on generative AI chatbot interventions targeting mental health issues, 2) quantify the effectiveness of these interventions via a meta-analysis of randomized controlled trials (RCTs), and examine key moderators of intervention effectiveness. Methods: This systematic review included 26 studies for narrative synthesis, out of which 12 randomized controlled trials were included in the meta-analysis. Results: The systematic synthesis revealed that 1) generative AI-chatbot interventions mostly took place in non-WEIRD countries (Western, Educated, Industrialized, Rich, and Democratic) and 2) there is a lack of studies focusing on young children and older adults. The meta-analysis showed a statistically significant effect (ES = 0.36, p = .039), which means that generative AI chatbots are, on average, effective in reducing negative mental health issues. Among moderators, we found statistically significant and higher effect sizes among interventions that have an active control group, conducted in WEIRD countries, recruited non-clinical populations, older age, majority female, non-personalized, with human assistance, and social-oriented. Conclusions: In conclusion, this comprehensive review has highlighted the potential of generative AI chatbots in addressing anxiety, depression, negative mood, and stress. The findings indicate that generative AI interventions are particularly beneficial in WEIRD countries, among non-clinical populations, older adults, and females. Human-assisted and social-oriented programs, as opposed to fully autonomous or task-oriented ones, demonstrate greater effectiveness. Meanwhile, non-personalized chatbots appear to yield more effective outcomes than personalized systems.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
The great AI hype correction of 2025
Some disillusionment was inevitable. When OpenAI released a free web app called ChatGPT in late 2022, it changed the course of an entire industry—and several world economies. Millions of people started talking to their computers, and their computers started talking back. We were enchanted, and we expected more.
Well, 2025 has been a year of reckoning. For a start, the heads of the top AI companies made promises they couldn’t keep. At the same time, updates to the core technology are no longer the step changes they once were.
This story is part of our new Hype Correction package, a collection of stories designed to help you reset your expectations about what AI makes possible—and what it doesn’t. Check out the rest of the package here, and you can read more about why it’s time to reset our expectations for AI in the latest edition of the Algorithm, our weekly AI newsletter. Sign up here to make sure you receive future editions straight to your inbox.
Quantum navigation could solve the military’s GPS jamming problem
Since the 2022 invasion of Ukraine, thousands of flights have been affected by a far-reaching Russian campaign of using radio transmissions that jammed its GPS system.
The growing inconvenience to air traffic and risk of a real disaster have highlighted the vulnerability of GPS and focused attention on more secure ways for planes to navigate the gauntlet of jamming and spoofing, the term for tricking a GPS receiver into thinking it’s somewhere else.
One approach that’s emerging from labs is quantum navigation: exploiting the quantum nature of light and atoms to build ultra-sensitive sensors that can allow vehicles to navigate independently, without depending on satellites. Read the full story.
—Amos Zeeberg
The must-reads
I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.
1 The Trump administration has launched its US Tech Force program In a bid to lure engineers away from Big Tech roles and straight into modernizing the government. (The Verge) + So, essentially replacing the IT workers that DOGE got rid of, then. (The Register)
2 Lawmakers are investigating how AI data centers affect electricity costs They want to get to the bottom of whether it’s being passed onto consumers. (NYT $) + Calculating AI’s water usage is far from straightforward, too. (Wired $) + AI is changing the grid. Could it help more than it harms? (MIT Technology Review)
3 Ford isn’t making a large all-electric truck after all After the US government’s support for EVs plummeted. (Wired $) + Instead, the F-150 Lightning pickup will be reborn as a plug-in hybrid. (The Information $) + Why Americans may be finally ready to embrace smaller cars. (Fast Company $) + The US could really use an affordable electric truck. (MIT Technology Review)
4 PayPal wants to become a bank in the US The Trump administration is very friendly to non-traditional financial companies, after all. (FT $) + It’s been a good year for the crypto industry when it comes to banking. (Economist $)
5 A tech trade deal between the US and UK has been put on ice America isn’t happy with the lack of progress Britain has made, apparently. (NYT $) + It’s a major setback in relations between the pair. (The Guardian)
6 Why does no one want to make the cure for dengue? A new antiviral pill appears to prevent infection—but its development has been abandoned. (Vox)
7 The majority of the world’s glaciers are forecast to disappear by 2100 At a rate of around 3,000 per year. (New Scientist $) + Inside a new quest to save the “doomsday glacier”. (MIT Technology Review)
8 Hollywood is split over AI While some filmmakers love it, actors are horrified by its inexorable rise. (Bloomberg $)
9 Corporate America is obsessed with hiring storytellers It’s essentially a rehashed media relations manager role overhauled for the AI age. (WSJ $)
10 The concept of hacking existed before the internet Just ask this bunch of teenage geeks. (IEEE Spectrum)
Quote of the day
“So the federal government deleted 18F, which was doing great work modernizing the government, and then replaced it with a clone? What is the point of all this?”
—Eugene Vinitsky, an assistant professor at New York University, takes aim at the US government’s decision to launch a new team to overhaul its approach to technology in a post on Bluesky.
One more thing
How DeepSeek became a fortune teller for China’s youth
As DeepSeek has emerged as a homegrown challenger to OpenAI, young people across the country have started using AI to revive fortune-telling practices that have deep roots in Chinese culture.
Across Chinese social media, users are sharing AI-generated readings, experimenting with fortune-telling prompt engineering, and revisiting ancient spiritual texts—all with the help of DeepSeek.
The surge in AI fortune-telling comes during a time of pervasive anxiety and pessimism in Chinese society. And as spiritual practices remain hidden underground thanks to the country’s regime, computers and phone screens are helping younger people to gain a sense of control over their lives. Read the full story.
—Caiwen Chen
We can still have nice things
A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)
+ Chess has been online as far back as the 1800s (no, really!) + Jane Austen was born 250 years ago today. How well do you know her writing? ($) + Rob Reiner, your work will live on forever. + I enjoyed this comprehensive guide to absolutely everything you could ever want to know about New England’s extensive seafood offerings.
Can I ask you a question: How do you feel about AI right now? Are you still excited? When you hear that OpenAI or Google just dropped a new model, do you still get that buzz? Or has the shine come off it, maybe just a teeny bit? Come on, you can be honest with me.
Truly, I feel kind of stupid even asking the question, like a spoiled brat who has too many toys at Christmas. AI is mind-blowing. It’s one of the most important technologies to have emerged in decades (despite all its many many drawbacks and flaws and, well, issues).
At the same time I can’t help feeling a little bit: Is that it?
If you feel the same way, there’s good reason for it: The hype we have been sold for the past few years has been overwhelming. We were told that AI would solve climate change. That it would reach human-level intelligence. That it would mean we no longer had to work!
Instead we got AI slop, chatbot psychosis, and tools that urgently prompt you to write better email newsletters. Maybe we got what we deserved. Or maybe we need to reevaluate what AI is for.
As my colleague Will Douglas Heaven puts it in the package’s intro essay, “You can’t help but wonder: When the wow factor is gone, what’s left? How will we view this technology a year or five from now? Will we think it was worth the colossal costs, both financial and environmental?”
Elsewhere in the package, James O’Donnell looks at Sam Altman, the ultimate AI hype man, through the medium of his own words. And Alex Heath explains the AI bubble, laying out for us what it all means and what we should look out for.
Michelle Kim analyzes one of the biggest claims in the AI hype cycle: that AI would completely eliminate the need for certain classes of jobs. If ChatGPT can pass the bar, surely that means it will replace lawyers? Well, not yet, and maybe not ever.
Similarly, Edd Gent tackles the big question around AI coding. Is it as good as it sounds? Turns out the jury is still out. And elsewhere David Rotman looks at the real-world work that needs to be done before AI materials discovery has its breakthrough ChatGPT moment.
Meanwhile, Garrison Lovely spends time with some of the biggest names in the AI safety world and asks: Are the doomers still okay? I mean, now that people are feeling a bit less scared about their impending demise at the hands of superintelligent AI? And Margaret Mitchell reminds us that hype around generative AI can blind us to the AI breakthroughs we should really celebrate.
Let’s remember: AI was here before ChatGPT and it will be here after. This hype cycle has been wild, and we don’t know what its lasting impact will be. But AI isn’t going anywhere. We shouldn’t be so surprised that those dreams we were sold haven’t come true—yet.
The more likely story is that the real winners, the killer apps, are still to come. And a lot of money is being bet on that prospect. So yes: The hype could never sustain itself over the short term. Where we’re at now is maybe the start of a post-hype phase. In an ideal world, this hype correction will reset expectations.
arXiv:2512.11912v1 Announce Type: new
Abstract: A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to sequence-to-sequence tasks, are remarkably resilient (for GPT-2, test NLL increases modestly from 2.87 to 3.59 despite 50% token corruption). By contrast, under the same levels of data corruption, class-conditional diffusion models degrade catastrophically (image-label consistency plummets by 56.81% relative to baseline), while classifiers show a moderate impact that diminishes with dataset scale. To explain these discrepancies, we analyze the results through a multi-perspective lens, integrating information theory, PAC learning, and gradient dynamics. These analyses suggest that robustness is heavily influenced by two key principles: the richness of conditioning information, which constrains the learning problem, and the absolute information content of the training data, which allows the signal from correct information to dominate statistical noise.
arXiv:2512.12413v1 Announce Type: new
Abstract: Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value. The present research conceptualises critical thinking in AI use as a dispositional tendency to verify the source and content of AI-generated information, to understand how models work and where they fail, and to reflect on the broader implications of relying on AI. Across six studies (N = 1365), we developed and validated the 13-item critical thinking in AI use scale and mapped its nomological network. Study 1 generated and content-validated scale items. Study 2 supported a three-factor structure (Verification, Motivation, and Reflection). Studies 3, 4, and 5 confirmed this higher-order model, demonstrated internal consistency and test-retest reliability, strong factor loadings, sex invariance, and convergent and discriminant validity. Studies 3 and 4 further revealed that critical thinking in AI use was positively associated with openness, extraversion, positive trait affect, and frequency of AI use. Lastly, Study 6 demonstrated criterion validity of the scale, with higher critical thinking in AI use scores predicting more frequent and diverse verification strategies, greater veracity-judgement accuracy in a novel and naturalistic ChatGPT-powered fact-checking task, and deeper reflection about responsible AI. Taken together, the current work clarifies why and how people exercise oversight over generative AI outputs and provides a validated scale and ecologically grounded task paradigm to support theory testing, cross-group, and longitudinal research on critical engagement with generative AI outputs.
arXiv:2512.12548v1 Announce Type: new
Abstract: Patch foraging involves the deliberate and planned process of determining the optimal time to depart from a resource-rich region and investigate potentially more beneficial alternatives. The Marginal Value Theorem (MVT) is frequently used to characterize this process, offering an optimality model for such foraging behaviors. Although this model has been widely used to make predictions in behavioral ecology, discovering the computational mechanisms that facilitate the emergence of optimal patch-foraging decisions in biological foragers remains under investigation. Here, we show that artificial foragers equipped with learned world models naturally converge to MVT-aligned strategies. Using a model-based reinforcement learning agent that acquires a parsimonious predictive representation of its environment, we demonstrate that anticipatory capabilities, rather than reward maximization alone, drive efficient patch-leaving behavior. Compared with standard model-free RL agents, these model-based agents exhibit decision patterns similar to many of their biological counterparts, suggesting that predictive world models can serve as a foundation for more explainable and biologically grounded decision-making in AI systems. Overall, our findings highlight the value of ecological optimality principles for advancing interpretable and adaptive AI.
arXiv:2512.12652v1 Announce Type: new
Abstract: This paper introduces the concept of value awareness in AI, which goes beyond the traditional value-alignment problem. Our definition of value awareness presents us with a concise and simplified roadmap for engineering value-aware AI. The roadmap is structured around three core pillars: (1) learning and representing human values using formal semantics, (2) ensuring the value alignment of both individual agents and multiagent systems, and (3) providing value-based explainability on behaviour. The paper presents a selection of our ongoing work on some of these topics, along with applications to real-life domains.
arXiv:2512.11814v1 Announce Type: cross
Abstract: Artificial intelligence (AI) scribes, systems that record and summarise patient-clinician interactions, are promoted as solutions to administrative overload. This paper argues that their significance lies not in efficiency gains but in how they reshape medical attention itself. Offering a conceptual analysis, it situates AI scribes within a broader philosophical lineage concerned with the externalisation of human thought and skill. Drawing on Iain McGilchrist's hemisphere theory and Lewis Mumford's philosophy of technics, the paper examines how technology embodies and amplifies a particular mode of attention. AI scribes, it contends, exemplify the dominance of a left-hemispheric, calculative mindset that privileges the measurable and procedural over the intuitive and relational. As this mode of attention becomes further embedded in medical practice, it risks narrowing the field of care, eroding clinical expertise, and reducing physicians to operators within an increasingly mechanised system.
arXiv:2512.12324v1 Announce Type: cross
Abstract: The rapid proliferation of Artificial Intelligence Generated Content has precipitated a crisis of trust and urgent regulatory demands. However, existing identification tools suffer from fragmentation and a lack of support for visible compliance marking. To address these gaps, we introduce the \textbf{UniMark}, an open-source, unified framework for multimodal content governance. Our system features a modular unified engine that abstracts complexities across text, image, audio, and video modalities. Crucially, we propose a novel dual-operation strategy, natively supporting both \emph{Hidden Watermarking} for copyright protection and \emph{Visible Marking} for regulatory compliance. Furthermore, we establish a standardized evaluation framework with three specialized benchmarks (Image/Video/Audio-Bench) to ensure rigorous performance assessment. This toolkit bridges the gap between advanced algorithms and engineering implementation, fostering a more transparent and secure digital ecosystem.
arXiv:2512.12500v1 Announce Type: cross
Abstract: Artificial intelligence (AI) is increasingly permeating healthcare, from physician assistants to consumer applications. Since AI algorithm's opacity challenges human interaction, explainable AI (XAI) addresses this by providing AI decision-making insight, but evidence suggests XAI can paradoxically induce over-reliance or bias. We present results from two large-scale experiments (623 lay people; 153 primary care physicians, PCPs) combining a fairness-based diagnosis AI model and different XAI explanations to examine how XAI assistance, particularly multimodal large language models (LLMs), influences diagnostic performance. AI assistance balanced across skin tones improved accuracy and reduced diagnostic disparities. However, LLM explanations yielded divergent effects: lay users showed higher automation bias - accuracy boosted when AI was correct, reduced when AI erred - while experienced PCPs remained resilient, benefiting irrespective of AI accuracy. Presenting AI suggestions first also led to worse outcomes when the AI was incorrect for both groups. These findings highlight XAI's varying impact based on expertise and timing, underscoring LLMs as a "double-edged sword" in medical AI and informing future human-AI collaborative system design.
arXiv:2512.12620v1 Announce Type: cross
Abstract: We study syllogistic reasoning in LLMs from the logical and natural language perspectives. In process, we explore fundamental reasoning capabilities of the LLMs and the direction this research is moving forward. To aid in our studies, we use 14 large language models and investigate their syllogistic reasoning capabilities in terms of symbolic inferences as well as natural language understanding. Even though this reasoning mechanism is not a uniform emergent property across LLMs, the perfect symbolic performances in certain models make us wonder whether LLMs are becoming more and more formal reasoning mechanisms, rather than making explicit the nuances of human reasoning.
arXiv:2512.12921v1 Announce Type: cross
Abstract: Artificial intelligence (AI) systems are being readily and rapidly adopted, increasingly permeating critical domains: from consumer platforms and enterprise software to networked systems with embedded agents. While this has unlocked potential for human productivity gains, the attack surface has expanded accordingly: threats now span content safety failures (e.g., harmful or deceptive outputs), model and data integrity compromise (e.g., poisoning, supply-chain tampering), runtime manipulations (e.g., prompt injection, tool and agent misuse), and ecosystem risks (e.g., orchestration abuse, multi-agent collusion). Existing frameworks such as MITRE ATLAS, National Institute of Standards and Technology (NIST) AI 100-2 Adversarial Machine Learning (AML) taxonomy, and OWASP Top 10s for Large Language Models (LLMs) and Agentic AI Applications provide valuable viewpoints, but each covers only slices of this multi-dimensional space.
This paper presents Cisco's Integrated AI Security and Safety Framework ("AI Security Framework"), a unified, lifecycle-aware taxonomy and operationalization framework that can be used to classify, integrate, and operationalize the full range of AI risks. It integrates AI security and AI safety across modalities, agents, pipelines, and the broader ecosystem. The AI Security Framework is designed to be practical for threat identification, red-teaming, risk prioritization, and it is comprehensive in scope and can be extensible to emerging deployments in multimodal contexts, humanoids, wearables, and sensory infrastructures. We analyze gaps in prevailing frameworks, discuss design principles for our framework, and demonstrate how the taxonomy provides structure for understanding how modern AI systems fail, how adversaries exploit these failures, and how organizations can build defenses across the AI lifecycle that evolve alongside capability advancements.
arXiv:2512.12950v1 Announce Type: cross
Abstract: Accurately mapping legal terminology across languages remains a significant challenge, especially for language pairs like Chinese and Japanese, which share a large number of homographs with different meanings. Existing resources and standardized tools for these languages are limited. To address this, we propose a human-AI collaborative approach for building a multilingual legal terminology database, based on a multi-agent framework. This approach integrates advanced large language models and legal domain experts throughout the entire process-from raw document preprocessing, article-level alignment, to terminology extraction, mapping, and quality assurance. Unlike a single automated pipeline, our approach places greater emphasis on how human experts participate in this multi-agent system. Humans and AI agents take on different roles: AI agents handle specific, repetitive tasks, such as OCR, text segmentation, semantic alignment, and initial terminology extraction, while human experts provide crucial oversight, review, and supervise the outputs with contextual knowledge and legal judgment. We tested the effectiveness of this framework using a trilingual parallel corpus comprising 35 key Chinese statutes, along with their English and Japanese translations. The experimental results show that this human-in-the-loop, multi-agent workflow not only improves the precision and consistency of multilingual legal terminology mapping but also offers greater scalability compared to traditional manual methods.
arXiv:2512.13438v1 Announce Type: cross
Abstract: While Large Language Model (LLM) agents show great potential for automated UI navigation such as automated UI testing and AI assistants, their efficiency has been largely overlooked. Our motivating study reveals that inefficient UI representation creates a critical performance bottleneck. However, UI representation optimization, formulated as the task of automatically generating programs that transform UI representations, faces two unique challenges. First, the lack of Boolean oracles, which traditional program synthesis uses to decisively validate semantic correctness, poses a fundamental challenge to co-optimization of token efficiency and completeness. Second, the need to process large, complex UI trees as input while generating long, compositional transformation programs, making the search space vast and error-prone. Toward addressing the preceding limitations, we present UIFormer, the first automated optimization framework that synthesizes UI transformation programs by conducting constraint-based optimization with structured decomposition of the complex synthesis task. First, UIFormer restricts the program space using a domain-specific language (DSL) that captures UI-specific operations. Second, UIFormer conducts LLM-based iterative refinement with correctness and efficiency rewards, providing guidance for achieving the efficiency-completeness co-optimization. UIFormer operates as a lightweight plugin that applies transformation programs for seamless integration with existing LLM agents, requiring minimal modifications to their core logic. Evaluations across three UI navigation benchmarks spanning Android and Web platforms with five LLMs demonstrate that UIFormer achieves 48.7% to 55.8% token reduction with minimal runtime overhead while maintaining or improving agent performance. Real-world industry deployment at WeChat further validates the practical impact of UIFormer.