❌

Reading view

Factors Influencing Continuance Intention for Online Consultations Among Survivors of Cancer: Grounded Theory Study

Background: Online consultation platforms have become an important component of survivorship care for patients with cancer, offering flexible access to oncology expertise between scheduled visits. However, evidence on what drives the willingness of survivors of cancer to continue using online consultations after initial adoption remains limited in China. A better understanding of continuance intention is needed to inform survivor-centered digital health strategies. Objective: This study aimed to explore the influencing factors of continued use of online consultations among survivors of cancer in southwest China and develop a grounded theoretical model explaining continuance intention. Methods: A grounded theory qualitative design was used. A total of 26 adult survivors of cancer with diverse demographic and clinical characteristics were purposively recruited from a tertiary cancer center in southwest China. All participants had used online consultations at least once in the preceding year. Semistructured telephone interviews were audio recorded; transcribed verbatim; and analyzed using open, axial, and selective coding with constant comparison until theoretical saturation was reached. During selective coding, categories and their relationships were integrated and iteratively refined to construct a grounded theoretical model of continuance intention. Results: Six interrelated domains influenced survivors’ continued use of online consultation platforms: platform quality, physician competence, user perception, individual condition, external context, and privacy concerns. Platform quality and physician competence influenced user perception of usefulness, reassurance, and trust, which functioned as a mediator of continued use. Individual condition, including health status, health literacy, and psychological needs, influenced both perceived usefulness and reliance on online consultations. External context, especially family encouragement, peer recommendations, and availability of local oncology services, directly facilitated or constrained continued use. Privacy concerns moderated how survivors balanced perceived benefits against risks of data misuse, stigma, and unwanted disclosure of cancer history. Survivors described online consultations as offering rapid guidance and emotional support that complemented hospital-based care but reported discontinuation when interactions were delayed or impersonal or when perceived privacy risks outweighed the benefits. Conclusions: The willingness of survivors of cancer to continue using online consultation platforms depends on multiple interrelated factors beyond traditional technological usability. Sustained engagement is shaped by survivors’ perceptions of usefulness and trust, physician empathy and timeliness, family encouragement, and acceptance of privacy trade-offs. The theoretical model advances understanding of digital health continuance in oncology and offers practical guidance for developing survivor-centered online consultation services.
  •  

Deploying a hybrid approach to Web3 in the AI era

When the concept of “Web 3.0” first emerged about a decade ago the idea was clear: Create a more user-controlled internet that lets you do everything you can now, except without servers or intermediaries to manage the flow of information.

Where Web2, which emerged in the early 2000s, relies on centralized systems to store data and supply compute, all owned—and monetized by—a handful of global conglomerates, Web3 turns that structure on its head. Instead, data and compute are decentralized through technologies like blockchain and peer-to-peer networks.

What was once a futuristic concept is quickly becoming a more concrete reality, even at a time when Web2 still dominates. Six out of ten Fortune 500 companies are exploring blockchain-based solutions, most taking a hybrid approach that combines traditional Web2 business models and infrastructure with the decentralized technologies and principles of Web3.

Popular use cases include cloud services, supply chain management, and, most notably financial services. In fact, at one point, the daily volume of transactions processed on decentralized finance exchanges exceeded $10 billion.

Gaining a Web3 edge

Among the advantages of Web3 for the enterprise are greater ownership and control of sensitive data, says Erman Tjiputra, founder and CEO of the AIOZ Network, which is building infrastructure for Web3, powered by decentralized physical infrastructure networks (DePIN), blockchain-based systems that govern physical infrastructure assets.

More cost-effective compute is another benefit, as is enhanced security and privacy as the cyberattack landscape grows more hostile, he adds. And it could even help protect companies from outages caused by a single point of failure, which can lead to downtime, data loss, and revenue deficits.

But perhaps the most exciting opportunity, says Tjiputra, is the ability to build and scale AI reliably and affordably. By leveraging a people-powered internet infrastructure, companies can far more easily access—and contribute to—shared resource like bandwidth, storage, and processing power to run AI inference, train models, and store data. All while using familiar developer tooling and open, usage-based incentives.

“We’re in a compute crunch where requirements are insatiable, and Web3 creates this ability to benefit while contributing,” explains Tjiputra.

In 2025, AIOZ Network launched a distributed compute platform and marketplace where developers and enterprises can access and monetize AI assets, and run AI inference or training on AIOZ Network’s more than 300,000 contributing devices. The model allows companies to move away from opaque datasets and models and scale flexibly, without centralized lock in.

Overcoming Web3 deployment challenges

Despite the promise, it is still early days for Web3, and core systemic challenges are leaving senior leadership and developers hesitant about its applicability at scale.

One hurdle is a lack of interoperability. The current fragmentation of blockchain networks creates a segregated ecosystem that makes it challenging to transfer assets or data between platforms. This often complicates transactions and introduces new security risks due to the reliance on mechanisms such as cross-chain bridges. These are tools that allow asset transfers between platforms but which have been shown to be vulnerable to targeted attacks.

“We have countless blockchains running on different protocols and consensus models,” says Tjiputra. “These blockchains need to work with each other so applications can communicate regardless of which chain they are on. This makes interoperability fundamental.”

Regulatory uncertainty is also a challenge. Outdated legal frameworks can sit at odds with decentralized infrastructures, especially when it comes to compliance with data protection and anti-money laundering regulations.

“Enterprises care about verifiability and compliance as much as innovation, so we need frameworks where on-chain transparency strengthens accountability instead of adding friction,” Tjiputra says.

And this is compounded by user experience (UX) challenges, says Tjiputra. “The biggest setback in Web3 today is UX,” he says. “For example, in Web2, if I forget my bank username or password, I can still contact the bank, log in and access my assets. The trade-off in Web3 is that, should that key be compromised or lost, we lose access to those assets. So, key recovery is a real problem.”

Building a bridge to Web3

Although such systemic challenges won’t be solved overnight, by leveraging DePIN networks, enterprises can bridge the gap between Web2 and Web3, without making a wholesale switch. This can minimize risk while harnessing much of the potential.

AIOZ Network’s own ecosystem includes capacity for media streaming, AI compute, and distributed storage that can be plugged into an existing Web2 tech stack. “You don’t need to go full Web3,” says Tjiputra. “You can start by plugging distributed storage into your workflow, test it, measure it, and see the benefits firsthand.”

The AIOZ Storage solution, for example, offers scalable distributed object storage by leveraging the global network of contributor devices on AIOZ DePIN. It is also compatible with existing storage systems or commonly used web application programming interfaces (APIs).

“Say we have a programmer or developer who uses Amazon S3 Storage or REST APIs, then all they need to do is just repoint the endpoints,” explains Tjiputra. “That’s it. It’s the same tools, it’s really simple. Even with media, with a single one-stop shop, developers can do transcoding and streaming with a simple REST API.”

Built on Cosmos, a network of hundreds of different blockchains that can communicate with each other, and a standardized framework enabled by Ethereum Virtual Machine (EVM), AIOZ Network has also prioritized interoperability. “Applications shouldn’t care which chain they’re on. Developers should target APIs without worrying about consensus mechanisms. That’s why we built on Cosmos and EVM—interoperability first.”

This hybrid model, which allows enterprises to use both Web2 and Web3 advantages in tandem, underpins what Tjiputra sees as the longer-term ambition for the much-hyped next iteration of the internet.

“Our vision is a truly peer-to-peer foundation for a people-powered internet, one that minimizes single points of failure through multi-region, multi-operator design,” says Tjiputra. “By distributing compute and storage across contributors, we gain both cost efficiency and end-to-end security by default.

“Ideally, we want to evolve the internet toward a more people-powered model, but we’re not there yet. We’re still at the starting point and growing.”

Indeed, Web3 isn’t quite snapping at the heels of the world’s Web2 giants, but its commercial advantages in an era of AI have become much harder to ignore. And with DePIN bridging the gap, enterprises and developers can step into that potential while keeping one foot on surer ground.

To learn more from AIOZ Network, you can read the AIOZ Network Vision Paper.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff.

This content was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

  •  

Adaptive therapy for perioperative non-small cell lung cancer: strategies guided by dynamic minimal residual disease adjustment

Transl Oncol. 2026 Jan 6;64:102660. doi: 10.1016/j.tranon.2025.102660. Online ahead of print.

ABSTRACT

Lung cancer remains the leading cause of cancer incidence and mortality worldwide, with non-small cell lung cancer (NSCLC) accounting for about 85% of cases. The low rate of early diagnosis and the high rate of occult metastases limit the survival benefits of conventional treatments. The current TNM staging system fails to fully reflect tumor heterogeneity or the dynamic molecular evolution of the disease, thus affecting the prediction of recurrence and the prognostic stratification. Some recent advances in minimal residual disease (MRD) detection, such as ultra-sensitive liquid biopsy technologies, have largely overcome the limitations of traditional imaging and offered a transformative approach for continuous, precision-based management of lung cancer. This review systematically summarized the technological evolution of MRD detection and highlighted its clinical significance in guiding adaptive therapy for NSCLC, including treatment escalation, de-escalation, and the emerging concept of precision-guided drug holidays. Moreover, the authors comprehensively discussed the "Four-Dimensional TNMB Staging System," which incorporates continuous molecular monitoring to address the static limitations of conventional staging and enhance the accuracy of prognostic stratification. Although ongoing challenges, such as the lack of standardized interpretation criteria and limited detection sensitivity, the combinations with the third-generation liquid biopsy platforms, multi-omics analyses, and multi-center prospective validation studies are expected to advance the clinical implementation of MRD-guided strategies. The paradigm change will enable the transition of NSCLC management from conventional standardized models to a precision-guided, closed-loop system of "monitoring-intervention-remonitoring," establishing a solid theoretical and practical foundation for comprehensive, molecularly driven management strategies.

PMID:41496417 | DOI:10.1016/j.tranon.2025.102660

  •  

Causal-Enhanced AI Agents for Medical Research Screening

arXiv:2601.02814v1 Announce Type: new Abstract: Systematic reviews are essential for evidence-based medicine, but reviewing 1.5 million+ annual publications manually is infeasible. Current AI approaches suffer from hallucinations in systematic review tasks, with studies reporting rates ranging from 28--40% for earlier models to 2--15% for modern implementations which is unacceptable when errors impact patient care. We present a causal graph-enhanced retrieval-augmented generation system integrating explicit causal reasoning with dual-level knowledge graphs. Our approach enforces evidence-first protocols where every causal claim traces to retrieved literature and automatically generates directed acyclic graphs visualizing intervention-outcome pathways. Evaluation on 234 dementia exercise abstracts shows CausalAgent achieves 95% accuracy, 100% retrieval success, and zero hallucinations versus 34% accuracy and 10% hallucinations for baseline AI. Automatic causal graphs enable explicit mechanism modeling, visual synthesis, and enhanced interpretability. While this proof-of-concept evaluation used ten questions focused on dementia exercise research, the architectural approach demonstrates transferable principles for trustworthy medical AI and causal reasoning's potential for high-stakes healthcare.
  •  

PatentMind: A Multi-Aspect Reasoning Graph for Patent Similarity Evaluation

arXiv:2505.19347v3 Announce Type: replace Abstract: Patent similarity evaluation plays a critical role in intellectual property analysis. However, existing methods often overlook the intricate structure of patent documents, which integrate technical specifications, legal boundaries, and application contexts. We introduce PatentMind, a novel framework for patent similarity assessment based on a Multi-Aspect Reasoning Graph (MARG). PatentMind decomposes patents into their three dimensions of technical features, application domains, and claim scopes, then dimension-specific similarity scores are calculated over the MARG. These scores are dynamically weighted through a context-aware reasoning process, which integrates contextual signals to emulate expert-level judgment. To support evaluation, we construct a human-annotated benchmark PatentSimBench, comprising 500 patent pairs. Experimental results demonstrate that the PatentMind-generated scores show a strong correlation ($r=0.938$) with expert annotations, significantly outperforming embedding-based models, patent-specific models, and advanced prompt engineering methods. Beyond computational linguistics, our framework provides a structured and semantically grounded foundation for real-world decision-making, particularly for tasks such as infringement risk assessment, underscoring its broader impact on both patent analytics and evaluation.
  •  

Organoids in translation: a bench-to-bedside framework for pancreatic cancer precision medicine

J Transl Med. 2026 Jan 6. doi: 10.1186/s12967-025-07596-8. Online ahead of print.

ABSTRACT

INTRODUCTION: Pancreatic ductal adenocarcinoma (PDAC) is one of the most lethal malignancies with a 5-year survival rate of < 13%. Standard treatments such as FOLFIRINOX or gemcitabine/nab-paclitaxel yield modest response rates, underscoring the urgent need for precision oncology approaches. Patient-derived organoids (PDOs) preserve the genomic, phenotypic, and histopathological features of the source tumor and offer a promising platform for drug screening, biomarker development, and personalized therapy. However, a systematic evaluation of their translational capacities is lacking.

METHODS: A systematic review was conducted according to the PRISMA 2020 guidelines (PROSPERO registration pending) using PubMed, EMBASE, and Cochrane CENTRAL (December 10, 2024) to identify English-language PDAC PDO studies that incorporated therapeutic testing. Ninety-five studies met the inclusion criteria. Data extraction captured >75 variables per study, including spanning culture methodology, therapeutic profiling, biomarker integration, and clinical correlation. A 13-domain weighted Translatability Scoring Framework adapted from Wehling et al. assessed predictive validity, biomarker strength, pharmacogenetics, and clinical trial alignment. Scores ranged from 0 to 5 and were categorized as good (>4.0), moderate (3.0-4.0), or low (<3.0) translational potential.

RESULTS: Of the 95 studies, 70.5% have been published since 2021, reflecting the rapid growth in this field. The mean PDO generation success rate was 89.7%, with the primary tumor tissue being the predominant source (48.4%). Only 24.8% were directly linked to clinical trials and 5.3% incorporated multi-omic profiling. The median translatability score was 3.13 (range, 1.72-4.59): 45.3% of the studies had low translatability, 50.5% moderate, and only 4.2% had good translational potential. High-scoring studies consistently combine multi-omic biomarker platforms, in vivo validation, clinical outcome correlation, and prospective trial integration. Conversely, the weakest domains were pharmacogenetics, endpoint strategies, and biomarker validation, limiting their overall clinical relevance.

CONCLUSIONS: PDOs have demonstrated strong feasibility and in vitro clinical correlation in PDAC; however, their clinical translation remains constrained by limited multi-omic integration, absence of pharmacogenomic modeling, and sparse clinical trial embedding. Standardization of protocols, adoption of harmonized and clinically relevant endpoints, and systematic incorporation of biomarker-driven co-clinical trial frameworks are urgently needed to transition PDOs from promising experimental surrogates to validating precision oncology tools capable of informing therapeutic decision-making in PDAC.

PMID:41495743 | DOI:10.1186/s12967-025-07596-8

  •  

Establishment and Optimization of a Patient-Reported Outcome–Based Electronic-Diary for Symptoms Evaluation in Patients With Gastroesophageal Reflux Disorder: Prospective Cohort Study

Background: Gastroesophageal reflux disease (GERD) symptoms significantly affect patients’ quality of life. Patient-reported outcome (PRO) instruments for symptoms measurement in GERD patients is advocated by regulatory authority. Current tools for GERD symptoms evaluation are limited and the results can be biased by the recall bias. To better characterize the GERD symptoms, an e-diary was developed for daily GERD symptom monitoring. Objective: To build up and optimize a PRO-based e-diary, and to investigate the effect of symptom frequency on adherence. Methods: The GERD e-diary evaluated 8 daytime (acid regurgitation, cough, heartburn, sour taste in the mouth, hiccups, hoarseness, dysphagia, and chest pain) and 2 nighttime symptoms (acid regurgitation and cough) for consecutive 8 weeks. The adherence of e-diary, defined as daily completing rate of e-diary, was evaluated and optimized from First Stage to Third Stage with no reminder implemented in First Stage, sending reminding SMS (Short Message Service) text messaging upon detecting missing data in Second Stage (no reminder during the first 3 to 5 days after enrollment), and immediate installation of reminding system at enrollment in Third Stage. GERD symptom frequency was obtained by summation of the symptomatic days in each week. A multiple regression analysis was performed to examine the effects of system optimization and GERD symptom frequency on patient adherence, while controlling for potential confounding variables. Results: 138 GERD patients (M/F=70/68; age: mean 52.9, SD 12.3 years) were recruited. At First Stage, the adherence was 47.2%, 40% and 57.6% for nighttime, daytime and overall symptom. System optimization significantly improved adherence with increased adherence of nighttime symptoms by 12.5% (P=.005) and 10.9% (P=.01), daytime symptom by 21.7% (P
  •  

STAT+: FDA announces sweeping changes to oversight of wearables, AI-enabled devices

LAS VEGAS — The Food and Drug Administration announced Tuesday that it will ease regulation of digital health products, following through on the Trump administration’s promises to deregulate artificial intelligence and promote its widespread use.

FDA Commissioner Marty Makary indicated that one of the agency’s priorities is fostering an environment that’s good for investors, and that FDA regulation needs to move “at Silicon Valley speed.” He announced the changes during an address to conference attendees at the Consumer Electronics Show.

The agency will soften its approach to the regulation of clinical decision support software, which include AI-enabled products that help doctors navigate diagnoses and treatment options. The agency previously considered products that delivered a single recommendation as FDA-regulated medical devices. Now, those products can enter the market without FDA review as long as they fulfill the agency’s other criteria for escaping regulation. 

Continue to STAT+ to read the full story…

© ANDREW CABALLERO-REYNOLDS/AFP via Getty Images

  •  

PRIME: an interpretable artificial intelligence model based on liquid biopsy improves prediction of progression risk in non-small cell lung cancer

Mil Med Res. 2026 Jan 6;12(1):94. doi: 10.1186/s40779-025-00679-z.

ABSTRACT

BACKGROUND: Despite the predictive impact of circulating tumor DNA (ctDNA) minimal residual disease (MRD), accurate prediction of failure risk after curative-intent treatments for early-stage or localized non-small cell lung cancer (NSCLC) patients to guide personalized therapy remains challenging. This study aimed to develop and validate an interpretable artificial intelligence-assisted model using global data resources.

METHODS: Liquid biopsy data, blood-based genomic alterations, clinicopathological features, and survival outcomes of stage I-III NSCLC patients who underwent surgery or definitive chemoradiotherapy were collected from 6 cohorts. PRIME (Progression Risk prediction by Interpretable Machine learning on ctDNA-MRD, Mutations, and clinical-therapeutic features) was trained by 6 machine learning algorithms across 4 cohorts and validated in 2 independent cohorts. Model performance was evaluated by the area under the curve (AUC) and interpreted by SHapley Additive exPlanations (SHAP). Whole-exome sequencing (WES) or whole-genome sequencing (WGS) of tumor tissue from 430 stage II-III NSCLC patients and RNA-sequencing (RNA-seq) data from 1149 subjects, sourced from The Cancer Genome Atlas, were used to validate the prognostic effect of mutations identified in peripheral blood and investigate the underlying mechanisms.

RESULTS: A global dataset encompassing 781 blood samples from 493 patients was analyzed. Clinical stage, pre-treatment ctDNA, post-treatment MRD, blood-based Kelch-like ECH-associated protein 1 (KEAP1), serine/threonine kinase 11 (STK11), and cyclin-dependent kinase inhibitor 2A (CDKN2A) mutations, and treatment modality were significantly associated with the risk of disease progression and were thereby included in the model training. WES/WGS and RNA-seq confirmed the poor prognostic effect of KEAP1, STK11, and CDKN2A mutations, which were characterized by the suppressive tumor microenvironment and attenuated humoral immunity. The neural network (NN) model exhibited optimal prediction of treatment failure risk in the training (AUC = 0.85, 95% CI 0.81-0.89) and validation sets (AUC = 0.82, 95% CI 0.74-0.89). SHAP analysis indicated that MRD (+0.306), treatment modality (+0.128), and pre-treatment ctDNA (+0.043) ranked in the top 3 contributions. NN-PRIME outperformed single liquid biopsy biomarkers and clinical-therapeutic signatures, and demonstrated consistent robustness across different clinical scenarios. High-risk patients identified by NN-PRIME had poorer prognoses but derived significant benefits from adjuvant therapy after surgery.

CONCLUSIONS: As an interpretable model integrating readily-accessible and crucial clinical-genomic predictors, PRIME achieves enhanced performance, allowing for early outcome prediction, refined risk stratification, and personalized clinical decision-making.

PMID:41491583 | PMC:PMC12771999 | DOI:10.1186/s40779-025-00679-z

  •  

Agentic AI for Autonomous, Explainable, and Real-Time Credit Risk Decision-Making

arXiv:2601.00818v1 Announce Type: new Abstract: Significant digitalization of financial services in a short period of time has led to an urgent demand to have autonomous, transparent and real-time credit risk decision making systems. The traditional machine learning models are effective in pattern recognition, but do not have the adaptive reasoning, situational awareness, and autonomy needed in modern financial operations. As a proposal, this paper presents an Agentic AI framework, or a system where AI agents view the world of dynamic credit independent of human observers, who then make actions based on their articulable decision-making paths. The research introduces a multi-agent system with reinforcing learning, natural language reasoning, explainable AI modules, and real-time data absorption pipelines as a means of assessing the risk profiles of borrowers with few humans being involved. The processes consist of agent collaboration protocol, risk-scoring engines, interpretability layers, and continuous feedback learning cycles. Findings indicate that decision speed, transparency and responsiveness is better than traditional credit scoring models. Nevertheless, there are still some practical limitations such as risks of model drift, inconsistencies in interpreting high dimensional data and regulatory uncertainties as well as infrastructure limitations in low-resource settings. The suggested system has a high prospective to transform credit analytics and future studies ought to be directed on dynamic regulatory compliance mobilizers, new agent teamwork, adversarial robustness, and large-scale implementation in cross-country credit ecosystems.
  •  

Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models

arXiv:2601.01321v1 Announce Type: new Abstract: Digital twins, as precise digital representations of physical systems, have evolved from passive simulation tools into intelligent and autonomous entities through the integration of artificial intelligence technologies. This paper presents a unified four-stage framework that systematically characterizes AI integration across the digital twin lifecycle, spanning modeling, mirroring, intervention, and autonomous management. By synthesizing existing technologies and practices, we distill a unified four-stage framework that systematically characterizes how AI methodologies are embedded across the digital twin lifecycle: (1) modeling the physical twin through physics-based and physics-informed AI approaches, (2) mirroring the physical system into a digital twin with real-time synchronization, (3) intervening in the physical twin through predictive modeling, anomaly detection, and optimization strategies, and (4) achieving autonomous management through large language models, foundation models, and intelligent agents. We analyze the synergy between physics-based modeling and data-driven learning, highlighting the shift from traditional numerical solvers to physics-informed and foundation models for physical systems. Furthermore, we examine how generative AI technologies, including large language models and generative world models, transform digital twins into proactive and self-improving cognitive systems capable of reasoning, communication, and creative scenario generation. Through a cross-domain review spanning eleven application domains, including healthcare, aerospace, smart manufacturing, robotics, and smart cities, we identify common challenges related to scalability, explainability, and trustworthiness, and outline directions for responsible AI-driven digital twin systems.
  •  

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

arXiv:2601.01330v1 Announce Type: new Abstract: Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM routing and aggregation at scale and identify three key bottlenecks: (1) current train-free routers are limited by a query-based paradigm focusing solely on textual similarity; (2) recent aggregation methods remain largely static, failing to select appropriate aggregators for different tasks;(3) the complementarity of routing and aggregation remains underutilized. To address these problems, we introduce JiSi, a novel framework designed to release the full potential of LLMs' collaboration through three innovations: (1) Query-Response Mixed Routing capturing both semantic information and problem difficulty; (2) Support-Set-based Aggregator Selection jointly evaluating the aggregation and domain capacity of aggregators; (3) Adaptive Routing-Aggregation Switch dynamically leveraging the advantages of routing and aggregation. Comprehensive experiments on nine benchmarks demonstrate that JiSi can surpass Gemini-3-Pro with only 47% costs by orchestrating ten open-source LLMs, while outperforming mainstream baselines. It suggests that collective intelligence represents a novel path towards Artificial General Intelligence (AGI).
  •  

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

arXiv:2601.01718v1 Announce Type: new Abstract: We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks while maintaining competitive capabilities on general-purpose tasks. To address the overthinking phenomenon commonly observed in Large Reasoning Models (LRMs), we propose Reflection-aware Adaptive Policy Optimization (RAPO), a novel RL training algorithm that effectively regulates overthinking behaviors. In enterprise-oriented tasks such as retrieval-augmented generation (RAG), complex table understanding, and summarization, Yuan3.0 Flash consistently achieves superior performance. Moreover, it also demonstrates strong reasoning capabilities in domains such as mathematics, science, etc., attaining accuracy comparable to frontier model while requiring only approximately 1/4 to 1/2 of the average tokens. Yuan3.0 Flash has been fully open-sourced to facilitate further research and real-world deployment: https://github.com/Yuan-lab-LLM/Yuan3.0.
  •  

Correctness isnt Efficiency: Runtime Memory Divergence in LLM-Generated Code

arXiv:2601.01215v1 Announce Type: cross Abstract: Large language models (LLMs) can generate programs that pass unit tests, but passing tests does not guarantee reliable runtime behavior. We find that different correct solutions to the same task can show very different memory and performance patterns, which can lead to hidden operational risks. We present a framework to measure execution-time memory stability across multiple correct generations. At the solution level, we introduce Dynamic Mean Pairwise Distance (DMPD), which uses Dynamic Time Warping to compare the shapes of memory-usage traces after converting them into Monotonic Peak Profiles (MPPs) to reduce transient noise. Aggregating DMPD across tasks yields a model-level Model Instability Score (MIS). Experiments on BigOBench and CodeContests show substantial runtime divergence among correct solutions. Instability often increases with higher sampling temperature even when pass@1 improves. We also observe correlations between our stability measures and software engineering indicators such as cognitive and cyclomatic complexity, suggesting links between operational behavior and maintainability. Our results support stability-aware selection among passing candidates in CI/CD to reduce operational risk without sacrificing correctness. Artifacts are available.
  •  

OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment

arXiv:2601.01576v1 Announce Type: cross Abstract: Evaluating novelty is critical yet challenging in peer review, as reviewers must assess submissions against a vast, rapidly evolving literature. This report presents OpenNovelty, an LLM-powered agentic system for transparent, evidence-based novelty analysis. The system operates through four phases: (1) extracting the core task and contribution claims to generate retrieval queries; (2) retrieving relevant prior work based on extracted queries via semantic search engine; (3) constructing a hierarchical taxonomy of core-task-related work and performing contribution-level full-text comparisons against each contribution; and (4) synthesizing all analyses into a structured novelty report with explicit citations and evidence snippets. Unlike naive LLM-based approaches, \textsc{OpenNovelty} grounds all assessments in retrieved real papers, ensuring verifiable judgments. We deploy our system on 500+ ICLR 2026 submissions with all reports publicly available on our website, and preliminary analysis suggests it can identify relevant prior work, including closely related papers that authors may overlook. OpenNovelty aims to empower the research community with a scalable tool that promotes fair, consistent, and evidence-backed peer review.
  •  

Multimodal Fact-Checking: An Agent-based Approach

arXiv:2512.22933v3 Announce Type: replace Abstract: The rapid spread of multimodal misinformation poses a growing challenge for automated fact-checking systems. Existing approaches, including large vision language models (LVLMs) and deep multimodal fusion methods, often fall short due to limited reasoning and shallow evidence utilization. A key bottleneck is the lack of dedicated datasets that provide complete real-world multimodal misinformation instances accompanied by annotated reasoning processes and verifiable evidence. To address this limitation, we introduce RW-Post, a high-quality and explainable dataset for real-world multimodal fact-checking. RW-Post aligns real-world multimodal claims with their original social media posts, preserving the rich contextual information in which the claims are made. In addition, the dataset includes detailed reasoning and explicitly linked evidence, which are derived from human written fact-checking articles via a large language model assisted extraction pipeline, enabling comprehensive verification and explanation. Building upon RW-Post, we propose AgentFact, an agent-based multimodal fact-checking framework designed to emulate the human verification workflow. AgentFact consists of five specialized agents that collaboratively handle key fact-checking subtasks, including strategy planning, high-quality evidence retrieval, visual analysis, reasoning, and explanation generation. These agents are orchestrated through an iterative workflow that alternates between evidence searching and task-aware evidence filtering and reasoning, facilitating strategic decision-making and systematic evidence analysis. Extensive experimental results demonstrate that the synergy between RW-Post and AgentFact substantially improves both the accuracy and interpretability of multimodal fact-checking.
  •  

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

arXiv:2506.05623v2 Announce Type: replace-cross Abstract: Infrastructure-as-Code (IaC) generation holds significant promise for automating cloud infrastructure provisioning. Recent advances in Large Language Models (LLMs) present a promising opportunity to democratize IaC development by generating deployable infrastructure templates from natural language descriptions. However, current evaluation focuses on syntactic correctness while ignoring deployability, the critical measure of the utility of IaC configuration files. Six state-of-the-art LLMs performed poorly on deployability, achieving only 20.8$\sim$30.2% deployment success rate on the first attempt. In this paper, we construct DPIaC-Eval, the first deployability-centric IaC template benchmark consisting of 153 real-world scenarios cross 58 unique services. Also, we propose an LLM-based deployability-centric framework, dubbed IaCGen, that uses iterative feedback mechanism encompassing format verification, syntax checking, and live deployment stages, thereby closely mirroring the real DevOps workflows. Results show that IaCGen can make 54.6$\sim$91.6% generated IaC templates from all evaluated models deployable in the first 10 iterations. Additionally, human-in-the-loop feedback that provide direct guidance for the deployability errors, can further boost the performance to over 90% passItr@25 on all evaluated LLMs. Furthermore, we explore the trustworthiness of the generated IaC templates on user intent alignment and security compliance. The poor performance (25.2% user requirement coverage and 8.4% security compliance rate) indicates a critical need for continued research in this domain.
  •  
❌