Normal view
-
Cell
-
Avoiding common failures in AI for health and medicine
Salaudeen et al. review common reliability failures in predictive and generative AI for healthcare, including erroneous model outputs, clinically unjustified performance differences, and deployment-time degradation. They examine why existing technical solutions fall short and argue for lifecycle-aware evaluation, continuous monitoring, and institutional governance.
-
InfoQ

-
Dropbox Evolves Riviera Content Processing Platform to Support AI Workloads
Dropbox has evolved Riviera from a file preview service into a universal content processing platform supporting more than 300 file formats and over 100 transformation capabilities. Processing hundreds of thousands of transformations per second, Riviera now supports Search, Replay, Sign, and Dash, while its APIs enable asynchronous content extraction for AI and RAG workflows. By Leela Kumili
Dropbox Evolves Riviera Content Processing Platform to Support AI Workloads
Dropbox has evolved Riviera from a file preview service into a universal content processing platform supporting more than 300 file formats and over 100 transformation capabilities. Processing hundreds of thousands of transformations per second, Riviera now supports Search, Replay, Sign, and Dash, while its APIs enable asynchronous content extraction for AI and RAG workflows.
By Leela Kumili-
Nature Medicine
-
Broadly neutralizing antibodies in adult males living with HIV undergoing analytical treatment interruption: secondary and exploratory outcomes of the phase II randomized controlled RIO trial
Nature Medicine, Published online: 15 September 2026; doi:10.1038/s41591-026-04644-8In the phase 2 RIO trial, there was delayed viral rebound and resistance to broadly neutralizing antibodies 3BNC117-LS and 10-1074-LS in adult males living with HIV undergoing analytical treatment interruption, and initial reservoir sensitivity to autologous antibodies was associated with a longer time to rebound.
Broadly neutralizing antibodies in adult males living with HIV undergoing analytical treatment interruption: secondary and exploratory outcomes of the phase II randomized controlled RIO trial
Nature Medicine, Published online: 15 September 2026; doi:10.1038/s41591-026-04644-8
In the phase 2 RIO trial, there was delayed viral rebound and resistance to broadly neutralizing antibodies 3BNC117-LS and 10-1074-LS in adult males living with HIV undergoing analytical treatment interruption, and initial reservoir sensitivity to autologous antibodies was associated with a longer time to rebound.-
InfoQ

-
Agoda Replaces 72-Shard SQL Server Price Cache with DragonflyDB
Agoda migrated its 1.5 TB hotel Price Cache from 72 SQL Server shards to DragonflyDB to handle growing read and write volumes. The migration used staged dual reads, parity validation, gradual traffic shifting, and decentralized failover detection. Agoda reports an approximately eightfold reduction in P99 read latency, with two DragonflyDB clusters providing high availability. By Leela Kumili
Agoda Replaces 72-Shard SQL Server Price Cache with DragonflyDB
Agoda migrated its 1.5 TB hotel Price Cache from 72 SQL Server shards to DragonflyDB to handle growing read and write volumes. The migration used staged dual reads, parity validation, gradual traffic shifting, and decentralized failover detection. Agoda reports an approximately eightfold reduction in P99 read latency, with two DragonflyDB clusters providing high availability.
By Leela Kumili-
cs.AI, q-bio.NC updates on arXiv.org
-
When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration
arXiv:2609.12482v1 Announce Type: new Abstract: We aim to characterise the value of artificial intelligence in the workplace. Current studies largely measure this value in terms of the current automation capabilities and public adoption of AI. However, such metrics ignore the greater impacts of human--agent collaboration in transforming the nature of work. To account for this, we must expand the scope of our analysis beyond atomised tasks of today, and instead focus on how AI can augment entire
When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration
-
cs.AI, q-bio.NC updates on arXiv.org
-
Confidence-Gated Transductive Test Generation for Code Reranking
arXiv:2609.12489v1 Announce Type: new Abstract: Test case synthesis is crucial for evaluating and ranking programs generated by large language models (LLMs). However, constructing high-quality test cases remains challenging because reliable expected outputs are often difficult to obtain. We propose Confidence-Gated Transductive Test Generation (CoTT), which first uses an efficient inductive procedure and invokes transductive generation only when inductive confidence is low. This adaptive design
Confidence-Gated Transductive Test Generation for Code Reranking
-
cs.AI, q-bio.NC updates on arXiv.org
-
MPT: Missing Prototype Tracking via Barycentric Reconstruction in Vehicular Federated Learning
arXiv:2609.12771v1 Announce Type: new Abstract: Cross-vehicle federated learning enables vehicles to collaboratively improve perception models while keeping locally collected driving data private. However, vehicle participation is transient, and a vehicle may depart before training converges while permanently taking its local data. When this departing vehicle holds most samples of a target class, the class becomes rare in the remaining FL network, and its recognition can silently degrade as the
MPT: Missing Prototype Tracking via Barycentric Reconstruction in Vehicular Federated Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Agentic TCAD Calibration Workflow for Oxide Semiconductor Transistors
arXiv:2609.12184v1 Announce Type: cross Abstract: Experimental TCAD calibration is essential for predictive technology modeling of emerging oxide semiconductor transistors. However, it remains time-consuming and expert dependent because of model ambiguity. Multiple physical models and parameter sets can reproduce the same measured transfer characteristics, while local fitting alone cannot uniquely identify the underlying device physics. We present the first demonstration of an agentic TCAD cali
Agentic TCAD Calibration Workflow for Oxide Semiconductor Transistors
-
cs.AI, q-bio.NC updates on arXiv.org
-
Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning
arXiv:2609.12584v1 Announce Type: cross Abstract: Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP ins
Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning
-
cs.AI, q-bio.NC updates on arXiv.org
-
Explaining Time Series Forecasting with Horizon-Resolved Attribution
arXiv:2609.12639v1 Announce Type: cross Abstract: Recent advances in explaining time series (TS) models have produced methods that identify which past values a prediction depends on. However, most existing methods return a single importance vector, assuming that every predicted step depends on the same past values. In this paper, we show that this assumption does not hold, as different forecast steps depend on different past values. Motivated by this observation, we propose Horizon-Resolved eXp
Explaining Time Series Forecasting with Horizon-Resolved Attribution
-
cs.AI, q-bio.NC updates on arXiv.org
-
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
arXiv:2609.13053v1 Announce Type: cross Abstract: Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
-
cs.AI, q-bio.NC updates on arXiv.org
-
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
arXiv:2609.13076v1 Announce Type: cross Abstract: Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interacti
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
-
cs.AI, q-bio.NC updates on arXiv.org
-
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
arXiv:2607.05915v4 Announce Type: replace Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still lag behind rule-based routers. We introduce PCBWorld, an open-source engine-grounded PCB routing environment built on KiCad, an electronic design automation (EDA) engine. As a human engineer does, agents in PCBWorld interactively route a board through the engine's native operations, guided by its Design Rule
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
-
cs.AI, q-bio.NC updates on arXiv.org
-
TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
arXiv:2504.04981v3 Announce Type: replace-cross Abstract: This paper studies continual test-time adaptation (CTTA), the task of adapting a model to constantly changing unseen domains in testing while preserving previously learned knowledge. Existing CTTA methods mostly focus on adaptation to the current test domain only, overlooking generalization to arbitrary test domains a model may face in the future. To tackle this limitation, we present a novel online test-time domain generalization framew
TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
-
cs.AI, q-bio.NC updates on arXiv.org
-
Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
arXiv:2604.16775v2 Announce Type: replace-cross Abstract: Generative medical event models use tokenized sequences of patient timelines as input, but practical guidance on the many decisions around tokenization is limited. We benchmark quantization granularity, reference-range anchoring, code--value fusion, numeric and temporal encodings, and native versus harmonized event representations from an expert-mapped common data model. Using both Llama and Qwen architectures, 156 models were trained fr
Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
-
cs.AI, q-bio.NC updates on arXiv.org
-
ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
arXiv:2604.19254v2 Announce Type: replace-cross Abstract: Popular low-rank parameter-efficient fine-tuning (PEFT) methods represent adaptation as separate updates to selected backbone weights, without maintaining an explicit task-specific state that is updated and reused across depth. These updates also require the backbone at inference and therefore cannot operate as standalone predictors. We propose ShadowPEFT, which consolidates trainable adaptation into a modular shadow component centered o
ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
-
cs.AI, q-bio.NC updates on arXiv.org
-
ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces
arXiv:2606.05402v2 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process. We introduce ReasoningFlow, a framework that captures the discourse structures of LRM reasoning traces into fine-grained directed acyclic graphs (DAGs). We develop and validate our annotation schema through careful manual annotation of 31 trace
ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces
-
cs.AI, q-bio.NC updates on arXiv.org
-
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
arXiv:2607.05276v2 Announce Type: replace-cross Abstract: Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications. We introduce ProPS, Prompted Profile Synthesis, a framework for generating distributions of speaker embeddings conditioned on natural language pr
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
-
cs.AI, q-bio.NC updates on arXiv.org
-
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
arXiv:2609.10410v2 Announce Type: replace-cross Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
-
Nature - Issue - nature.com science feeds
-
UK universities are on the brink — what is their greatest source of stress?
Nature, Published online: 14 September 2026; doi:10.1038/d41586-026-02608-zNature asked eight leaders in higher education what they thought about the state of UK institutions. Here’s how they answered.
UK universities are on the brink — what is their greatest source of stress?
Nature, Published online: 14 September 2026; doi:10.1038/d41586-026-02608-z
Nature asked eight leaders in higher education what they thought about the state of UK institutions. Here’s how they answered.