❌

Normal view

Off-Policy Evaluation and Learning for Survival Outcomes under Censoring

arXiv:2603.22900v1 Announce Type: cross Abstract: Optimizing survival outcomes, such as patient survival or customer retention, is a critical objective in data-driven decision-making. Off-Policy Evaluation~(OPE) provides a powerful framework for assessing such decision-making policies using logged data alone, without the need for costly or risky online experiments in high-stakes applications. However, typical estimators are not designed to handle right-censored survival outcomes, as they ignore unobserved survival times beyond the censoring time, leading to systematic underestimation of the true policy performance. To address this issue, we propose a novel framework for OPE and Off-Policy Learning~(OPL) tailored for survival outcomes under censoring. Specifically, we introduce IPCW-IPS and IPCW-DR, which employ the Inverse Probability of Censoring Weighting technique to explicitly deal with censoring bias. We theoretically establish that our estimators are unbiased and that IPCW-DR achieves double robustness, ensuring consistency if either the propensity score or the outcome model is correct. Furthermore, we extend this framework to constrained OPL to optimize policy value under budget constraints. We demonstrate the effectiveness of our proposed methods through simulation studies and illustrate their practical impacts using public real-world data for both evaluation and learning tasks.

Orientability of Causal Relations in Time Series using Summary Causal Graphs and Faithful Distributions

arXiv:2508.21742v2 Announce Type: replace Abstract: Understanding causal relations between temporal variables is a central challenge in time series analysis, particularly when the full causal structure is unknown. Even when the full causal structure cannot be fully specified, experts often succeed in providing a high-level abstraction of the causal graph, known as a summary causal graph, which captures the main causal relations between different time series while abstracting away micro-level details. In this work, we present conditions that guarantee the orientability of micro-level edges between temporal variables given the background knowledge encoded in a summary causal graph and assuming having access to a faithful and causally sufficient distribution with respect to the true unknown graph. Our results provide theoretical guarantees for edge orientation at the micro-level, even in the presence of cycles or bidirected edges at the macro-level. These findings offer practical guidance for leveraging SCGs to inform causal discovery in complex temporal systems and highlight the value of incorporating expert knowledge to improve causal inference from observational time series data.
  • βœ‡cs.AI, q-bio.NC updates on arXiv.org
  • Conformal Tradeoffs: Guarantees Beyond Coverage Petrus H. Zwart
    arXiv:2602.18045v2 Announce Type: replace-cross Abstract: Deployed conformal predictors are long-lived decision infrastructure reused over finite operational windows. In practice, stakeholders care not only about marginal coverage, but also about operational quantities: how often the system commits versus defers, and what error exposure it induces when it acts. These deployment-facing quantities are not determined by coverage alone: identical calibrated thresholds can yield markedly different o
     

Conformal Tradeoffs: Guarantees Beyond Coverage

arXiv:2602.18045v2 Announce Type: replace-cross Abstract: Deployed conformal predictors are long-lived decision infrastructure reused over finite operational windows. In practice, stakeholders care not only about marginal coverage, but also about operational quantities: how often the system commits versus defers, and what error exposure it induces when it acts. These deployment-facing quantities are not determined by coverage alone: identical calibrated thresholds can yield markedly different operational profiles depending on score geometry. We develop tools for operational certification and planning beyond coverage for split conformal prediction. First, Small-Sample Beta Correction (SSBC) inverts the exact finite-sample rank/Beta law to map a user request $(\alpha^\star,\delta)$ to a concrete calibration grid point with PAC-style semantics, yielding explicit finite-window coverage guarantees for a reused deployed rule. Second, because no distribution-free pivot exists beyond coverage, we propose Calibrate-and-Audit: an independent audit set supports certified finite-window predictive envelopes (Binomial/Beta-Binomial) for key operational quantities -- commitment frequency, deferral, and decisive error exposure -- and related metrics via linear projection, without committing to a scalar objective. Third, we give a geometric characterization of the feasibility constraints and regime boundaries induced by a fixed conformal partition, clarifying why operational quantities are coupled and how calibration navigation trades them off. The result is an operational menu rthat traces attainable operational profiles (Pareto trade-offs) and attach finite-window uncertainty envelopes to each regime. We illustrate the approach on benchmark molecular toxicity and aqueous solubility datasets.

Detecting Structural Heart Disease from Electrocardiograms via a Generalized Additive Model of Interpretable Foundation-Model Predictors

arXiv:2603.02616v1 Announce Type: cross Abstract: Structural heart disease (SHD) is a prevalent condition with many undiagnosed cases, and early detection is often limited by the high cost and accessibility constraints of echocardiography (ECHO). Recent studies show that artificial intelligence (AI)-based analysis of electrocardiograms (ECGs) can detect SHD, offering a scalable alternative. However, existing methods are fully black-box models, limiting interpretability and clinical adoption. To address these challenges, we propose an interpretable and effective framework that integrates clinically meaningful ECG foundation-model predictors within a generalized additive model, enabling transparent risk attribution while maintaining strong predictive performance. Using the EchoNext benchmark of over 80,000 ECG-ECHO pairs, the method demonstrates relative improvements of +0.98% in AUROC, +1.01% in AUPRC, and +1.41% in F1 score over the latest state-of-the-art deep-learning baseline, while achieving slightly better performance even with only 30% of the training data. Subgroup analyses confirm robust performance across heterogeneous populations, and the estimated entry-wise functions provide interpretable insights into the relationships between risks of traditional ECG diagnoses and SHD. This work illustrates a complementary paradigm between classical statistical modeling and modern AI, offering a pathway to interpretable, high-performing, and clinically actionable ECG-based SHD screening.
  • βœ‡cs.AI, q-bio.NC updates on arXiv.org
  • On the Granularity of Causal Effect Identifiability Yizuo Chen Β· Adnan Darwiche
    arXiv:2510.16703v2 Announce Type: replace-cross Abstract: The classical notion of causal effect identifiability is defined in terms of treatment and outcome variables. In this paper, we consider the identifiability of state-based causal effects: how an intervention on a particular state of treatment variables affects a particular state of outcome variables. We demonstrate that state-based causal effects may be identifiable even when variable-based causal effects may not. Moreover, we show that
     

On the Granularity of Causal Effect Identifiability

arXiv:2510.16703v2 Announce Type: replace-cross Abstract: The classical notion of causal effect identifiability is defined in terms of treatment and outcome variables. In this paper, we consider the identifiability of state-based causal effects: how an intervention on a particular state of treatment variables affects a particular state of outcome variables. We demonstrate that state-based causal effects may be identifiable even when variable-based causal effects may not. Moreover, we show that this separation occurs only when additional knowledge -- such as context-specific independencies -- is available. We further examine knowledge that constrains the states of variables, and show that such knowledge can improve both variable-based and state-based identifiability when combined with other knowledge such as context-specific independencies. We finally propose an approach for identifying causal effects under these additional constraints, and conduct empirical studies to further illustrate the separations between the two levels of identifiability.

Information-Theoretic Causal Bounds under Unmeasured Confounding

arXiv:2601.17160v3 Announce Type: replace-cross Abstract: We develop a data-driven information-theoretic framework for sharp partial identification of causal effects under unmeasured confounding. Existing approaches often rely on restrictive assumptions, such as bounded or discrete outcomes; require external inputs (for example, instrumental variables, proxies, or user-specified sensitivity parameters); necessitate full structural causal model specifications; or focus solely on population-level averages while neglecting covariate-conditional effects. We overcome all four limitations simultaneously by establishing novel information-theoretic, data-driven divergence bounds. Our key theoretical contribution shows that the f-divergence between the observational distribution P(Y | A = a, X = x) and the interventional distribution P(Y | do(A = a), X = x) is upper bounded by a function of the propensity score alone. This result enables sharp partial identification of conditional causal effects directly from observational data, without requiring external sensitivity parameters, auxiliary variables, full structural specifications, or outcome boundedness assumptions. For practical implementation, we develop a semiparametric estimator satisfying Neyman orthogonality (Chernozhukov et al., 2018), which ensures root-n consistent inference even when nuisance functions are estimated via flexible machine learning methods. Simulation studies and real-world data applications, implemented in the GitHub repository (https://github.com/yonghanjung/Information-Theretic-Bounds), demonstrate that our framework provides tight and valid causal bounds across a wide range of data-generating processes.

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

arXiv:2602.14862v1 Announce Type: cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilistic models. It is mostly used in two contexts: improving the calibration of classifiers and tuning the stochasticity of large language models (LLMs). In both cases, temperature scaling is the most popular method for the job. Despite its popularity, a rigorous theoretical analysis of the properties of temperature scaling has remained elusive. We investigate here some of these properties. For classification, we show that increasing the temperature increases the uncertainty in the model in a very general sense (and in particular increases its entropy). However, for LLMs, we challenge the common claim that increasing temperature increases diversity. Furthermore, we introduce two new characterisations of temperature scaling. The first one is geometric: the tempered model is shown to be the information projection of the original model onto the set of models with a given entropy. The second characterisation clarifies the role of temperature scaling as a submodel of more general linear scalers such as matrix scaling and Dirichlet calibration: we show that temperature scaling is the only linear scaler that does not change the hard predictions of the model.
❌