Program
SAVI 2026 is over! Thanks to everyone who participated and for making this an awesome experience!
You can also subscribe to the iCal version.
| Time | Monday | Tuesday | Wednesday | Thursday | Friday |
|---|---|---|---|---|---|
| 09:00 | Opening | KeynoteSnigdha Panigrahi | Talks C1 | KeynoteGert de Cooman | KeynoteVictor H. de la Peña |
| ca. 09:15KeynoteGergely Neu | |||||
| 09:30 | |||||
| 10:00 | Coffee Break | Poster Pitches | Coffee Break | Coffee Break | |
| Coffee Break | |||||
| 10:30 | Talks B1 | Posters C | Talks D1 | Talks E1 | |
| ca. 10:45Talks A1 | |||||
| 11:00 | Open Problem Session | ||||
| 11:30 | |||||
| ca. 11:45KeynoteMorgane Austern | |||||
| 12:00 | Closing | ||||
| Lunch | |||||
| 12:30 | Lunch | Lunch | Lunch | ||
| Lunch | |||||
| 13:00 | |||||
| 13:30 | Talks A2 | Talks B2 | Talks D2 | End | |
| Excursion | |||||
| 14:00 | |||||
| 14:30 | Poster Pitches | Poster Pitches | Poster Pitches | ||
| 15:00 | Posters A | Posters B | Posters D | ||
| 15:30 | |||||
| 16:00 | |||||
| 16:30 | Talks A3 | Talks B3 | Talks D3 | ||
| 17:00 | |||||
| 17:30 | Banquet | ||||
| 18:00 | End | End | End |
- Monday: Bandits & allocation · Game-theoretic statistics · Conformal prediction & forecasting · Confidence sequences & concentration · Sequential testing and adaptive inference
- Tuesday: Multiple testing & FDR · E-process optimality · Closed testing and e-closure · Conditional/sequential inference · Clinical and applied anytime-valid inference
- Wednesday: Foundations of e-values · Universal inference & likelihood ratios · Sequential testing and changepoints · Post-hoc inference & evidence synthesis · Formal methods and e-statistics · Social activity & dinner
- Thursday: Imprecise probabilities · Clinical trials and adaptive designs · Evidence, post-hoc inference & optimal e-values · Decision-making, voting and auditing · Prediction, forecasting and ML applications · Concentration, self-normalization & conditional inference
- Friday: Foundations of anytime-valid inference and future directions
Printing Posters
Posters can be printed at the University of Twente Xerox shop. Send an email with your poster as pdf to xeroxcarre@utwente.nl. Xerox has indicated that they are very busy in the week before and during the conference, so it can take a few days for a poster to be processed and printed. Xerox is situated in the building Carré and is open from 9:15 AM – 1 PM and 1:30 PM – 4:30 PM on Weekdays.
UPDATE from experience (by Rianne): Printing takes 20 minutes if you are the first in the queue. On Monday there was no queue. Walking there took me 8 minutes (one way). If you send your poster by mail, they will e-mail you when it’s ready. When you enter the main entrance of Hal B, go straight ahead up the first flight of stairs, and then immediately to the right.
Excursion and Banquet
The excursion start at 15.15 at Oyfo Techniekmuseum, Hazemeijerstraat 300, Hengelo. The organisers are going there by different means of transport, and are happy to take you along. We’ll keep you posted about where to gather at what time. The Banquet follows the excursion (it starts at 18.00) and takes place at: Lust, Bakery en Bistro.
Getting there
You can join (one of) your favourite organiser(s) going by your favourite means of transport.
Information from the organizers during the conference
Signal Group
Join our SAVI 2026 Signal Group for connecting informally with other participants.
Open problem session
You can write your name on the flipover for signing up for presenting an open problem on Friday.
Group picture
We will make a group picture on Thursday at 15.00 (right after the poster pitches) at the blue stairs just outside the terrace of of the restaurant.
Impromptu dinner groups on Thursday:
- Group with Rianne (exactly the 30 people who signed up Thu morning): leave together (for the bus) at 18.20 at the UPark entrance, or arrive yourself at 19.00 at Het Paradijs, Nicolaas Beetsstraat 48, Enschede.
- Group with Wouter (random number of participants, open for more participants): follow Wouter.
Certificate of attendance
Send an e-mail to Rianne if you need one with all your personal details that should be on there.
Abstracts
Keynotes
Snigdha Panigrahi (Flexible selective inference with conditional validity)
When the same dataset is used both to select a model or parameters and to perform inference, naïve tests and interval estimates can be severely misleading. Data carving offers a principled way to reuse the data, conditioning on the event that the model was selected and basing inference on the resulting conditional distribution. In practice, however, data carving has been notoriously difficult to use. Existing methods require an analytic characterization of the selection event, which is available only for certain selection procedures and often requires highly selection-specific derivations. This raises the question: can we still provide selective inference through data carving without ever requiring an analytic description of the selection event?
In this talk, I will present two new approaches to selective inference that accomplish exactly this. These methods: (i) enable conditional inference without requiring an analytic description of the selection event, (ii) avoid trading selection quality for inferential power, unlike data splitting approaches, and (iii) eliminate the unnecessary conditioning that limits many existing approaches.
Gergely Neu (Online-to-X conversions: A recipe for algorithmic statistics)
A recent line of work in mathematical statistics has uncovered a curious recipe for proving complex statistical claims via a decomposition into two parts: a “simple” statistical claim and a “more complex” claim that can be addressed using the theory of algorithms. This technique has been successfully used for addressing a variety of classic problems such as mean estimation, parameter-estimation of linear models, or generalization bounds for statistical learning theory. In this talk, I will survey some existing use cases and attempt to connect the dots by dreaming up a theory of “algorithmic statistics”.
Gert de Cooman (Views from the border)
For more than a decade now, there have been efforts to combine ideas from the fields of imprecise probabilities and game-theoretic probability. This has enriched both fields, but has also resulted in advances in stochastic processes (imprecise Markov chains) and algorithmic randomness (imprecise randomness). I want to argue that the connection can also help foster ideas in statistics, by drawing from imprecise probabilities, algorithmic randomness and robust statistics. My aim in the talk will be to focus on ideas rather than on the very technical details.
Morgane Austern (Efficient finite sample bounds via optimal transport)
Finite sample bounds are ubiquitous in statistics and machine learning, underpinning applications ranging from multi-armed bandit problems to early stopping rules. However, classical bounds are often overly conservative, leading to suboptimal algorithms. In this talk, I will propose a method for deriving sharper bounds by bridging the gap between asymptotic limit theorems and finite-sample concentration. We achieve this by exploiting recent advances in Optimal Transport, Stein method and information theory. The resulting bounds are efficient, strictly valid in the finite-sample regime, and significantly tighter than the state-of-the-art. I will demonstrate how these bounds lead to direct algorithmic improvements. I will then explore how we can generalize those ideas to valid anytime guarantees, to dependent data and to more complex estimators.
Victor H. de la Peña (From Decoupling to Ville’s Equality: A Personal History of Foundational Inequalities for Anytime-Valid Inference)
In this keynote I revisit, from a personal vantage point, the development of inequalities now foundational to anytime-valid inference. The first part begins with decoupling: by comparing dependent processes with their conditionally independent (tangent) counterparts, I obtained a first general class of exponential inequalities for self-normalized martingales. I then describe how pseudo-maximization, which extends Robbins’ method of mixtures, sharpened these (with M. J. Klass and T. L. Lai) into refined, time-uniform inequalities underlying modern confidence sequences and sequential tests. A recollection joins the two parts: around 1992, J. L. Doob suggested I turn to self-normalized inequalities, advice that shaped much of what followed. The second part, Beyond Wald’s Equation and the Optional Sampling Theorem to Ville’s Equality, extends Doob’s optional sampling theorem for martingales, charting a route from Wald’s equation for randomly stopped sums to Ville’s equality — a coherent foundation for inference that remains valid under optional stopping.
Talks A1
Anytime Detection of Strategic Deviations in Multi-Agent Systems (Etienne Gauthier (Speaker), Francis Bach, Michael I. Jordan)
In many multi-agent systems, agents interact repeatedly and are expected to settle into stable, rational behavior over time. Yet in practice, behavior often drifts, and detecting such deviations in real time remains an open challenge. We introduce a sequential testing framework that monitors whether observed play is consistent with a benchmark of strategic behavior, without assuming a fixed sample size. Our approach builds on the e-value framework for safe anytime-valid inference: by “betting” against the benchmark, we construct a test supermartingale that accumulates evidence whenever observed payoffs systematically violate the expected conditions. For repeated normal-form games, we take equilibrium as the benchmark, yielding a statistically sound, interpretable measure of departure from equilibrium that can be monitored online; our framework unifies the treatment of Nash, correlated, and coarse correlated equilibria, offering finite-time guarantees and a detailed analysis of detection times. We also leverage Benjamini-Hochberg-type procedures to increase detection power in large games while rigorously controlling the false discovery rate. Finally, we extend our method to stochastic games, verifying online whether observed trajectories adhere to a specified target policy, such as a computed equilibrium, broadening the framework’s applicability to dynamic, state-dependent settings.
See also Posters A.
Predictive and confidence regions in causal inference (Vladimir Vovk (Speaker), Ruodu Wang)
We extend conformal e-prediction to design an algorithm for producing prediction sets in causal inference. Our algorithm works in the cases covered by the back-door criterion and those covered by the front-door criterion. However, its weakness is that it assumes that the training set is IID, or at least not too far from being IID, in some sense. In particular, our algorithm is not directly applicable in the setting of sequential decision making. To deal with this setting, we apply SAVI methods based on the law of the iterated algorithm to find explicit confidence intervals and confidence sequences. Once we have confidence intervals, we can use them to derive prediction sets; even if these prediction sets are less efficient for an IID training set, they are applicable in highly non-IID settings such as that of sequential decision making.
See also Posters A.
Brian Lee
abstract tba
See also Posters A.
Talks A2
Sequentially Valid Forecast Comparisons (Timo Dimitriadis, Jan-Lukas Wermuth (Speaker), Johanna Ziegel)
Time series forecasting and its evaluation are inherently sequential. However, their classical comparison via the Diebold–Mariano test is valid only for fixed evaluation horizons and does not allow for continuous monitoring. We develop sequentially valid inference on forecast performance for general, potentially unbounded loss differentials in multi-step-ahead forecasting based on asymptotic confidence sequences. These confidence sequences provide asymptotic coverage guarantees uniformly over time, thereby enabling online monitoring and tracking of loss differentials while permitting data peeking. We document the favorable finite-sample properties of our procedure in simulations and demonstrate their practical relevance in an application to volatility forecasting.
See also Posters A.
Optimal prediction with E-values (Nick Koning (Speaker), Sam van Meer)
Prediction sets offer a binary inclusion / exclusion for each element at the same fixed confidence level. We generalize this to fuzzy prediction sets, which exclude elements at their own data-driven confidence level. Our key insight is that a fuzzy prediction set is equivalent to a single E-value, capturing precisely what E-values bring to prediction. Fuzzy prediction sets inherit the merging properties of their E-value, and offer richer guarantees to decision makers. We show in what sense optimal E-values give rise to optimal (fuzzy) prediction sets. We apply our results to conformal prediction, deriving optimal conformal prediction sets, and characterizing in what sense classical conformal prediction is optimal.
See also Posters A.
Talks A3
Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences (Liviu Aolaritei (Speaker), Michael I. Jordan)
We study a basic but unresolved question in stochastic optimization: when should stochastic gradient descent (SGD) be stopped based only on its observed trajectory? We develop anytime-valid confidence sequences for stochastic gradient methods that remain valid under continuous monitoring and directly yield statistically valid stopping rules. In convex optimization, they certify weighted suboptimality under general stepsize schedules; in nonconvex optimization, they certify weighted first-order stationarity. The result is a unified framework for online stopping of SGD with provable complexity guarantees.
See also Posters B.
Bayes-Assisted Confidence Sequences for Bounded Means (Valentin Kilian, Stefano Cortinovis, François Caron (Speaker))
Confidence sequences based on test martingales provide time-uniform uncertainty quantification for the mean of bounded IID observations without parametric distributional assumptions. Their practical efficiency, however, depends strongly on the choice of martingale updates, and many existing constructions do not exploit prior information about plausible data-generating distributions or mean values. We propose a Bayes-assisted framework that uses a Bayesian working predictive model to adaptively construct confidence sequences. For each candidate mean and time point, the predictive distribution selects, among valid one-step martingale factors, the update maximising predictive expected log-growth; validity is therefore preserved even when the prior or working model is misspecified. We prove that if the predictive distribution is Wasserstein-consistent, the resulting procedure is asymptotically log-optimal, matching the per-sample log-growth of an oracle procedure with access to the true distribution. We instantiate the framework using robust predictives based on Dirichlet-process mixtures and Bayesian exponentially tilted empirical likelihood. Experiments on synthetic data, sequential best-arm identification for LLM evaluation, and prediction-powered inference show that informative priors can substantially reduce confidence-sequence width and sampling effort while retaining anytime-valid coverage.
See also Posters B.
Bet on Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters (Ivane Antonov*, Sohom Mukherjee* (Speaker), Richard Pibernik, Yo Joong Choe)
Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor’s information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained, “feature-aware” evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.
See also Posters B.
Posters A
Best Arm Identification for Bandits with Shifting Means (Lukas Zierahn, Wouter M. Koolen, Christina Katsimerou, Shubhada Agrawal and Dirk van der Hoeven)
We consider the problem of bandits with shifting means in the fixed confidence setting. We propose a parameterized family of e-values leveraging importance weighting. We then optimize their design for stopping early, and obtain an efficient sampling rule.
Aymeric Capitaine, Antoine Scheid, Etienne Boursier, Alain Durmus, Michael I. Jordan
abstract tba
Interactive Edge Orientation Identification: An Optimization Approach (Fabian Damken, Wouter M. Koolen, Rianne de Heide)
Prior work in pure exploration for multi-armed bandits has largely focused on best-arm identification and ranking, both in the fixed-confidence and fixed-budget setting. In many real-world applications, however, one is often interested in more than just the best option, yet a full ranking of options is typically unnecessary. Instead, questions such as “is A better than B? is B better than C?” are prevailing, but this problem of remains largely unstudied. We reframe this problem as one of learning the orientation of edges in a graph and propose a novel algorithm for tackling these problems in the fixed-confidence setting based on directly optimizing for the minimum number of samples.
E-Scores for (In)Correctness Assessment of Generative Model Outputs (Guneet S. Dhillon, Javier González, Teodora Pandeva, Alicia Curth)
While generative models, especially large language models (LLMs), are ubiquitous in today’s world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of LLM responses where the probability of including an incorrect response, or error, is capped at a user-defined tolerance level. However, since these methods are based on p-values, they are susceptible to p-hacking, i.e., choosing the tolerance level post-hoc can invalidate the guarantees. We therefore leverage e-values to complement generative model outputs with e-scores as measures of incorrectness. In addition to achieving the guarantees as before, e-scores further provide users with the flexibility of choosing data-dependent tolerance levels while upper bounding size distortion, a post-hoc notion of error. We experimentally demonstrate their efficacy in assessing LLM outputs under different forms of correctness: mathematical factuality and property constraints satisfaction.
Multi-Armed Sequential Hypothesis Testing by Betting (Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan)
We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms. We consider the composite global null hypothesis \(\mathcal{P}\) that all arms are null in a certain sense (e.g. all dosages of a treatment are ineffective) and we are interested in rejecting \(\mathcal{P}\) in favor of a composite alternative \(\mathcal{Q}\) where at least one arm is non-null (e.g. there exists an effective treatment dosage). We posit an optimality desideratum that we describe informally as follows: even if several arms are non-null, we seek \(e\)-processes and sequential tests whose performance are as strong as the ones that have oracle knowledge about which arm generates the most evidence against \(\mathcal{P}\). Formally, we generalize notions of log-optimality and expected rejection time optimality to more than one arm, obtaining matching lower and upper bounds for both. A key technical device in this optimality analysis is a modified upper-confidence-bound-like algorithm for unobservable but sufficiently “estimable” rewards. In the design of this algorithm, we derive nonasymptotic concentration inequalities for optimal wealth growth rates in the sense of Kelly (1956). These may be of independent interest.
Parameter-Free and Group Conditional Online Conformal Prediction (Beepul Bharti, Ambar Pal, Jacopo Teneggi, Jeremias Sulam)
Uncertainty quantification (UQ) is critical for the deployment of machine learning predictors in real-world scenarios where the data distribution may shift over time (i.e., data may not be exchangeable). Current online conformal prediction (OCP) methods address this issue at the expense of either (i) group-wise error control or (ii) learning-rate independent implementation. Group-conditional coverage is essential for fairness across different collections of data points and for providing finer UQ guarantees. Parameter-free optimization is crucial for robustness to adversarial and unknown data shifts. We propose a parameter-free algorithm for group-conditional OCP and demonstrate that it achieves the best group-conditional coverage guarantees. We evaluate our algorithm on synthetic and real-world data, demonstrating that our method not only improves the reliability of existing parameter-free OCP methods but also provides prediction intervals that are comparable in size to well-tuned group-conditional approaches. By unifying group-conditional coverage with parameter-free online algorithms, our work lays a foundation for fair and robust uncertainty quantification in shifting environments.
Eventually LIL Regret: Almost sure ln ln (T) Regret for a Sub-Gaussian Mixture on Unbounded Data (Shubhada Agrawal, Aaditya Ramdas)
abstract tba
Anytime Detection of Strategic Deviations in Multi-Agent Systems (Etienne Gauthier, Francis Bach, Michael I. Jordan)
See Talks A1.
Predictive and confidence regions in causal inference (Vladimir Vovk, Ruodu Wang)
See Talks A1.
Jessica Dai, Nika Haghtalab, Brian W. Lee
See Talks A1.
Sequentially Valid Forecast Comparisons (Timo Dimitriadis, Jan-Lukas Wermuth (Speaker), Johanna Ziegel)
See Talks A1.
Optimal prediction with E-values (Nick Koning, Sam van Meer)
See Talks A2.
Talks B1
Sequential Testing of Conditional Means: From Worst-Case to Instance-Dependent Optimality (Noah Liniger (Speaker), Antoine Scheid, Ian Waudby-Smith, Alain Durmus and Michael I. Jordan)
We develop sequential tests for conditional means of bounded random variables. These tests apply to several hypothesis testing problems, including testing the conditional coverage and calibration of “black-box” machine learning predictors, two-sample testing, and the evaluation of conditional treatment effects. We first identify the class of all admissible e-processes for conditional mean testing, and subsequently restrict our attention to this class. We then analyze the power of these e-processes and their downstream sequential tests through their asymptotic growth rates. For the maximin notion of power known as Growth Rate Optimality in the Worst Case (GROW), we derive optimal e-processes for one-sided and fixed-mean alternatives. For two-sided alternatives, we construct asymptotically GROW e-processes using mixture methods and sublinear-regret algorithms. We then study instance-dependent optimality guarantees and prove a hardness result showing that instance-dependent growth rate optimality is impossible over the full alternative. This motivates two weaker, composite-alternative-specific notions of optimality called mean growth rate optimality and variance growth rate optimality. These notions lie between worst-case and instance-dependent optimality by being worst case over progressively smaller subsets of the alternative containing the instance. We show that both notions are achievable over the full alternative and provide e-processes that attain them.
See also Posters B.
False discovery rates of refreshing monitoring procedures (Qiuqi Wang (Speaker), Ruodu Wang, Zhenyuan Zhang)
Practical monitoring procedures require anytime validity. In this paper, we examine the rationality of monitoring procedures that restart whenever rejections occur. Specifically, we study an online false discovery rate (FDR) control problem of refreshing monitoring procedures based on test (super)martingales and e-processes. We obtain explicit FDR bounds with fixed rejection thresholds under the independence and general settings. Moreover, we improve refreshing monitoring procedures by dropping less informative data points with an FDR guarantee. Numerical studies will be performed to illustrate the theoretical findings.
See also Posters B.
A Uniform Improvement of the Benjamini-Hochberg Procedure via e-Closure (Jelle Goeman)
This paper presents closed BH, a uniform improvement of the False Discovery Rate controlling method of Benjamini and Hochberg (BH). Closed BH is valid under the same assumption of Positive Regression Dependency on a Subset (PRDS) as BH, but also under an alternative and weaker minimal sufficient condition. As a uniform improvement, closed BH never rejects fewer hypotheses than BH, but it may reject quite a few more. An increase in power is observed especially when the number of false null hypotheses is large. The novel method is constructed using the e-Closure principle, a recently derived general principle for multiple testing. The method is implemented in the eClosure package in R.
See also Posters B.
A complete characterization of testable hypotheses (Johannes Ruf (Speaker), M. Larsson and A. Ramdas)
We revisit a fundamental question in hypothesis testing: given two sets of probability measures \(\mathcal{P}\) and \(\mathcal{Q}\), when does a nontrivial (i.e. strictly unbiased) test for \(\mathcal{P}\) against \(\mathcal{Q}\) exist? Le Cam showed that, when \(\mathcal{P}\) and \(\mathcal{Q}\) have a common dominating measure, a test that has power exceeding its level by more than \( \varepsilon \) exists if and only if the convex hulls of \(\mathcal{P}\) and \(\mathcal{Q}\) are separated in total-variation distance by more than \(\varepsilon\). The requirement of a dominating measure is frequently violated in nonparametric statistics. In a passing remark, Le Cam described an approach to address more general scenarios, but he stopped short of stating a formal theorem. This work completes Le Cam’s program, by presenting a matching necessary and sufficient condition for testability: for the aforementioned theorem to hold without assumptions, one must take the closures of the convex hulls of \(\mathcal{P}\) and \(\mathcal{Q}\) in the space of bounded finitely additive measures. We provide simple elucidating examples, and elaborate on various subtle measure theoretic and topological points regarding compactness and achievability.
See also Posters B.
Talks B2
Sample-efficient multiple testing with adaptive data collection (Zhimei Ren)
This talk concerns adaptive experimental design for multiple testing with e-values, where experimenters adaptively query hypotheses and form rejection sets based on sequentially updated e-values. We propose a general framework for sampling and constructing test statistics centered around e-values, which ensures anytime-valid FDR control under arbitrary dependence among the data collected across different hypotheses and discovers true effects in a sample-efficient manner. We establish general sample-efficiency guarantees of the proposed method in terms of the growth rate of the underlying e-processes, and instantiate the theory in a range of testing problems. We demonstrate the effectiveness of our approach through simulations and real-data experiments, showing that it ensures error control while maintaining a low sampling budget in online multiple testing scenarios.
See also Posters B.
Multiple Testing with Anytime-Valid Evidence (Yury Tavyrikov (speaker), Jelle Goeman and Rianne de Heide)
Sequential experiments generate evidence over time, and their strongest evidence may occur at different stopping times. For a single hypothesis, the running maximum of an e-process is anytime-valid through Ville’s inequality. Yet these maxima are not themselves e-values: they can have infinite expectation, making their direct use in multiple-testing procedures problematic. We develop sharp, dependence-robust tail bounds for aggregating running maxima of arbitrarily dependent e-processes. Our main result shows that the sum of \(K\) maxima has a \(1/\lambda\) tail bound with only a harmonic, logarithmic-in-\(K\) penalty. This penalty is unavoidable: an explicit dependent construction achieves the matching harmonic rate. The proof gathers asynchronous peaks into a single collection of stopped e-process values, where a rank-weighted argument reveals the harmonic structure. The approach extends beyond sums to general monotone aggregators, including power transforms that can reduce or eliminate the growing harmonic penalty. We illustrate two applications: anytime-valid local tests for closed testing and FDR control, and a global test for detecting coordinated signal in an unknown subset without paying an exponential price for scanning all groups.
See also Posters B.
Talks B3
Uniform log-optimality and conditional likelihood ratios (Dante de Roos)
For exponential families, Hao et al. (2024) and Hao & Grünwald (2026) have proposed to use conditional likelihood ratios whenever the log-optimal (GRO) e-variable is hard to compute. Here, the conditioning variable is a nontrivial statistic that is sufficient for both the null and alternative hypotheses, and the resulting e-variable is approximately log-optimal. Inspired by their results, we study conditional likelihood ratios from a completely abstract perspective, and identify the cases where conditional likelihood ratios are log-optimal. Most notably, the null hypothesis can always, in some sense, be canonically extended, and the conditional likelihood ratio is *uniformly* log-optimal for testing this extended null against the alternative. This notion of optimality is stronger than (RE)GROW and, to the best of our knowledge, has been mostly overlooked in the literature. Many previously studied e-variables fit nicely into our framework, which demonstrates the conceptual strength of our theory.
This talk does not have a poster.
Preferences over E-processes (Sam van Meer (Speaker), Nick Koning)
We formulate preferences over e-processes. Imposing natural axioms on how different degrees of evidence are valued over time, we derive a family of utilities indexed by a single parameter that can be interpreted as a discount factor. The resulting optimal e-process is proportional to a power of the likelihood ratio process, and optimal uniformly over stopping rules. The log-optimal e-process, often presented as a universal choice, is obtained as a special case that corresponds to linear utility in time.
See also Posters C.
Beyond First-order Asymptotics in Sequential Mean Testing (Vikas Deep, Shubhada Agrawal (Speaker))
We revisit the problem of sequentially testing the mean of bounded distributions in a level-\(\alpha\) power-one framework. We study a KL-inf-based sequential test that is known to attain the information-theoretic lower bound on the expected stopping time with exact constants as \(\alpha \to 0\). Going beyond first-order asymptotics, we establish a central limit theorem (CLT) for the stopping time of this test. Our analysis proceeds in two steps. First, we prove a novel CLT for the KL-inf statistic itself, characterizing its fluctuations around its deterministic limit. We then leverage this result to show that the stopping time, centered appropriately and scaled by \(\sqrt{\mathrm{log}(1/\alpha)}\), converges in distribution to a Gaussian limit with an explicit variance. This yields a second-order characterization of an asymptotically optimal sequential test for bounded distributions. Finally, we present numerical experiments that corroborate our theoretical findings.
This talk is based on https://arxiv.org/abs/2606.04520, a joint work with Vikas Deep (NUS, Singapore).
See also Posters C.
Posters B
Permutation-Based FDR Control via the e-Closure Principle (Rovanos Tsafack Nzanguim, Aurele Mingam, Jelle Goeman, Rianne de Heide)
While permutation methods are widely used for constructing valid p-values under dependence, there is currently no general approach for controlling the false discovery rate within the permutation framework. We address this gap by combining permutation based p-values with the e-closure principle to obtain a procedure that controls the false discovery rate at a prescribed level \(\alpha\). We consider testing problems for which the data is permutation-invariant under the joint null hypothesis and we compute the raw p-values for each permuted data set. These marginal p-values are then merged using Simes’ merging function and the corresponding global p-values are computed. The global p-values are thus transformed into global e-values using the SU calibrator and incorporated into the e-closure procedure.
Bringing Flexibility to the Benjamini Hochberg Procedure (Aurele Mingam, Rovanos Tsafack Nzanguim, Rianne de Heide, Jelle Goeman)
In multiple testing, post-hoc selection of the rejection set is possible while retaining rigorous guarantees under family-wise error rate (FWER) or false discovery proportion (FDP) tail-probability control. Within this framework, the error criterion itself can also be chosen after observing the data. Recently, Xu et al. (2026) extended these flexibility guarantees to false discovery rate (FDR) control through the e-closure principle. While this approach yielded substantial improvements for several FDR-controlling procedures, including BY, SU and e-BH, it provided little additional flexibility for the famous Benjamini–Hochberg (BH) procedure. We introduce a general framework to improve a broad class of FDR-controlling procedures. Our approach exploits the conservativeness of procedures whose FDR is bounded by \(\alpha \lvert N \rvert / m\), leaving an expectation gap whenever not all hypotheses are null. We show how this unused budget can be converted into additional post-hoc flexibility by augmenting the associated e-collection with carefully designed all-or-nothing e-variables. This construction preserves the original BH rejection set while enabling flexibility.
Global Sequential Testing for Multi-Stream Auditing (Beepul Bharti, Ambar Pal, Jeremias Sulam)
Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed. This auditing task can be modeled as a sequential hypothesis testing problem with \(k\) data streams and a global null hypothesis that asserts the system operates as intended across all \(k\) streams. Under the alternative, the standard global sequential test, which uses a Bonferroni correction, has an expected stopping time of \(\mathcal{O}\left(\ln \frac{k}{\alpha}\right)\) for large \(k\) and significance level \(\alpha\). In this work, we demonstrate that efficient sequential tests, relying on merging martingales via averaging and product rules, provide improved stopping times, and thus more powerful tests against the null. Using these results, we show that a balanced test can match the Bonferroni rate of \(\mathcal{O}\left(\ln \frac{k}{\alpha}\right)\) in the sparse regime (just a few non-null streams) while achieving \(\mathcal{O}\left(\frac{1}{k}\ln \frac{1}{\alpha}\right)\) under dense alternatives (many non-null steams).
Optimal e-variables for bounded mean testing under composite alternatives (Eugenio Clerico, Sebastian Arnold)
We derive the unique e-values with optimal (relative) growth rate in the worst-case for testing the mean of a bounded random variable, hereby contributing with the first application beyond the assumption of mutually absolutely continuous hypotheses of the (RE)GROW quality criteria for e-variables originally proposed by Grünwald et al. (2024). For both criteria, we characterise explicitly the alternatives for which it is most difficult to test against. These worst-case distributions admit a meaningful interpretation and reduce the original highly nonparametric problem to an associated Bernoulli model. We show that REGROW yields meaningful optimal e-variables, whereas GROW becomes trivial for one- and two-sided mean testing without a minimal effect size assumption. Our analysis relies on a result from convex analysis that represents worst-case expectations via convex minorants.
Linear regression testing with Model-X (Sebastian Arias, Peter Grünwald)
We evaluate the performance of the Model-X e-variable (Grünwald et al., 2023) in the context of linear regression testing. The Model-X e-variable was originally developed for conditional independence testing, and it relies on partial knowledge of the covariate distribution (e.g., knowledge of the allocation procedure assigning participants to experimental groups). We compare its e-power against that of the growth-rate optimal e-variable for the linear regression null, proposed by Grünwald et al. (2025), which instead leverages linearity and distributional assumptions about the target variable only.
In scenarios where both e-variables are applicable – including testing treatment effects via t-tests, ANOVA, and ANCOVA analyses – we show that the Model-X e-variable can attain greater e-power on finite samples, while it is asymptotically always weaker. Additionally, we identify a “generalized ANCOVA” setting where the asymptotic e-power gap is of order \(O( \Vert \delta \Vert^6 )\) as the effect size vector \(\delta\) tends to zero, suggesting that the difference may be practically negligible for typical effect sizes.
D. Hop, N. Koning, S. van der Meer
abstract tba
Sequential Testing in \(2 \times 2\) Contingency Tables with Conditional e-Variables (Yonqqi Wang, Sebastian Arnold, Francesca Giuffrida, Peter Grünwald, Thorsten Dickhaus)
Much remains unknown about the design of optimal e-processes and e-variables for the simple yet fundamental problem of testing association in \(2 \times 2\) contingency tables. The e-process introduced by Turner et al. (2024) is growth-rate optimal (GRO) for independent Bernoulli streams with fixed means, but it is powerless when applied to a single large table. Conditional e-variables, by contrast, perform significantly better in the single-table setting, outperforming both the UI e-variable and an e-variable obtained by calibrating Fisher’s exact test. When tables arrive sequentially, Turner’s approach is generally asymptotically optimal under fixed-mean conditions, but conditional e-variables demonstrate robustness to varying means while the odds ratio remains fixed, as measured by growth rates and expected stopping times.
Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences (Liviu Aolaritei, Michael I. Jordan)
See Talks A3.
Bayes-Assisted Confidence Sequences for Bounded Means (Valentin Kilian, Stefano Cortinovis, François Caron)
See Talks A3.
Features: Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters (Ivane Antonov*, Sohom Mukherjee*, Richard Pibernik, Yo Joong Choe)
See Talks A3.
A complete characterization of testable hypotheses (Johannes Ruf, M. Larsson and A. Ramdas)
See Talks B1.
False discovery rates of refreshing monitoring procedures (Qiuqi Wang, Ruodu Wang, Zhenyuan Zhang)
See Talks B1.
Sequential Testing of Conditional Means: From Worst-Case to Instance-Dependent Optimality (Noah Liniger, Antoine Scheid, Ian Waudby-Smith, Alain Durmus and Michael I. Jordan)
See Talks B1.
A Uniform Improvement of the Benjamini-Hochberg Procedure via e-Closure (Jelle Goeman)
See Talks B2.
Multiple Testing with Anytime-Valid Evidence (Yury Tavyrikov, Jelle Goeman and Rianne de Heide)
See Talks B1.
Sample-efficient multiple testing with adaptive data collection (Zhimei Ren)
See Talks B2.
Talks C1
Analyzing the relationship between universal and classical likelihood ratio inference: Almost sure-relations and asymptotic rates (Lorenz Matz (Speaker), Hannes Leeb)
Universal Inference is a general method for generating confidence sets and tests which is based on a specific e-value construction. Unlike procedures built on asymptotic results or most resampling methods, its confidence sets/tests have finite-sample guarantees and require virtually no regularity conditions on the model or the set of null distributions. However, its relation to classical inference approaches in terms of power and confidence set diameters is still poorly understood. To shed some light on this, we study the relationship between the split LR and the classical LR confidence set in general settings. For one- and two-dimensional models, we show that the classical LR set is almost surely contained in the split LR set when the maximum likelihood estimator is used for the latter. Furthermore, we find conditions under which the diameters of both confidence sets shrink at the same rate by proving some general results about these rates for ‘M-type’ sets.
See also Posters C.
Time-sensitive anytime-valid testing (Eugenio Clerico (Speaker), Iskander Azangulov, and Patrick Rebeschini)
Abstract: Anytime-valid tests allow evidence to be checked during data collection: one can either continue testing or stop and reject the null while still controlling type-I error. Yet, in many applications rejection is useful only if it comes soon enough. We introduce a time-sensitive testing-by-betting framework that favours early rejection by assigning rewards to rejection times and maximising their expected value under a given alternative. This encompasses hard deadlines and softer time preferences. The resulting optimal control problem admits a Bellman representation in terms only of time and evidence against the null, rather than the full history. For hard deadlines, the simple-vs-simple case reduces to a finite-horizon Neyman-Pearson problem, with a corresponding optimal e-process. Furthermore, we show that exponentially decaying rewards admit a stationary approximation, yielding the exponential-decay-optimal (EDO) criterion: a finite-time-scale counterpart to the classical growth-rate-optimal (GRO) viewpoint in anytime-valid statistics, with the GRO criterion recovered in the large-time-scale limit.
See also Posters D.
Posters C
When generalized likelihood ratios are – and are not – e-processes (Stephan Bongers, Peter Grünwald)
Generalized likelihood ratios (GLRs) are among the most widely used tools for composite hypothesis testing. By replacing unknown nuisance parameters with their maximum likelihood estimates, they provide a natural and powerful way to accumulate evidence against the null hypothesis. Yet, despite their widespread use, GLRs are not automatically anytime-valid and need not form e-processes. We show that the behavior of the GLR in exponential families depends on the covariance ordering of the null and alternative. In one regime, the GLR is automatically an e-process; in the other, it generally is not. We introduce a calibrated GLR that restores e-process validity and often achieves substantially higher e-power than the sequential RIP and UI e-processes.
E-values in the Lean theorem prover (Gaëtan Serré, Rémy Degenne)
We build an e-value library for the Lean theorem prover. Lean is a proof assistant, a programming language in which one can write mathematical theorems and proofs. We build a new library on top of Mathlib, Lean’s mathematical library, in which we implement definitions for e-variables, their utility, the numeraire/GRO, duality with the reverse information projection and Kullback-Leibler divergences, and other tools needed for working with e-values. To illustrate the use of our library, we prove information theoretic results in Lean using the e-variables definitions: data-processing inequality and tensorization equalities for e-variables, and lower bounds for hypothesis testing.
M-estimation with e-statistics (Hongjian Wang)
We present a theory of point estimation with e-statistics (e-values and e-processes) by introducing the “ME-estimator”: the parameter that minimizes the corresponding e-statistic, or the evidence against it. Our approach is based on the intuitive idea of e-statistics as a measure of evidence and betting pay-off, and naturally generalizes the classical method of maximum likelihood estimation. First, we establish the consistency as well as the almost sure convergence rate for ME-estimators relating to the high-probability bounds on the size of the confidence set derived from thresholding the e-statistics, an approach that sets ME-estimators apart from traditional M-estimators. Second, we conduct classical M-estimator-style analysis on the consistency and asymptotic normality of ME-estimators in the bounded mean estimation setting, discussing the notion of efficiency (or lack thereof) from various choices of betting strategy. Our work brings e-statistics, a fundamental tool for inference and uncertainty quantification, to the space of estimation.
Asymptotic REGROW e-variables for exponential families (Dante de Roos, Peter Grünwald, Sebastian Arnold)
Optimal e-variables for exponential families have recently been studied by Hao and Grünwald (Bernoulli, 2026). Extending their work, we study testing problems in which \(H_0\) and \(H_1\) are given as exponential families with a shared sufficient statistic. We extend the theory to the case where \(H_1\) is higher-dimensional than \(H_0\), covering the important setting of testing a parameter of interest in the presence of nuisance. We show that the loss in expected growth of a proposed conditional type of e-variable compared to the optimal “oracle” GRO is of order \((d/2) \mathrm{log}(n) + O(1)\) uniformly over the parameters with d denoting difference in dimensionality between null and alternative. This asymptotically optimal (“relative GROW”) rate is reminiscent of the classical BIC derivation from which we adapt several proof ideas; our setting is technically more challenging though through the presence of conditional likelihoods.
Extending the exponential martingale to testing equality of \(K\)-samples (Alexander Ly, Udo Boehm, Wouter Koolen, Peter Grünwald)
This work addresses anytime-valid inference for the \(K\)-sample problem in the unbalanced setting where the sample sizes differ across groups. The null hypothesis assumes that all \(K\) groups have the same, but unknown, population mean, as is commonly the case in applications such as multi-arm clinical trials and A/B/(n) testing. We developed a test martingale that generalises the standard \(K = 1\) exponential martingale, ensuring adaptability to varying sample configurations. Guided by (i) labelling invariance, which guarantees the test’s robustness to the choice of reference group, and (ii) a minimal clinically relevant mean difference, we derive the \(Q\)-optimal Gaussian mixed \(K\)-sample exponential martingale. Test inversion leads to simultaneous anytime-valid confidence ellipsoids for the mean differences.
Sequential Conditional Independence Testing with Machine Learning Models (Angel Reyero Lobo, Sebastian Arias, Michele Meziu)
Conditional Independence Testing is a ubiquitous problem in scientific discovery. The widely employed Model-X assumption shifts the modelling burden from the dependency of the output given the inputs, to the dependencies within inputs. Growth-rate optimal e-variables have been studied in this setting (Grünwald et al., 2023), but it remains unclear how to optimally include machine learning models in these tests. In this work, we exploit the performance drop of a model when a given feature is removed to construct an exponential e-variable that approximates the GRO, and compare its e-power to that of a class of antisymmetric e-variables (Shaer et. al., 2023). We show theoretically and experimentally that antisymmetric e-variables can outperform the exponential in low signal regimes, while the exponential dominates when the signal is large and the loss is well specified. Finally, we provide an actionable algorithm to approximate the growth-rate optimal among the antisymmetric e-variables via kernel density estimation of the loss distribution.
Preferences over E-processes (Sam van Meer, Nick Koning)
See Talks B3.
Beyond First-order Asymptotics in Sequential Mean Testing (Vikas Deep, Shubhada Agrawal)
See Talks B3.
Analyzing the relationship between universal and classical likelihood ratio inference: Almost sure-relations and asymptotic rates (Lorenz Matz, Hannes Leeb)
See Talks C1.
Time-sensitive anytime-valid testing (Eugenio Clerico, Iskander Azangulov, and Patrick Rebeschini)
See Talks C1.
Talks D1
Downstream Post-Hoc Hypothesis Testing or: how E-Values generalize De Finetti’s Probability (Peter Grünwald (Speaker), Ben Chugg, Aaditya Ramdas)
Chugg et al. (IJAR, 2026) formalized post-hoc hypothesis testing as a game between Adversary, setting the loss function, and Statistician, deciding whether to reject \(H_0\). They showed that, under some conditions on Adversary, “admissible” decision rules must be based on e-variables; and under further conditions, these must have a certain form, e.g. in some settings they must be increasing in the LR. However, both the conditions and their analysis were exceedingly complicated. Here we show that by changing the game slightly, with a different, arguably more reasonable, action space for Statistician, we can obtain a much crisper result, which has the classical Neyman-Pearson lemma as a simple special case. The resulting re-interpretation of e-values shows that they are really direct generalizations of De Finetti’s conceptualization of “upper probability”. This allows to give real meaning to crazy statements like “the probability that the empirical average of \(\tau\) outcomes is outside the set \(A\) is bounded by \(\mathrm{exp}(-1.5 \tau)\)” where \(\tau\) is a random (!) stopping time.
See also Posters D.
e-values as evidence (Ben Chugg, Peter Grünwald, Aaditya Ramdas)
A recurring debate in the philosophy of statistics concerns what, exactly, should count as a measure of evidence for or against a given hypothesis. P-values, likelihood ratios, and Bayes factors all have their defenders. In this paper we add two additional candidates to this list: the e-value and its sequential analogue, the e-process. E-values enjoy several desirable properties as measures of evidence: they combine naturally across studies, handle composite hypotheses, provide long-run error rates, and admit a useful interpretation as the wealth accrued by a bettor in a game against the null distribution. E-processes additionally handle optional stopping and optional continuation. This work examines the extent to which e-values and e-processes satisfy the evidential desiderata of different statistical traditions, concluding that they combine attractive features of p-values, likelihood ratios, and Bayes factors, and merit serious consideration as interpretable and intuitive measures of statistical evidence.
See also Posters D.
Optimal Posterior E-values with Non-Convex Parameter Sets and Applications to Voting Systems (Timothée Mathieu (Speaker), Adrienne Tuynman)
In this talk, I present our work on using e-values for sequential testing to identify a preference-feedback winner (e.g. Condorcet, Borda, …). The practical motivation of our work is political polls, in which one wants to identify (or test for) the winner while collecting as few samples as possible. This setting led us to consider posterior optimal e-variables, defined as a generalisation of the GRO e-variable to composite \(H_1\), using a Bayesian prior on the alternative. We provide a Frank Wolfe algorithm that computes the Reverse Information Projection for non-convex parameter sets, which helps us to compute this e-variable. Finally, we give theoretical guarantees and practical illustrations of the power and sample complexity of the associated testing methods.
See also Posters D.
SSBBs: A method for determining the sample size when E-Backtesting the Expected Shortfall (Dennis Oestmann (Speaker), Thorsten Dickhaus)
This talk presents an approach for determining sample sizes required to detect underestimations of the Expected Shortfall with a prescribed power when applying the SAVI-based backtesting procedure introduced by Wang et al. in their recent paper “E-Backtesting”.
We consider scenarios in which the Value-at-Risk at level \(p\) is always estimated correctly, while the difference between the true Expected Shortfall and Value-at-Risk is underestimated by a given factor \(r\). We show that exploiting the structure of the backtest e-statistic \(\mathrm{max} \{ 0, x - z \} / (1 - p) / (r - z)\), proposed for backtesting the Expected Shortfall at level \(p\), enables the derivation of approximate lower bounds for the required sample sizes by considering a simple sequence of i.i.d. Bernoulli-distributed random variables.
We also discuss potential limitations of this approximation and compare the resulting sample size requirements with those obtained in practical applications using Monte Carlo simulations.
See also Posters D.
Talks D2
Design-Based Anytime-Valid Inference for Randomized Experiments with Delayed Outcomes and Staggered Entry (Michael Lindon (Speaker), Nathan Kallus)
Delayed outcomes are ubiquitous in online experimentation: treatment can affect whether an outcome occurs, when it occurs, and its realized value. To accommodate staggered entry while remaining robust to environmental nonstationarity and unit-level heterogeneity, we adopt a design-based perspective and target the sample cumulative reward in each arm as a function of calendar time. Our confidence sequences allow practitioners to continuously monitor the counterfactual incremental reward, such as revenue, that would have been realized by calendar time t had all entered units been assigned to treatment rather than control. The main technical challenge is the choice of design-based filtration, complicated by the presence of asynchronous potential outcome times. We show that the IPW treatment-effect estimation error is not a martingale with respect to any filtration, while each arm-specific IPW estimation error is a martingale with respect to a carefully chosen arm-specific event-time filtration. We therefore construct a confidence sequence for the treatment effect by combining two arm-level confidence sequences with a union bound, and further demonstrate that this can outperform the traditional design-based variance upper bound. Finally, we characterize the class of augmentations for which the per-arm AIPW estimation error remains a martingale.
See also Posters D.
Adaptive clinical trials based on design-optimal e-values: Application to single-arm trials (Stef Baas (Speaker), Joost van Rosmalen, Judith ter Schure)
e-values are a promising method for adaptive clinical trials due to their anytime validity, facilitate repeated interim analyses, complex stopping rules, and valid inference under protocol deviations. The e-value literature focuses mostly on asymptotic optimality; however, sample sizes in clinical trials are often limited. To this end, we investigate e-value-based designs with finite-horizon optimality for single-arm multistage clinical trials with binary data (relevant in early-phase cancer trials). We use the betting interpretation of e-values to construct e-values that minimize the expected sample size, stopping for futility or efficacy with constraints on the minimum power. We construct these designs through constrained dynamic programming based on the currently observed e-value, the maximum sample size, and the pre-specified significance level. Using exact calculations, we show that, next to robustness, e-value-based designs can provide competitive operating characteristics to standard (non-)adaptive designs with and without futility stopping and outperform growth-rate-optimal e-values in finite samples. In addition, small e-values automatically indicate trial continuation is futile, e.g., an e-value of zero indicates the impossibility of an efficacy conclusion. Hence, e-value-based designs provide a viable alternative to the current state-of-the-art in single-arm binary trials, warranting extension to other adaptive clinical trial settings such as multi-arm multi-stage and response-adaptive designs.
See also Posters D.
Talks D3
Betting on Bets: Anytime-Valid Tests for Stochastic Dominance (Sebastian Arnold* (Speaker), YJ Choe* (Speaker), Marco Scarsini, Ilia Tsetlin)
How can we monitor, in real time, whether one uncertain prospect has any upside over another? To answer this question, we develop a novel family of sequential, anytime-valid tests for stochastic dominance (SD), a classical and popular notion for comparing entire distribution functions. The problem is distinct from the popular problem of testing for dominance in means, which would not capture distributional differences beyond the mean. We first derive powerful, nonparametric e-processes that quantify evidence against the null hypothesis that one prospect is dominated by another. For first-order SD, these e-processes are constructed as a mixture of asymptotically growth-rate optimal e-variables and yield a test of power one that retains validity under continuous monitoring. The approach further generalizes to sequential testing for SD beyond the first order, including any higher-order SD. Empirically, we demonstrate that the resulting sequential tests are competitive with classical, non-anytime-valid SD tests in terms of power, and include a real-world application to baseball analytics, examining a controversial phenomenon known as third-time-through-the-order penalty. Finally, we sketch the complementary and challenging problem of testing whether a prospect has a definite upside, and describe the conditions under which we can derive a nontrivial anytime-valid test.
See also Posters D.
Verifying Elections with Adaptively Weighted Test Supermartingales (Alexander Ek (Speaker), Michelle Blom, Philip B. Stark, Peter J. Stuckey, Damjan Vukcevic)
An increasingly important part of trustworthy elections is the risk-limiting audit (RLA): a statistical audit that efficiently confirms the reported outcome when it is correct or with high probability corrects the reported outcome if it is incorrect. An election audit is risk-limiting if the chance it fails to correct a wrong electoral outcome is at most a pre-specified limit.
We present recent advances in RLAs for instant-runoff voting (IRV), a ranked-choice system used in Australia, USA, and other countries. To test that the reported outcome is wrong, we need a composite hypothesis written as a union of \(O(k!)\) intersections of \(O(2^k)\) simple hypotheses when there are \(k\) candidates. Prior work addressed this by pre-selecting a subset of the simple hypotheses covering all unions. While this eliminates multiplicity, it necessitates reliable estimates of the votes cast. Instead, we use predictably weighted averages of adaptive test supermartingales, removing the need for pre-selection entirely. We call this the Adaptively Weighted Audits of Instant-Runoff Elections (AWAIRE) framework.
This work involves new statistical and computational methods in sequential anytime-valid inference. Naively testing all simple hypotheses is intractable when \(k > 6\). We take advantage of the tree structure of the tally algorithm to start by testing a few simple hypotheses, adaptively adding more as needed to confirm the results, using branch-and-bound to prune the search space. This makes it feasible to audit contests with more than 50 candidates. The approach may be useful in other situations where the null hypothesis can be expressed as a union of intersections of simple hypotheses.
See also Posters D.
On vector-valued self-normalized concentration inequalities (Diego Martinez Taboada (Speaker), Tomas Gonzalez, Aaditya Ramdas)
The study of self-normalized processes plays a crucial role in a wide range of applications, from sequential decision-making to econometrics. While the behavior of self-normalized concentration has been widely investigated for scalar-valued processes, vector-valued processes remain comparatively underexplored, especially outside of the sub-Gaussian framework. We provide concentration bounds for self-normalized processes with light tails beyond sub-Gaussianity (such as Bennett or Bernstein bounds), and we illustrate the relevance of our results in the context of online linear regression.
See also Posters D.
Posters D
The Winner’s Curse as an e-Process Compensator with Applications to Safe Anytime-Valid Testing for Dose-Ranging Trials (Victor K. de la Peña, Fangyuan Lin, Demissie Alemayehu, Victor H. de la Peña)
Phase II dose-ranging trials often report the largest observed dose–control effect while inspecting accumulating data repeatedly. This creates two coupled distortions: selection optimism from choosing the empirical winner, known as the winner’s curse, and Type-I error inflation from multiplicity across doses and interim looks. We develop an anytime-valid procedure for testing whether the best true dose effect exceeds a clinically meaningful margin. The mathematical starting point is a recent selection-premium identity for the running maximum: for dose–control scores, the expected gain from re-selecting the current leader becomes a predictable selection charge. Subtracting this charge gives a residual with nonpositive drift under the composite null; applying a one-sided mixture-exponential construction then yields an e-process and hence an anytime-valid global test. The resulting rule has a transparent ledger form: raw best effect minus selection charge minus monitoring margin, and a “go” decision is made only when the remaining evidence still exceeds the clinical margin. We give plug-in implementations for Gaussian and binary outcomes, prove finite-sample Type-I control and anytime lower confidence bounds for the best dose effect, and illustrate the method through worked examples and simulations.
Anytime-valid log-rank testing for randomised trials (Joren Brunekreef, Renee Menezes, Rianne de Heide)
Anytime-valid (AV) log-rank tests allow continuous monitoring of survival trials with Type-I control under optional stopping and continuation. However, their practical deployment is less explored. We characterise the fixed-\(\delta\) AV log-rank by simulations across representative trial designs, varying the design and true hazard ratios. Anytime-validity admits unrestricted stopping, so performance is measured via “power-over-time”: the rejection probability by each follow-up moment, extending classical fixed-\(n\) power. Under simple stopping rules, the AV log-rank’s power-over-time approaches the classical fixed-\(n\) power at the design horizon and can exceed it under continuation. We deploy the framework retrospectively on a real randomised trial with patient-level data, where a cohort-complete construction produces a calendar-date announcement directly comparable to the trial’s analysis cutoff.
Clinical trial problems looking for anytime-valid solutions (Judith ter Schure)
Clinical trials sometimes operate under tightly planned stopping rules, but often stop for other reasons. Data on safety outcomes, secondary outcomes, or data from other clinical trials might make a trial stop early. Or patient inclusion is so slow that completing the trial seems impossible, which is also often related to the results observed in undisclosed ways. My poster invites you to contact me if you would like to work on e-value methods that you can see in action on real data. Specifically, a soon to start clinical trial in emergency medicine will probably need a way to switch from O’Brien-Fleming alpha-spending to optional continuation. Contact me if you would like to be a collaborator that works on new concrete theory that I can directly put in practice. I have a little budget for a visit to Amsterdam, and the team includes excellent trialists (clinicians running the trial) with a lot of ambition for future flexibilization of the trials (platform trials) and a lot of interest in statistics. My poster lists a few more settings with unexpected stopping that fit the e-value theory perfectly, but still lack optimal of-the-shelf anytime-valid confidence intervals.
E-values for early decisions in clinical trials (Rianne de Heide, based in part on joint work with Nynke Luijten and Vincent van der Noort)
This poster presents two projects of early phase clinical trials where an e-value based design provides an improvement in expected sample size with respect to the trials that are currently conducted in many medical trials. Firstly, we study combined pilot-definitive or feasibility-definitive trials. Rather than treating the pilot or feasibility study as separate from the main analysis, an e-process can be monitored during early recruitment, while preserving Type I error control. A paired binary-outcome example illustrates a vast sample size improvement. Ongoing work studies power, sample size, estimation, and confidence intervals under many diffent widely used settings. Secondly, we compare an e-value based single-arm phase-II design with Simon’s optimal and minimax two-stage designs for binary response outcomes, such as used in the ongoing DRUP study (https://drupstudy.nl). The e-value based design uses two e-processes, one for efficacy and one for futility, and simulations indicate lower average stopping samples sizes than Simon’s designs in relevant clinical settings. Ongoing work studies optimal e-values for this setting and a real-world application in oncology.
Incentive aware AI regulations: A credal characterisation (Anurag Singh, Rajeev Verma, Julian Rodemann, Siu Lun Chau, Krikamol Muandet)
While high-stakes ML applications demand strict regulations, strategic ML providers often evade them to lower development costs. To address this challenge, we cast AI regulation as a mechanism design problem under uncertainty and introduce regulation mechanisms: a framework that maps empirical evidence from models to a license for some market share. The providers can select from a set of licenses, effectively forcing them to bet on their model’s ability to fulfil regulation. We aim at regulation mechanisms that achieve perfect market outcome, i.e. (a) drive non-compliant providers to self-exclude, and (b) ensure participation from compliant providers. We prove that a mechanism has perfect market outcome if and only if the set of non-compliant distributions forms a credal set, i.e., a closed, convex set of probability measures. This result connects mechanism design and imprecise probability by establishing a duality between regulation mechanisms and the set of non-compliant distributions. We also demonstrate these mechanisms in practice via experiments on regulating use of spurious features for prediction and fairness. Our framework provides new insights at the intersection of mechanism design and imprecise probability, offering a foundation for development of enforceable AI regulations.
Bayes-Assisted Confidence Sequences for Cumulative Distribution Functions (Valentin Kilian*, Shirley Xiaoqi Liu*, Judith Rousseau, Francesco Orabona, Patrick Rebeschini)
We study how Bayesian priors can tighten confidence bands for the CDF of a bounded random variable in sequential inference. Moving beyond pointwise mean estimation, we aim to estimate the entire CDF from sequential observations. By placing a prior on the unknown CDF, we construct mixture wealth processes under the betting formulation leveraging prior information to reduce regret and hence confidence band width, while preserving anytime-valid guarantees uniformly over time and across the entire support. Theoretical and empirical analyses demonstrate tighter intervals when the prior is well-aligned with the true distribution. This is work in progress, joint with Valentin Kilian, Judith Rousseau, Francesco Orabona, and Patrick Rebeschini.
e-values instead of p-values in clinical trials: What happens? (Yongxi Long, Alexander ly, Jelle Goeman, Nikos Ignatiadis, Peter Grünwald, Erik van Zwet)
Introduction: e-values have been proposed as a flexible alternative to p-values for anytime-valid inference. In the context of clinical trials, they enable continuous monitoring of accumulating evidence and optional continuation of recruitment beyond the originally planned sample size. Of course, this flexibility comes at the cost of reduced statistical power relative to fixed-sample testing. The goal of this paper is to quantify this trade-off empirically in the context of clinical trials. This is complicated by the fact that there is a wide class of valid e-processes with different alpha-spending behaviors.
Methods: We estimated the distribution of effect sizes from more than 20,000 randomized controlled trials from the Cochrane Database of Systematic Reviews (CDSR). We then generated synthetic trial trajectories up to and beyond their original sample sizes. At the original sample size, these trajectories hit the original effect estimate and its standard error. We can use this synthetic dataset to benchmark any given e-process against the fixed sample p-value. We focus on the probability of significance (PoS) and the time until rejection.
We evaluate several e-processes. First, we construct an empirical Bayes e-process that considers the estimated effect size distribution relative to the null. Second, we used point-prior e-processes that are informed by each trial’s minimum clinically important difference (MCID). Lastly, we used e-processes with method-of-moment (MoM) priors whose modes are informed by trial MCIDs.
Results: The MCID-point prior e-process performs most efficiently across the CDSR cohort: it requires planning for approximately 1.8 times the original sample size to match the PoS of a fixed-sample p-value, but the expected sample size is reduced to approximately 1.5 times due to optional stopping. The empirical Bayes e-process is highly conservative because it reserves its alpha-budget for detecting many small-effect trials in the CDSR cohort at large sample sizes. Smoothing the e-processes using MoM priors improves robustness to effect-size misspecification for individual trials, but pays a larger efficiency price on average across the CDSR cohort.
Conclusion: The great flexibility of e-values offers attractive advantages, particularly in enabling anytime-valid testing. Their practical implementation in clinical trials requires a balance between flexibility and power. In a setting characterized by many small effects and resource constraints, more aggressive e-process constructions yield better average efficiency.
Equivalence testing with data-dependent and post-hoc equivalence margins (Stan Koobs, Nick Koning)
We consider the problem of assessing whether an unknown quantity \(\mu \geq 0\) is negligible. This is classically treated as a testing problem, comparing the hypothesis that \(\mu\) is larger than some “margin” \(\Delta\) against the alternative that it is smaller: negligible. A longstanding problem is that this margin must be specified in advance, but many applications do not admit a single natural margin below which \(\mu\) is negligible. Our conceptual contribution is to reason backwards from a stylized application to argue that one should not treat this as a testing problem, unless a natural margin is available. Our methodological contribution is to propose reporting a data-dependent margin \(\widehat{\Delta}_\alpha\), that bounds \(\mu\) with probability \(1 - \alpha\). We generalize this to a curve of margins \(\alpha \mapsto \widehat{\Delta}_\alpha\), uniformly valid under the post-hoc selection of the margin. These rely on e-values, and our technical contribution is to derive such e-values for models that are strictly totally positive of order 3, nesting the classical z-test and t-test settings.
Dynamic evidence synthesis with e-values: Anytime-valid meta-analysis for the presence or absence of effects (Alexander Ly, Sebastian Arias, Sebastian Arnold, Udo Boehm, Stephan Bongers, Michele Meziu, Angel Reyero Lobo, Dante de Roos, Meike Steinhilber, Yongqi Wang, Peter Grünwald)
In large-scale reproducibility efforts in the social sciences, such as ManyLabs2, influential studies are replicated across laboratories around the world, an expensive and time-consuming process. Previous applications of E-processes to such a real-world meta-analysis were hindered by the absence of a built-in method to stop for futility, i.e., when there is compelling evidence that the original effect is not reproducible. We remedy this with a Wald-style sequential test, but with anytime-valid decision boundaries. This procedure races the E-process \(S\) and \(R\), where \(S\) refers to the ordinary E-process which will incorrectly yield e-values larger than \(1/\alpha\) with no more than alpha probability under the null. On the other hand, \(R\) is constructed such that \(1/R\) is a test martingale against the modified alternative hypothesis that the effect size parameter is at least as large as a minimal clinical relevant effect size. The first time \(S \geq 1/\alpha\) and \(R \leq \beta\) then defines the Wald-like test with the sought-after type I and type II error control over time. To quantify the gain in efficiency we apply the procedure to 26 meta-analyses from ManyLabs2, and show how off-the-shelf e-values lead to massive reductions in the required number of replication attempts and their sample sizes, while leading to more robust conclusions.
Deliberative Prediction (Rajeev Verma, Rabanus Derr, Christian A. Naesseth, Eric Nalisnick, Krikamol Muandet)
Consequential institutions now routinely use data-driven predictors to forecast relevant outcomes, such as whether an applicant will succeed in college. However, the purpose of forecasting is to enable better decisions and outcomes, rather than merely producing accurate predictions. Standard criteria for forecast quality, such as calibration and multi-calibration, are framed as intrinsic properties of the predictor, leaving no role for the decision-maker to deliberate on its outputs. We argue that forecast quality is relativistic: each criterion expresses a randomness statement relative to an evaluator defined by an information set, and the appropriate criterion follows once the evaluator is specified. Under institutional separation between the model designer and the decision-maker, the natural evaluator is the deploying institution, which often possesses outcome-relevant information unavailable to the model designer. We identify a data-side condition, under-determinism, in which the decision-maker’s information strictly refines the forecast and residual heterogeneity remains non-trivial. Under this condition, no precise forecast can be sound, whereas an imprecise forecast can satisfy an appropriate notion of imprecise randomness. Rather than suppressing this heterogeneity, it should be reflected in the forecast itself. We term this approach deliberative prediction: the forecaster reports what its information cannot resolve, leaving the residual, including normative considerations such as fairness, to the decision-maker. Thus, imprecision becomes a formal trace of what the forecaster leaves open for the decision-maker to deliberate.
Post-detection Inference for Sequential Changepoint Localization (Aytijhya Saha, Aaditya Ramdas)
We address a fundamental but largely unexplored challenge in sequential changepoint analysis: conducting inference following a detected change. We develop a very general framework to construct confidence sets for the unknown changepoint using only the data observed up to a data-dependent stopping time at which an arbitrary sequential detection algorithm declares a change. Our framework applies broadly, encompassing both classical settings with specified pre- and post-change distribution classes and fully distribution-free settings. We provide theoretical guarantees on the width of our confidence intervals. Extensive simulations demonstrate that the produced sets have reasonable size and slightly conservative coverage. In summary, we present the first general method for sequential changepoint localization, which is theoretically sound and broadly applicable in practice.
Betting on Bets: Anytime-Valid Tests for Stochastic Dominance (Sebastian Arnold*, YJ Choe*, Marco Scarsini, Ilia Tsetlin)
See Talks D3.
Verifying Elections with Adaptively Weighted Test Supermartingales (Alexander Ek, Michelle Blom, Philip B. Stark, Peter J. Stuckey, Damjan Vukcevic)
See Talks D3.
On vector-valued self-normalized concentration inequalities (Diego Martinez Taboada (Speaker), Tomas Gonzalez, Aaditya Ramdas)
See Talks D3.
Design-Based Anytime-Valid Inference for Randomized Experiments with Delayed Outcomes and Staggered Entry (Michael Lindon, Nathan Kallus)
See Talks D2.
Adaptive clinical trials based on design-optimal e-values: Application to single-arm trials (Stef Baas, Joost van Rosmalen, Judith ter Schure)
See Talks D2.
Downstream Post-Hoc Hypothesis Testing or: how E-Values generalize De Finetti’s Probability (Peter Grünwald, Ben Chugg, Aaditya Ramdas)
See Talks D1.
e-values as evidence (Ben Chugg, Peter Grünwald, Aaditya Ramdas)
See Talks D1.
Optimal Posterior E-values with Non-Convex Parameter Sets and Applications to Voting Systems (Timothée Mathieu, Adrienne Tuynman)
See Talks D1.
SSBBs: A method for determining the sample size when E-Backtesting the Expected Shortfall (Dennis Oestmann (Speaker), Thorsten Dickhaus)
See Talks D1.
Talks E1
Georgii Potapov
abstract tba
List of Participants
- Shubhada Agrawal
- Demissie Alemayehu
- Liviu Aolaritei
- Sebastian Arias
- Sebastian Arnold
- Morgane Austern
- Stef Baas
- Beepul Bharti
- Stephan Bongers
- Bastiaan Braams
- Joren Brunekreef
- François Caron
- Yo Joong Choe
- Ben Chugg
- Eugenio Clerico
- Gert de Cooman
- Fabian Damken
- Rabanus Derr
- Guneet Dhillon
- Alexander Ek
- Yixuan Fan
- Etienne Gauthier
- Jelle Goeman
- Peter Grünwald
- Rianne de Heide
- Roel Hulsman
- Nick Koning
- Stan Koobs
- Wouter Koolen
- Brian Lee
- Michael Lindon
- Noah Liniger
- Xiaoqi Shirley Liu
- Yongxi Long
- Nynke Luijten
- Alexander Ly
- Diego Martinez Taboada
- Valentina Masarotto
- Timothée Mathieu
- Lorenz Matz
- Sam van Meer
- Michele Meziu
- Aurèle Mingam
- Sohom Mukherjee
- Gergely Neu
- Dennis Oestmann
- Snigdha Panigrahi
- Victor H. de la Peña
- Victor K. de la Peña
- Georgii Potapov
- Abed Razawy
- Zhimei Ren
- Angel David Reyero Lobo
- Dante de Roos
- Joost van Rosmalen
- Johannes Ruf
- Ricardo Sandoval
- Antoine Scheid
- Gaëtan Serré
- Anurag Singh
- Martijn Sprokkereef
- Yury Tavyrikov
- Rovanos Tsafack Nzanguim
- Adrienne Tuynman
- Rajeev Verma
- Vladimir Vovk
- Qiuqi Wang
- Ruodu Wang
- Yongqi Wang
- Jan-Lukas Wermuth
- Redouane Yagouti