Abstract
Generating extreme events poses two distinct obstacles, and generative models may easily fail at both. The difficulty is not confined to heavy tails, though those are the hardest case, and in the multivariate setting different directions may sit in different tail regimes, none of them known in advance.
The first obstacle is in the generative objective itself, where every standard choice imposes regularity on the measures: KL and f-divergences require absolute continuity and can easily fail on singular or heavy-tailed targets, while Wasserstein distances require moments of both arguments and may not have well-defined variational derivatives. Lipschitz-regularized divergences, by contrast, are infimal convolutions of a divergence with a transport cost, i.e. a proximal regularization in the Wasserstein-1 geometry, and inherit the best of both, and more: the only assumption needed is on the source, which we design (a finite first moment), and, most importantly, none at all on the target, singular or heavy-tailed alike. The variational derivatives then exist, are finite and unique, and the Wasserstein gradient flow is well posed on targets other objectives cannot handle. Its dissipation inequality also yields a principled, data-driven stopping criterion in place of a preset number of steps.
The second obstacle is that the transport itself cannot be regular for learning heavier tails than the source's. A Lipschitz map sends a light-tailed source to a light-tailed output, so a well-behaved flow never reaches the tail. CVaR, which measures tail risk, is blind to the bulk and aggregates the whole tail into a single scalar observable, so relatively few tail samples suffice to estimate it reliably, where densities and their gradients cannot be resolved at all; adding it to the generative objective - the two terms together forming a free energy - revives the transport in the sparsely sampled tail, through a bounded but non-Lipschitz velocity. The resulting correction is tail-agnostic and applies to the samples of any pre-trained model without access to its architecture; across synthetic and real target data it sharply improves both bulk and tail accuracy over pre-trained baselines and over recent tail-adapted GANs, diffusions and normalizing flows.