Joint VaR–ES Scoring: Fissler–Ziegel Forecast Evaluation
Learn why Value at Risk and Expected Shortfall need a joint scoring rule, how Fissler–Ziegel scores compare forecasts, and why scores answer a different question from VaR exception tests.
In this guideStart with the forecast pair and the tail convention
Short summary
VaR is a quantile and can be scored on its own with a quantile loss. Expected Shortfall is different: over broad distribution classes it is not elicitable as a standalone point forecast. Fissler and Ziegel show that VaR and ES can be evaluated together with a strictly consistent joint score under suitable conditions. That makes paired forecast comparison possible; it does not turn a low score into proof that a risk model is correct.
Start with the forecast pair and the tail convention
Let L_t be a loss, so larger values are worse, and let c be a confidence level such as 97.5%. Define the conditional loss quantile v(c,t) as the smallest x for which F(L_t ≤ x | information at t−1) reaches c. Under the upper-tail loss convention, Expected Shortfall is e(c,t) = [1/(1−c)] × the integral of v(u,t) from c to 1, when the tail mean is finite. With a continuous distribution, this is the average loss beyond the VaR threshold. If losses can have a mass at the threshold, the quantile-integral definition handles the required fraction of that mass.
The score must use the same convention as the forecasts. Some papers define VaR and ES for lower-tail returns, often with negative numbers, and write the tail probability as α = 1−c. For a loss-positive version, the same risk forecast is in the upper tail and ES is at least as large as VaR. Mixing a return convention with a loss convention can reverse inequalities and exception indicators. Record the sign, confidence level, horizon, and whether a forecast is conditional on information available before the outcome.
Why VaR and ES do not have the same standalone score
A quantile has a familiar strictly consistent score. For a loss observation L, one pinball score for its c-quantile v is S_VaR(v,L) = (I{L ≤ v}−c)(v−L). Across repeated comparable forecasts, its expected value is minimized at the target quantile, subject to the usual quantile conditions. This makes the score useful for evaluating VaR by itself.
There is no analogous strictly consistent score that elicits ES alone over the broad classes of loss distributions used in risk management. That statement has a specific scope: it is about a standalone point forecast and a score whose expected value uniquely selects ES. It does not say that ES is meaningless, untestable, or impossible to compare as part of a larger forecast.
Fissler and Ziegel establish the key distinction: although ES is not elicitable on its own in this sense, the pair (VaR, ES) is jointly elicitable under mild regularity and moment conditions. Their 2016 paper gives necessary and sufficient conditions for broad classes of strictly consistent scores and discusses the applied VaR–ES result (Fissler and Ziegel, 2016). “Jointly” means one scoring function takes both forecast components and the realized outcome as inputs. It does not mean any two scores can be added together and called a valid ES score.
What a Fissler–Ziegel score evaluates
Write S_c(v,e;L) for an admissible joint score at confidence c. Strict consistency means that, for the target distribution, the expected score is minimized at the correct pair: (VaR_c, ES_c) = arg min over (v,e) of E[S_c(v,e;L)]. Under the relevant conditions, the minimizer is unique. The score combines information about whether the VaR boundary was crossed with information about the realized tail loss and the ES forecast. The two forecasts therefore have to be coherent as a pair; for loss-positive upper tails, ES should not be below VaR.
Fissler–Ziegel describe a class of such scores, not a single mandatory formula. Patton, Ziegel, and Chen use that result to estimate and compare dynamic VaR–ES forecasts (Patton, Ziegel, and Chen, 2019). One commonly used FZ0 member is written for lower-tail returns Y = −L, α = 1−c, and negative forecasts e ≤ v < 0: S_FZ0(Y,v,e;α) = −I{Y ≤ v}(v−Y)/(αe) + v/e + ln(−e) − 1. That particular expression depends on its return sign and negativity assumptions. It is not a plug-in formula for every positive-loss dataset or every admissible joint score. In practice, document the chosen score and verify that its assumptions match the data and forecast convention.
Read score levels and score differences carefully
For model m, calculate one joint score for each forecast–outcome pair and then the average score: mean_S_m = (1/T) Σ_t S_c(v_hat(m,t), e_hat(m,t); L_t). When the same admissible score, dates, target, and conventions are used, the lower average is better relative to that scoring rule. A score is not a probability, a VaR amount, or an estimate of dollars lost. Its numeric level depends on the chosen score and scale.
For a paired comparison, form d_t = S_c(v_hat(A,t),e_hat(A,t);L_t) − S_c(v_hat(B,t),e_hat(B,t);L_t). A negative average difference favors A on the selected score. Because both forecasts face the same realization, this paired difference is more informative than comparing separately rounded average scores. But the difference can still be noisy, especially when only a few observations reach the tail.
Compare models on identical forecast origins, portfolio or target definition, horizon, confidence level, and data vintage. Forecasts must be recorded before the outcomes they are scored against. If horizons overlap, or score differences remain serially dependent, the uncertainty calculation needs to reflect that dependence. Report the mean difference with an appropriate standard error or interval; a small raw gap is not automatically meaningful.
<!-- learn:illustration -->

Why two standalone losses are not a substitute
It can be tempting to compute a VaR pinball loss and add an ES error such as the absolute difference between forecast ES and a proxy. That second term may not be a valid ES score, and an arbitrary weighted sum need not have the joint-consistency property. A joint FZ score is built so the pair of forecasts targets the VaR–ES functional together. Choosing weights or transformations without that property changes the question being evaluated.
There are multiple scores in the Fissler–Ziegel class. They can emphasize different aspects and have different scales, so a score difference from one choice is not directly comparable with a difference from another. A forecast ranking is meaningful only after the scoring rule has been selected and reported. Patton, Ziegel, and Chen discuss score design for practical risk forecasts, including an FZ0 choice with scale properties under additional sign assumptions. Those properties do not make FZ0 universally best.
The forecast pair also matters. A numerically attractive ES forecast paired with an incompatible VaR forecast may score poorly or violate the intended quantile-tail relationship. Conversely, a well-scored pair does not prove that every part of the conditional distribution is correct. It evaluates the stated functional under the chosen score and evaluation design.
Same VaR exceptions can hide different tail losses
A VaR exception indicator is binary: I_t = 1 when L_t exceeds its VaR forecast, and I_t = 0 otherwise. Two models can produce the same number and timing of exceptions while their observations beyond the threshold differ greatly in size. A count of four breaches records four breaches; it does not record whether their losses exceeded VaR by a small amount or by a much larger amount. ES is meant to describe average severity in the selected tail, so an evaluation that ignores the outcomes’ magnitudes cannot assess the VaR–ES pair.
A joint score uses the threshold and ES forecasts together with the realized loss, so tail outcomes can affect the score beyond the hit count. The chosen FZ score determines exactly how each observation contributes. Do not infer a numeric winner from one extreme loss without computing the published score on the full common sample. Nor should one rare event dominate a broad conclusion without sensitivity analysis.
This is a forecast-comparison use of scoring, not the same exercise as checking whether one model’s VaR exceptions have the target frequency. It is useful for asking which forecast pair performed better under a common rule; it is not, by itself, a complete calibration certificate.
Exception tests and ES backtests ask different questions
Kupiec-style VaR tests compare the number of threshold exceptions with a target rate; Christoffersen-style extensions also examine their dependence or clustering. These tests reduce each outcome to a hit or non-hit, so they do not measure the amount of loss beyond VaR and do not directly validate an ES forecast. The existing VaR and Expected Shortfall guide explains the two risk measures and why exception counts omit tail severity.
Joint scoring supports relative forecast comparison, while backtesting asks whether a forecast is adequately calibrated for a stated procedure. The questions are related but not interchangeable. Acerbi and Szekely proposed nonparametric ES backtests and argued that elicitability is relevant to model selection, not a prerequisite for testing a risk measure (Acerbi and Szekely, 2014). Bayer and Dimitriadis later developed regression-based ES backtests using joint VaR–ES structure; their variants have their own assumptions and covariance requirements (Bayer and Dimitriadis, 2022). A score ranking should not be presented as a substitute for a dedicated backtest, and a backtest non-rejection is not proof that a forecast pair is accurate.
Plan paired inference and prevent selection from leaking in
Treat the time series of score differences as data. If the evaluation design has non-overlapping one-step forecasts and suitable dependence assumptions, a paired test of the mean difference may be considered. With serial dependence or overlapping horizons, use inference designed for that dependence, such as an appropriate long-run variance estimate or block resampling. State the method and bandwidth or block rule rather than reporting only a p-value.
Keep score choice, tail level, forecast horizon, evaluation dates, and model set fixed before ranking results. If researchers try many scores, windows, assets, or forecast variants and report only the best one, the winning average score is affected by selection. Reserve a later evaluation period where possible and disclose how alternatives were screened. If forecasts were tuned on the reported test sample, the reported uncertainty does not account for that search automatically.
A sound comparison should therefore report both average scores and their paired difference with uncertainty, the forecasts’ VaR and ES conventions, the score definition, the common out-of-sample dates, and the number of tail observations. If rankings change under reasonable score choices or periods, show that sensitivity instead of hiding it.
Limits and a practical reporting checklist
High confidence levels leave few tail observations, so model ranking can be unstable even when the time series is long. Conditional heteroskedasticity, regime changes, overlapping returns, parameter estimation, data errors, and changes in portfolio composition can all affect score differences. The joint-elicitation result does not erase these problems; it supplies a principled score class under mathematical conditions.
Before interpreting a result, check that the realized loss matches the forecast’s portfolio, horizon, valuation method, and sign; that VaR and ES use the same tail probability; that forecasts were genuinely ex ante; and that the score is in the admissible class for the stated assumptions. Report sample length, realized tail counts, score formula, paired uncertainty procedure, and any score or model search. Explain that the conclusion is conditional on this design.
A lower joint score is evidence about comparative predictive performance under a particular proper scoring rule. It does not establish that a model captures every tail event, that its risk estimate is safe, or that future results will match the sample. This is statistical material, not investment advice or a recommendation to select a portfolio or risk model.
Common questions
Q1Can I compare Expected Shortfall forecasts with absolute error?
Not as a general replacement for a strictly consistent scoring rule. ES alone is not elicitable over broad distribution classes; a principled point-forecast comparison typically scores ES jointly with VaR, or uses a dedicated ES backtest for a calibration question.
Q2Does a joint VaR–ES score tell me whether a risk model is calibrated?
It evaluates relative performance under a selected joint score. Calibration and misspecification are separate questions, so pair the comparison with suitable backtests and diagnostics.
Q3Can two models tie on VaR exceptions but receive different joint scores?
Yes. Their hit indicators can match while the realized losses beyond VaR, or the ES forecasts, differ. A joint score uses both forecasts and outcomes; the exact effect depends on the score.
Q4Does a low score mean the model is safe to use?
No. It is a sample-based comparison under chosen assumptions and a chosen score. Tail data are sparse, and regime changes, model selection, or data problems can alter the result. This article is educational and is not investment advice.
Sources and further reading
Report an issue
We’ll prepare an email with this article link. Mark receives the report only after you send it
Quick check
Read the guide? Check yourself with 3 questions
Question 01
Why is Expected Shortfall commonly scored together with VaR?
Choose an answer to see the explanation
Options glossary
The average loss from a chosen VaR quantile through the worst tail under a precise convention; it measures severity beyond the threshold rather than only its location.
Read the deeper guideAssignmentThe process that requires an option writer to fulfill the contract after an exercise notice is allocated; it can create or remove an underlying position.
Read the deeper guideBid-ask spreadThe gap between the best displayed bid and ask, which is a practical trading cost and a signal of how uncertain an immediate fill may be.
Read the deeper guide