ROBUST INFERENCE FOR DIFFERENCES OF SHARPE RATIOS
=================================================

Michael Wolf
University of Zurich
September 2026


This collection of R routines provides robust inference for differences
of Sharpe ratios. It supersedes all previous versions of these routines.

The routines are designed for inference on the difference between two
Sharpe ratios. They do not provide inference for an individual Sharpe ratio.


GETTING STARTED
===============

The file Sharpe.RData contains all routines as well as the two datasets
used in the empirical applications of Ledoit and Wolf (2008). To load
everything into the R workspace, use

> load("Sharpe.RData")

The individual .R files containing the source code for all routines are
also included in the distribution. They are not needed if Sharpe.RData
has been loaded, but are provided for users who wish to inspect or modify
the code.


MAIN ROUTINES
=============

hac.sharpe
--------------------

Carries out HAC inference for the difference between two Sharpe ratios.

The routine reports the two estimated Sharpe ratios, their estimated
difference, HAC standard errors, and corresponding two-sided p-values.
Both the standard HAC estimator and its prewhitened version are reported.
The Parzen kernel is used.

Example:

> hac.sharpe(ret.agg)


bootstrap.sharpe
----------------

Carries out studentized circular block bootstrap inference for the
difference between two Sharpe ratios.

The user supplies the block size b and the number M of bootstrap
replications. The default null hypothesis is that the difference between
the two Sharpe ratios is zero.

Example:

> bootstrap.sharpe(ret.agg, b = 6, M = 9999)

The routine returns the estimated difference in Sharpe ratios and the
two-sided bootstrap p-value.

A data-dependent block size can be obtained using
block.size.calibrate.sharpe.


block.size.calibrate.sharpe
---------------------------

Selects the block size for the circular block bootstrap using the
calibration procedure of Ledoit and Wolf (2008).

The default candidate block sizes are

    b.vec = c(1, 3, 6, 10)

the default number of bootstrap replications used for each test is

    M = 499

and the default number of calibration samples is

    K = 1000.

Example:

> block.size.calibrate.sharpe(ret.agg)

The routine reports the estimated rejection probability for each
candidate block size and the selected block size.

The calibration is computationally more demanding than the other
routines and may take several minutes. We recommend running this
function separately and then supplying the selected block size to
bootstrap.sharpe.

Users should inspect the estimated rejection probabilities rather than
rely mechanically on the selected block size. In particular, if the
selected block size is at the boundary of b.vec and the rejection
probabilities are still moving toward the nominal significance level,
the candidate set should be extended and the calibration rerun.

boot.stats.sharpe
-----------------

For a given circular block bootstrap sample and block size, computes

    Delta.hat.star    the bootstrap difference in sample Sharpe ratios

and

    se.star           the corresponding bootstrap standard error.

This is primarily a supporting routine and is called internally by
bootstrap.sharpe.

It can also be useful when the Sharpe-ratio inference provided by these
routines is embedded in a multiple-testing procedure. In particular,
hac.inference.sharpe and boot.stats.sharpe provide low-level ingredients
needed to construct studentized statistics for a collection of
Sharpe-ratio differences. These quantities can then be used with
higher-level multiple-testing procedures such as the Romano-Wolf
stepdown methodology; see Romano and Wolf (2016).

To preserve dependence across hypotheses, bootstrap samples must be
generated jointly across all assets or strategies involved. The
function cbb.sequence can be used to generate a common bootstrap index
sequence, which is then used to re-index the rows of the original data
matrix.

Users interested only in bootstrap inference for a single difference
of Sharpe ratios do not need to call boot.stats.sharpe directly.


SUPPORTING ROUTINES
===================

The distribution contains a number of supporting routines used internally
for HAC estimation, prewhitening, circular and stationary block bootstrap
sampling, block-size calibration, and computation of differences in
Sharpe ratios.

The source code for each routine is supplied in a separate .R file.

Users interested only in standard inference will normally need to call
only hac.inference.sharpe, bootstrap.sharpe, and, if desired,
block.size.calibrate.sharpe.

Users implementing their own bootstrap or multiple-testing procedures
may additionally find boot.stats.sharpe useful.


IMPLEMENTATION
==============

The HAC routines use the Parzen kernel. Both standard and prewhitened
HAC inference are provided.

Bootstrap inference uses the studentized circular block bootstrap.

The block-size calibration uses the stationary bootstrap to generate
the calibration samples.

The current implementation uses vectorized R code where useful for
computational efficiency.


DATA
====

The input return matrix for the main inference routines must have two
columns, one for each return series. Returns should be in excess of the
relevant risk-free rate. Thus, ret must be a T x 2 matrix rather than a
2 x T matrix.

The dataset ret.agg contains the mutual-fund data used in the empirical
application of Ledoit and Wolf (2008).

The dataset ret.hedge contains the hedge-fund data used in the empirical
application of Ledoit and Wolf (2008).

Both datasets are included in Sharpe.RData.


FILES
=====

Sharpe.RData contains all routines and the two empirical datasets.

In addition, the distribution contains one .R source file for each
routine. These source files are provided for transparency and for users
who wish to inspect or modify the implementation.

Loading Sharpe.RData is sufficient for normal use of the routines.


REFERENCES
==========

Ledoit, O. and Wolf, M. (2008).
Robust performance hypothesis testing with the Sharpe ratio.
Journal of Empirical Finance 15, 850-859.

Romano, J. P. and Wolf, M. (2016).
Efficient computation of adjusted p-values for resampling-based stepdown
multiple testing.
Statistics & Probability Letters 113, 38-40.
