You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Implementations of covariance estimators and confidence intervals
Introduction
This package provides numerically stable, online (streaming) estimators
for means, covariance matrices, and correlation matrices,
along with methods for computing confidence intervals on the covariance estimates.
Estimators included:
OnlineCovariance - Welford's one-pass algorithm for incremental mean and covariance estimation, with support for merging independent estimators
EMACovariance - Exponential moving average covariance, parameterized by alpha, halflife, or span
SMACovariance - Simple moving average covariance over a fixed rolling window
All estimators support a geometric=True mode that applies a log transform
to observations, suitable for multiplicative processes like financial returns,
and a frequency parameter for annualizing results (e.g. frequency=252 for daily data).
Confidence interval methods for the sample covariance matrix include:
asymptotic - Normal approximation using the Wishart variance formula
wishart - Monte Carlo sampling from the Wishart distribution
bootstrap - Bootstrap resampling
parametric_mc - Parametric Monte Carlo from a fitted normal
Basic usage
importnumpyasnpfromcovariance_calculators.estimatorsimportOnlineCovariancefromcovariance_calculators.intervalsimportcalc_covariance_intervals## Generate toy data: 500 samples from a 3d normal distributionrng=np.random.default_rng(42)
true_cov=np.array([[1.0, 0.5, 0.2],
[0.5, 2.0, 0.3],
[0.2, 0.3, 1.5]])
L=np.linalg.cholesky(true_cov)
data=rng.standard_normal((500, 3)) @ L.T## Stream data through the online estimatoroc=OnlineCovariance(order=3)
forrowindata:
oc.add(row)
print("mean =")
print(oc.mean)
print("cov =")
print(oc.cov)
print("corr =")
print(oc.corr)
## Compute 95% confidence intervals on the covariance estimatecov, ci_lower, ci_upper=calc_covariance_intervals(
data=data,
confidence_level=0.95,
method="asymptotic",
)
print("ci_lower =")
print(ci_lower)
print("ci_upper =")
print(ci_upper)
This will create a .venv virtualenv (if one doesn't already exist),
install the dependencies from requirements.txt, and add the parent
directory to your PYTHONPATH.
Run the tests:
make test
You can also run the linter with:
make lint
Coverage tests
Shows that the confidence intervals are well calibrated at 1, 2, 3, and 4 $\sigma$ confidence levels:
Note that the update term for the online covariance is a term in a scatter matrix, $S$,
using the currently observed data, $x_{n}$, and the previous means, $\mu_{n-1}$.
But also note that the $\delta \delta^\prime$ form is also
convenient because it comes naturally normalized and can be readily generalized for weighting.
where by summing a geometric series, one can show that for exponential weighting, $W_{n} = 1$,
so $\hat{V} = C_{n}$.
Theory of confidence intervals
Cochran's theorem
For $n$ i.i.d. samples from a normal distribution, $x_i \sim N(\mu, \sigma^2)$,
Cochran's theorem gives the sampling distribution of the MLE variance estimator: