Class
NumCosmoMathStatsDist
Description [src]
abstract class NumCosmoMath.StatsDist : GObject.Object
{
/* No available fields */
}
Base class for N-dimensional probability distributions reconstructed from samples.
Represents a distribution as a mixture of radial-basis kernels placed on a set of sample points, evaluates it and draws from it. The theory, and what the shrinkage and kernel-selection options do, are on the Kernel Mixture Densities page.
This is an abstract class. NcmStatsDistKDE gives all kernels a common bandwidth,
NcmStatsDistVKDE lets it vary between sample points; both build the interpolation
matrix and the covariance decompositions. The kernel profile comes from a
NcmStatsDistKernel, either NcmStatsDistKernelGauss or NcmStatsDistKernelST.
Add the sample points with ncm_stats_dist_add_obs(), then call ncm_stats_dist_prepare():
without the sample’s $-2\ln g(x_i)$ it weights every kernel equally, with them it solves
a non-negative least-squares problem for the weights. Either way the object is then
ready for ncm_stats_dist_eval() and ncm_stats_dist_sample().
Nothing else has to be set. NcmStatsDist:over-smooth sets the bandwidth,
NcmStatsDist:CV-type the cross-validation that can fit it and
NcmStatsDist:split-frac the out-of-sample fraction; NcmStatsDist:center-shrink
contracts the mixture so its covariance matches the sample, and
NcmStatsDist:auto-kernel fits the tail of the NcmStatsDistKernelST the object was
built with along with the bandwidth, which requires a cross-validation that fits it.
NcmStatsDist:defensive-frac mixes a wide Student-t component into the proposal so
that its density has no holes; it is off by default.
Center shrinkage requires a kernel with a finite covariance: a NcmStatsDistKernelST
with $\nu \leq 2$, the Cauchy kernel included, aborts. It also raises the lower
bound on a fitted $\nu$ from $1$ to $2.5$.
Instance methods
ncm_stats_dist_add_obs
Adds a copy of y to the sample, which the next ncm_stats_dist_prepare() uses.
ncm_stats_dist_eval
Evaluates the density at $\vec{x} = $ x. Use ncm_stats_dist_eval_m2lnp() where it
may underflow.
ncm_stats_dist_eval_m2lnp
Evaluates $-2\ln P(\vec{x})$ at $\vec{x} = $ x, summing the kernels in the
logarithm so that it does not underflow.
ncm_stats_dist_eval_m2lnp_vec
Evaluates the distribution at every point of x_a, writing $-2\ln P(\vec{x}_i)$ into
element $i$ of m2lnp. Equivalent to calling ncm_stats_dist_eval_m2lnp() once per
point, but subclasses may share work across the batch, so the two need not agree to
the last bit.
ncm_stats_dist_get_Ki
Gets kernel i: its center, its scale matrix at the applied bandwidth (the covariance
is $\kappa$ times it), its normalization and its weight.
ncm_stats_dist_get_center_shrink_factor
Gets the scalar part $a = \det(A)^{1/d}$ of the center shrinkage transform $A$
applied in the last preparation, see the class description. The bandwidth in use is
$a$ times the nominal one. It is $1$ when center shrinkage is disabled, and for
NcmStatsDistKDE with the sample covariance $A = a\,I$ exactly.
ncm_stats_dist_get_href
Gets the bandwidth applied to the kernels in the last preparation, the same one
used by ncm_stats_dist_get_lnnorm() and ncm_stats_dist_get_Ki(). It is the nominal
bandwidth the subclass derives from NcmStatsDist:over-smooth, see the class
description, times the center shrinkage factor
ncm_stats_dist_get_center_shrink_factor(), which is one unless
NcmStatsDist:center-shrink is set. It is zero before the first preparation.
ncm_stats_dist_get_lnnorm
Gets the logarithm of the normalization $N_i$ of kernel i at the applied bandwidth,
$K_i(x) = \bar{K}(\chi^2) / N_i$, see NcmStatsDistKernel.
ncm_stats_dist_get_rnorm
Gets the squared residual $|M w - 1|^2$ of the last non-negative least-squares weight
fit, with $M$ the interpolation matrix whose rows are divided by the target density
and $w$ the weights before normalization. Zero when the last
ncm_stats_dist_prepare() did not fit the weights.
ncm_stats_dist_peek_center_array
Gets the kernel centers $c_i$ used in the last preparation, one per kernel. They
coincide with the first ncm_stats_dist_get_n_kernels() sample points unless center
shrinkage is enabled, see the class description.
ncm_stats_dist_peek_center_shrink_matrix
Gets the $d \times d$ center shrinkage transform $A$ applied in the last
preparation: the centers are $c_i = \mu + A (x_i - \mu)$ and the kernel scale
matrices $\hat A \Sigma_i \hat A^T$ with $\hat A = A / a$, see the class
description. It is the identity when center shrinkage is disabled, and NULL before
the first preparation.
ncm_stats_dist_peek_cov_decomp
Gets the upper-triangular Cholesky factor $U_i$ of the scale matrix
$\Sigma_i = U_i^T U_i$ of kernel i, as applied in the last preparation. The kernel
covariance is $\kappa h^2 \Sigma_i$, with $h$ from ncm_stats_dist_get_href().
ncm_stats_dist_peek_full_cov
Gets the subclass’s global scale matrix: for NcmStatsDistKDE the scale matrix shared
by all kernels (the sample, fixed or robust covariance).
ncm_stats_dist_peek_full_cov_decomp
Gets the upper-triangular Cholesky factor of the subclass’s global scale matrix, ncm_stats_dist_peek_full_cov(), as applied in the last preparation.
ncm_stats_dist_prepare
Builds the estimator from the sample: the kernel shapes, the bandwidth (fitted by the
cross-validation chosen with NcmStatsDist:CV-type) and the weights. Leaves the object
ready for ncm_stats_dist_eval() and ncm_stats_dist_sample().
ncm_stats_dist_prepare_shapes
Runs the first stage of a prepare alone: the per-kernel covariance structures the
subclass builds from the sample, before any bandwidth is applied, over the kernel count
of the last ncm_stats_dist_prepare(). For inspecting those structures;
ncm_stats_dist_prepare() runs it as part of the whole pipeline, and only after that is
the object ready for ncm_stats_dist_eval().
ncm_stats_dist_sample
Draws a point from the mixture, including the wide component of
NcmStatsDist:defensive-frac, and stores it in x.
ncm_stats_dist_set_auto_kernel
Sets whether the kernel is chosen together with the over-smooth factor, see
NcmStatsDist:auto-kernel.
ncm_stats_dist_set_center_shrink
Enables or disables center shrinkage, see the class description. Takes effect at the next call to ncm_stats_dist_prepare().
ncm_stats_dist_set_cv_type
Sets NcmStatsDist:CV-type, the objective that fits the bandwidth at the next
ncm_stats_dist_prepare(). With #NCM_STATS_DIST_CV_NONE the bandwidth is the
rule-of-thumb one times NcmStatsDist:over-smooth.
ncm_stats_dist_set_defensive_frac
Sets NcmStatsDist:defensive-frac. Takes effect at the next preparation: until then
evaluation and sampling use the fraction of the last one.
ncm_stats_dist_set_defensive_nu
Sets NcmStatsDist:defensive-nu. Takes effect at the next preparation.
ncm_stats_dist_set_defensive_scale
Sets NcmStatsDist:defensive-scale. Takes effect at the next preparation.
ncm_stats_dist_set_kernel
Sets the kernel of the mixture, NcmStatsDistKernelGauss or NcmStatsDistKernelST,
and the dimension to the kernel’s.
ncm_stats_dist_set_over_smooth
Sets NcmStatsDist:over-smooth. With a cross-validation it is the starting point of
the fit.
ncm_stats_dist_set_split_frac
Sets NcmStatsDist:split-frac, read by the split cross-validations. A fraction that
leaves no out-of-sample point aborts at the next ncm_stats_dist_prepare().
ncm_stats_dist_set_uniform_weights
Sets NcmStatsDist:uniform-weights. Takes effect at the next preparation.
Properties
NumCosmoMath.StatsDist:CV-type
The NcmStatsDistCV that fits the bandwidth at ncm_stats_dist_prepare().
Default: #NCM_STATS_DIST_CV_NONE.
NumCosmoMath.StatsDist:auto-kernel
Whether the kernel is chosen together with the over-smooth factor, by the same
objective. It applies to every NcmStatsDist:CV-type except
NCM_STATS_DIST_CV_NONE, which fits nothing. The kernel must be the
NcmStatsDistKernelST the object was built with; any other kernel aborts at
prepare. Its degrees of freedom $\nu$ are fitted in place, jointly with the
over-smooth factor, over $\nu \in [\nu_\mathrm{min}, 10^4]$, with
$\nu_\mathrm{min} = 2.5$ when NcmStatsDist:center-shrink is set and $1$
otherwise. At $\nu = 10^4$ the log-density differs from the Gaussian kernel’s by
about $[(\chi^2 - d)^2 - 2d] / (4\nu)$: over the central 99.8% of $\chi^2$, from
$2 \times 10^{-3}$ at $d = 1$ to $6 \times 10^{-2}$ at $d = 100$. Default: FALSE.
NumCosmoMath.StatsDist:center-shrink
Whether to shrink the kernel centers toward the sample mean so that the covariance of the kernel mixture matches the sample covariance for any bandwidth, see the class description. Default: FALSE.
NumCosmoMath.StatsDist:defensive-frac
Weight $\epsilon$ of a wide Student-t component mixed into the proposal,
$q = (1 - \epsilon)\, q_\mathrm{mixture} + \epsilon\, t_\nu(\mu, c\,C)$, the
Student-t density with location $\mu$ and scale matrix $c\,C$, where $\mu$ and $C$
are the sample mean and covariance, $c$ is NcmStatsDist:defensive-scale and $\nu$
is NcmStatsDist:defensive-nu. It bounds the proposal density from below where the
kernels leave holes, so that a walker there can still move. Zero disables it and
leaves every evaluation and draw unchanged. Default: 0.
NumCosmoMath.StatsDist:defensive-nu
Degrees of freedom $\nu$ of the wide Student-t component of
NcmStatsDist:defensive-frac. Default: 3.
NumCosmoMath.StatsDist:defensive-scale
Factor $c$ multiplying the sample covariance in the scale matrix of the wide
component of NcmStatsDist:defensive-frac. Default: 4.
NumCosmoMath.StatsDist:over-smooth
Factor multiplying the kernel rule-of-thumb bandwidth, ncm_stats_dist_kernel_get_rot_bandwidth(). A cross-validation other than
NCM_STATS_DIST_CV_NONE fits it in $[10^{-2}, 20]$, starting from the current
value. NcmStatsDistVKDE uses it as the bandwidth itself unless
NcmStatsDistVKDE:use-rot-href is set. Default: 1.
NumCosmoMath.StatsDist:print-fit
Whether the bandwidth fit prints its progress to the standard output. Default: FALSE.
NumCosmoMath.StatsDist:split-frac
Fraction $f$ of the sample used as kernel centers by the split cross-validations,
NCM_STATS_DIST_CV_SPLIT_M2LNP and #NCM_STATS_DIST_CV_SPLIT_ACCEPT: the first
$\lceil f n \rceil$ points, in the order they were added, are the centers and the others score the bandwidth. Default: 0.5.
NumCosmoMath.StatsDist:uniform-weights
Whether ncm_stats_dist_prepare() keeps uniform kernel weights instead of
fitting them by NNLS. The sample’s $-2\ln L$ is then used only by the bandwidth
objectives that need it and by the outlier cut. Default: FALSE.
NumCosmoMath.StatsDist:use-threads
Whether NcmStatsDistVKDE runs its OpenMP loops in parallel; the other classes do
not read it. Default: FALSE.
Signals
Signals inherited from GObject (1)
GObject::notify
The notify signal is emitted on an object when one of its properties has its value set through g_object_set_property(), g_object_set(), et al.
Class structure
struct NumCosmoMathStatsDistClass {
void (* set_dim) (
NcmStatsDist* sd,
const guint dim
);
gdouble (* bandwidth) (
NcmStatsDist* sd
);
void (* prepare_shapes) (
NcmStatsDist* sd,
GPtrArray* sample_array
);
void (* prepare_kernels) (
NcmStatsDist* sd
);
void (* compute_IM) (
NcmStatsDist* sd,
NcmMatrix* IM
);
gdouble (* amise) (
NcmStatsDist* sd
);
NcmMatrix* (* peek_cov_decomp) (
NcmStatsDist* sd,
guint i
);
NcmMatrix* (* peek_full_cov_decomp) (
NcmStatsDist* sd
);
NcmMatrix* (* peek_full_cov) (
NcmStatsDist* sd
);
gdouble (* get_lnnorm) (
NcmStatsDist* sd,
guint i
);
gdouble (* eval_weights) (
NcmStatsDist* sd,
NcmVector* weights,
NcmVector* x
);
gdouble (* eval_weights_m2lnp) (
NcmStatsDist* sd,
NcmVector* weights,
NcmVector* x
);
void (* eval_weights_m2lnp_vec) (
NcmStatsDist* sd,
NcmVector* weights,
GPtrArray* x_a,
NcmVector* m2lnp
);
void (* eval_weights_m2lnp_loo) (
NcmStatsDist* sd,
NcmVector* weights,
GPtrArray* x_a,
NcmVector* m2lnp
);
void (* reset) (
NcmStatsDist* sd
);
}
No description available.
Class members
set_dim: void (* set_dim) ( NcmStatsDist* sd, const guint dim )No description available.
bandwidth: gdouble (* bandwidth) ( NcmStatsDist* sd )No description available.
prepare_shapes: void (* prepare_shapes) ( NcmStatsDist* sd, GPtrArray* sample_array )No description available.
prepare_kernels: void (* prepare_kernels) ( NcmStatsDist* sd )No description available.
compute_IM: void (* compute_IM) ( NcmStatsDist* sd, NcmMatrix* IM )No description available.
amise: gdouble (* amise) ( NcmStatsDist* sd )No description available.
peek_cov_decomp: NcmMatrix* (* peek_cov_decomp) ( NcmStatsDist* sd, guint i )No description available.
peek_full_cov_decomp: NcmMatrix* (* peek_full_cov_decomp) ( NcmStatsDist* sd )No description available.
peek_full_cov: NcmMatrix* (* peek_full_cov) ( NcmStatsDist* sd )No description available.
get_lnnorm: gdouble (* get_lnnorm) ( NcmStatsDist* sd, guint i )No description available.
eval_weights: gdouble (* eval_weights) ( NcmStatsDist* sd, NcmVector* weights, NcmVector* x )No description available.
eval_weights_m2lnp: gdouble (* eval_weights_m2lnp) ( NcmStatsDist* sd, NcmVector* weights, NcmVector* x )No description available.
eval_weights_m2lnp_vec: void (* eval_weights_m2lnp_vec) ( NcmStatsDist* sd, NcmVector* weights, GPtrArray* x_a, NcmVector* m2lnp )No description available.
eval_weights_m2lnp_loo: void (* eval_weights_m2lnp_loo) ( NcmStatsDist* sd, NcmVector* weights, GPtrArray* x_a, NcmVector* m2lnp )No description available.
reset: void (* reset) ( NcmStatsDist* sd )No description available.
Virtual methods
NumCosmoMath.StatsDistClass.get_lnnorm
Gets the logarithm of the normalization $N_i$ of kernel i at the applied bandwidth,
$K_i(x) = \bar{K}(\chi^2) / N_i$, see NcmStatsDistKernel.
NumCosmoMath.StatsDistClass.peek_cov_decomp
Gets the upper-triangular Cholesky factor $U_i$ of the scale matrix
$\Sigma_i = U_i^T U_i$ of kernel i, as applied in the last preparation. The kernel
covariance is $\kappa h^2 \Sigma_i$, with $h$ from ncm_stats_dist_get_href().
NumCosmoMath.StatsDistClass.peek_full_cov
Gets the subclass’s global scale matrix: for NcmStatsDistKDE the scale matrix shared
by all kernels (the sample, fixed or robust covariance).
NumCosmoMath.StatsDistClass.peek_full_cov_decomp
Gets the upper-triangular Cholesky factor of the subclass’s global scale matrix, ncm_stats_dist_peek_full_cov(), as applied in the last preparation.
NumCosmoMath.StatsDistClass.prepare_shapes
Runs the first stage of a prepare alone: the per-kernel covariance structures the
subclass builds from the sample, before any bandwidth is applied, over the kernel count
of the last ncm_stats_dist_prepare(). For inspecting those structures;
ncm_stats_dist_prepare() runs it as part of the whole pipeline, and only after that is
the object ready for ncm_stats_dist_eval().