Class

NumCosmoMathStatsDist

Description [src]

abstract class NumCosmoMath.StatsDist : GObject.Object
{
  /* No available fields */
}

Base class for N-dimensional probability distributions reconstructed from samples.

Represents a distribution as a mixture of radial-basis kernels placed on a set of sample points, evaluates it and draws from it. The theory, and what the shrinkage and kernel-selection options do, are on the Kernel Mixture Densities page.

This is an abstract class. NcmStatsDistKDE gives all kernels a common bandwidth, NcmStatsDistVKDE lets it vary between sample points; both build the interpolation matrix and the covariance decompositions. The kernel profile comes from a NcmStatsDistKernel, either NcmStatsDistKernelGauss or NcmStatsDistKernelST.

Add the sample points with ncm_stats_dist_add_obs(), then call ncm_stats_dist_prepare(): without the sample’s $-2\ln g(x_i)$ it weights every kernel equally, with them it solves a non-negative least-squares problem for the weights. Either way the object is then ready for ncm_stats_dist_eval() and ncm_stats_dist_sample().

Nothing else has to be set. NcmStatsDist:over-smooth sets the bandwidth, NcmStatsDist:CV-type the cross-validation that can fit it and NcmStatsDist:split-frac the out-of-sample fraction; NcmStatsDist:center-shrink contracts the mixture so its covariance matches the sample, and NcmStatsDist:auto-kernel fits the tail of the NcmStatsDistKernelST the object was built with along with the bandwidth, which requires a cross-validation that fits it.

NcmStatsDist:defensive-frac mixes a wide Student-t component into the proposal so that its density has no holes; it is off by default.

Center shrinkage requires a kernel with a finite covariance: a NcmStatsDistKernelST with $\nu \leq 2$, the Cauchy kernel included, aborts. It also raises the lower bound on a fitted $\nu$ from $1$ to $2.5$.

Ancestors

Descendants

Functions

ncm_stats_dist_clear

Decreases the reference count of sd and sets the pointer sd to NULL.

Instance methods

ncm_stats_dist_add_obs

Adds a copy of y to the sample, which the next ncm_stats_dist_prepare() uses.

ncm_stats_dist_eval

Evaluates the density at $\vec{x} = $ x. Use ncm_stats_dist_eval_m2lnp() where it may underflow.

ncm_stats_dist_eval_m2lnp

Evaluates $-2\ln P(\vec{x})$ at $\vec{x} = $ x, summing the kernels in the logarithm so that it does not underflow.

ncm_stats_dist_eval_m2lnp_vec

Evaluates the distribution at every point of x_a, writing $-2\ln P(\vec{x}_i)$ into element $i$ of m2lnp. Equivalent to calling ncm_stats_dist_eval_m2lnp() once per point, but subclasses may share work across the batch, so the two need not agree to the last bit.

ncm_stats_dist_free

Decreases the reference count of sd.

ncm_stats_dist_get_Ki

Gets kernel i: its center, its scale matrix at the applied bandwidth (the covariance is $\kappa$ times it), its normalization and its weight.

ncm_stats_dist_get_auto_kernel
No description available.

ncm_stats_dist_get_center_shrink
No description available.

ncm_stats_dist_get_center_shrink_factor

Gets the scalar part $a = \det(A)^{1/d}$ of the center shrinkage transform $A$ applied in the last preparation, see the class description. The bandwidth in use is $a$ times the nominal one. It is $1$ when center shrinkage is disabled, and for NcmStatsDistKDE with the sample covariance $A = a\,I$ exactly.

ncm_stats_dist_get_cv_type
No description available.

ncm_stats_dist_get_defensive_frac
No description available.

ncm_stats_dist_get_defensive_nu
No description available.

ncm_stats_dist_get_defensive_scale
No description available.

ncm_stats_dist_get_dim
No description available.

ncm_stats_dist_get_href

Gets the bandwidth applied to the kernels in the last preparation, the same one used by ncm_stats_dist_get_lnnorm() and ncm_stats_dist_get_Ki(). It is the nominal bandwidth the subclass derives from NcmStatsDist:over-smooth, see the class description, times the center shrinkage factor ncm_stats_dist_get_center_shrink_factor(), which is one unless NcmStatsDist:center-shrink is set. It is zero before the first preparation.

ncm_stats_dist_get_kernel
No description available.

ncm_stats_dist_get_lnnorm

Gets the logarithm of the normalization $N_i$ of kernel i at the applied bandwidth, $K_i(x) = \bar{K}(\chi^2) / N_i$, see NcmStatsDistKernel.

ncm_stats_dist_get_n_kernels
No description available.

ncm_stats_dist_get_over_smooth
No description available.

ncm_stats_dist_get_print_fit
No description available.

ncm_stats_dist_get_rnorm

Gets the squared residual $|M w - 1|^2$ of the last non-negative least-squares weight fit, with $M$ the interpolation matrix whose rows are divided by the target density and $w$ the weights before normalization. Zero when the last ncm_stats_dist_prepare() did not fit the weights.

ncm_stats_dist_get_sample_size
No description available.

ncm_stats_dist_get_split_frac
No description available.

ncm_stats_dist_get_uniform_weights
No description available.

ncm_stats_dist_get_use_threads
No description available.

ncm_stats_dist_kernel_choose

Draws a kernel index with probability equal to its weight.

ncm_stats_dist_peek_center_array

Gets the kernel centers $c_i$ used in the last preparation, one per kernel. They coincide with the first ncm_stats_dist_get_n_kernels() sample points unless center shrinkage is enabled, see the class description.

ncm_stats_dist_peek_center_shrink_matrix

Gets the $d \times d$ center shrinkage transform $A$ applied in the last preparation: the centers are $c_i = \mu + A (x_i - \mu)$ and the kernel scale matrices $\hat A \Sigma_i \hat A^T$ with $\hat A = A / a$, see the class description. It is the identity when center shrinkage is disabled, and NULL before the first preparation.

ncm_stats_dist_peek_cov_decomp

Gets the upper-triangular Cholesky factor $U_i$ of the scale matrix $\Sigma_i = U_i^T U_i$ of kernel i, as applied in the last preparation. The kernel covariance is $\kappa h^2 \Sigma_i$, with $h$ from ncm_stats_dist_get_href().

ncm_stats_dist_peek_full_cov

Gets the subclass’s global scale matrix: for NcmStatsDistKDE the scale matrix shared by all kernels (the sample, fixed or robust covariance).

ncm_stats_dist_peek_full_cov_decomp

Gets the upper-triangular Cholesky factor of the subclass’s global scale matrix, ncm_stats_dist_peek_full_cov(), as applied in the last preparation.

ncm_stats_dist_peek_kernel
No description available.

ncm_stats_dist_peek_sample_array
No description available.

ncm_stats_dist_peek_weights
No description available.

ncm_stats_dist_prepare

Builds the estimator from the sample: the kernel shapes, the bandwidth (fitted by the cross-validation chosen with NcmStatsDist:CV-type) and the weights. Leaves the object ready for ncm_stats_dist_eval() and ncm_stats_dist_sample().

ncm_stats_dist_prepare_shapes

Runs the first stage of a prepare alone: the per-kernel covariance structures the subclass builds from the sample, before any bandwidth is applied, over the kernel count of the last ncm_stats_dist_prepare(). For inspecting those structures; ncm_stats_dist_prepare() runs it as part of the whole pipeline, and only after that is the object ready for ncm_stats_dist_eval().

ncm_stats_dist_ref

Increases the reference count of sd.

ncm_stats_dist_reset

Discards all sample points added so far.

ncm_stats_dist_sample

Draws a point from the mixture, including the wide component of NcmStatsDist:defensive-frac, and stores it in x.

ncm_stats_dist_set_auto_kernel

Sets whether the kernel is chosen together with the over-smooth factor, see NcmStatsDist:auto-kernel.

ncm_stats_dist_set_center_shrink

Enables or disables center shrinkage, see the class description. Takes effect at the next call to ncm_stats_dist_prepare().

ncm_stats_dist_set_cv_type

Sets NcmStatsDist:CV-type, the objective that fits the bandwidth at the next ncm_stats_dist_prepare(). With #NCM_STATS_DIST_CV_NONE the bandwidth is the rule-of-thumb one times NcmStatsDist:over-smooth.

ncm_stats_dist_set_defensive_frac

Sets NcmStatsDist:defensive-frac. Takes effect at the next preparation: until then evaluation and sampling use the fraction of the last one.

ncm_stats_dist_set_defensive_nu

Sets NcmStatsDist:defensive-nu. Takes effect at the next preparation.

ncm_stats_dist_set_defensive_scale

Sets NcmStatsDist:defensive-scale. Takes effect at the next preparation.

ncm_stats_dist_set_kernel

Sets the kernel of the mixture, NcmStatsDistKernelGauss or NcmStatsDistKernelST, and the dimension to the kernel’s.

ncm_stats_dist_set_over_smooth

Sets NcmStatsDist:over-smooth. With a cross-validation it is the starting point of the fit.

ncm_stats_dist_set_print_fit

Sets NcmStatsDist:print-fit.

ncm_stats_dist_set_split_frac

Sets NcmStatsDist:split-frac, read by the split cross-validations. A fraction that leaves no out-of-sample point aborts at the next ncm_stats_dist_prepare().

ncm_stats_dist_set_uniform_weights

Sets NcmStatsDist:uniform-weights. Takes effect at the next preparation.

ncm_stats_dist_set_use_threads

Sets NcmStatsDist:use-threads.

Methods inherited from GObject (43)

Please see GObject for a full list of methods.

Properties

NumCosmoMath.StatsDist:CV-type

The NcmStatsDistCV that fits the bandwidth at ncm_stats_dist_prepare(). Default: #NCM_STATS_DIST_CV_NONE.

NumCosmoMath.StatsDist:N

Number of sample points held, ncm_stats_dist_get_sample_size().

NumCosmoMath.StatsDist:auto-kernel

Whether the kernel is chosen together with the over-smooth factor, by the same objective. It applies to every NcmStatsDist:CV-type except

NCM_STATS_DIST_CV_NONE, which fits nothing. The kernel must be the

NcmStatsDistKernelST the object was built with; any other kernel aborts at prepare. Its degrees of freedom $\nu$ are fitted in place, jointly with the over-smooth factor, over $\nu \in [\nu_\mathrm{min}, 10^4]$, with $\nu_\mathrm{min} = 2.5$ when NcmStatsDist:center-shrink is set and $1$ otherwise. At $\nu = 10^4$ the log-density differs from the Gaussian kernel’s by about $[(\chi^2 - d)^2 - 2d] / (4\nu)$: over the central 99.8% of $\chi^2$, from $2 \times 10^{-3}$ at $d = 1$ to $6 \times 10^{-2}$ at $d = 100$. Default: FALSE.

NumCosmoMath.StatsDist:center-shrink

Whether to shrink the kernel centers toward the sample mean so that the covariance of the kernel mixture matches the sample covariance for any bandwidth, see the class description. Default: FALSE.

NumCosmoMath.StatsDist:defensive-frac

Weight $\epsilon$ of a wide Student-t component mixed into the proposal, $q = (1 - \epsilon)\, q_\mathrm{mixture} + \epsilon\, t_\nu(\mu, c\,C)$, the Student-t density with location $\mu$ and scale matrix $c\,C$, where $\mu$ and $C$ are the sample mean and covariance, $c$ is NcmStatsDist:defensive-scale and $\nu$ is NcmStatsDist:defensive-nu. It bounds the proposal density from below where the kernels leave holes, so that a walker there can still move. Zero disables it and leaves every evaluation and draw unchanged. Default: 0.

NumCosmoMath.StatsDist:defensive-nu

Degrees of freedom $\nu$ of the wide Student-t component of NcmStatsDist:defensive-frac. Default: 3.

NumCosmoMath.StatsDist:defensive-scale

Factor $c$ multiplying the sample covariance in the scale matrix of the wide component of NcmStatsDist:defensive-frac. Default: 4.

NumCosmoMath.StatsDist:kernel

The NcmStatsDistKernel of the mixture.

NumCosmoMath.StatsDist:over-smooth

Factor multiplying the kernel rule-of-thumb bandwidth, ncm_stats_dist_kernel_get_rot_bandwidth(). A cross-validation other than

NCM_STATS_DIST_CV_NONE fits it in $[10^{-2}, 20]$, starting from the current

value. NcmStatsDistVKDE uses it as the bandwidth itself unless NcmStatsDistVKDE:use-rot-href is set. Default: 1.

NumCosmoMath.StatsDist:print-fit

Whether the bandwidth fit prints its progress to the standard output. Default: FALSE.

NumCosmoMath.StatsDist:split-frac

Fraction $f$ of the sample used as kernel centers by the split cross-validations,

NCM_STATS_DIST_CV_SPLIT_M2LNP and #NCM_STATS_DIST_CV_SPLIT_ACCEPT: the first

$\lceil f n \rceil$ points, in the order they were added, are the centers and the others score the bandwidth. Default: 0.5.

NumCosmoMath.StatsDist:uniform-weights

Whether ncm_stats_dist_prepare() keeps uniform kernel weights instead of fitting them by NNLS. The sample’s $-2\ln L$ is then used only by the bandwidth objectives that need it and by the outlier cut. Default: FALSE.

NumCosmoMath.StatsDist:use-threads

Whether NcmStatsDistVKDE runs its OpenMP loops in parallel; the other classes do not read it. Default: FALSE.

Signals

Signals inherited from GObject (1)
GObject::notify

The notify signal is emitted on an object when one of its properties has its value set through g_object_set_property(), g_object_set(), et al.

Class structure

struct NumCosmoMathStatsDistClass {
  void (* set_dim) (
    NcmStatsDist* sd,
    const guint dim
  );
  gdouble (* bandwidth) (
    NcmStatsDist* sd
  );
  void (* prepare_shapes) (
    NcmStatsDist* sd,
    GPtrArray* sample_array
  );
  void (* prepare_kernels) (
    NcmStatsDist* sd
  );
  void (* compute_IM) (
    NcmStatsDist* sd,
    NcmMatrix* IM
  );
  gdouble (* amise) (
    NcmStatsDist* sd
  );
  NcmMatrix* (* peek_cov_decomp) (
    NcmStatsDist* sd,
    guint i
  );
  NcmMatrix* (* peek_full_cov_decomp) (
    NcmStatsDist* sd
  );
  NcmMatrix* (* peek_full_cov) (
    NcmStatsDist* sd
  );
  gdouble (* get_lnnorm) (
    NcmStatsDist* sd,
    guint i
  );
  gdouble (* eval_weights) (
    NcmStatsDist* sd,
    NcmVector* weights,
    NcmVector* x
  );
  gdouble (* eval_weights_m2lnp) (
    NcmStatsDist* sd,
    NcmVector* weights,
    NcmVector* x
  );
  void (* eval_weights_m2lnp_vec) (
    NcmStatsDist* sd,
    NcmVector* weights,
    GPtrArray* x_a,
    NcmVector* m2lnp
  );
  void (* eval_weights_m2lnp_loo) (
    NcmStatsDist* sd,
    NcmVector* weights,
    GPtrArray* x_a,
    NcmVector* m2lnp
  );
  void (* reset) (
    NcmStatsDist* sd
  );
  
}

No description available.

Class members
set_dim: void (* set_dim) ( NcmStatsDist* sd, const guint dim )

No description available.

bandwidth: gdouble (* bandwidth) ( NcmStatsDist* sd )

No description available.

prepare_shapes: void (* prepare_shapes) ( NcmStatsDist* sd, GPtrArray* sample_array )

No description available.

prepare_kernels: void (* prepare_kernels) ( NcmStatsDist* sd )

No description available.

compute_IM: void (* compute_IM) ( NcmStatsDist* sd, NcmMatrix* IM )

No description available.

amise: gdouble (* amise) ( NcmStatsDist* sd )

No description available.

peek_cov_decomp: NcmMatrix* (* peek_cov_decomp) ( NcmStatsDist* sd, guint i )

No description available.

peek_full_cov_decomp: NcmMatrix* (* peek_full_cov_decomp) ( NcmStatsDist* sd )

No description available.

peek_full_cov: NcmMatrix* (* peek_full_cov) ( NcmStatsDist* sd )

No description available.

get_lnnorm: gdouble (* get_lnnorm) ( NcmStatsDist* sd, guint i )

No description available.

eval_weights: gdouble (* eval_weights) ( NcmStatsDist* sd, NcmVector* weights, NcmVector* x )

No description available.

eval_weights_m2lnp: gdouble (* eval_weights_m2lnp) ( NcmStatsDist* sd, NcmVector* weights, NcmVector* x )

No description available.

eval_weights_m2lnp_vec: void (* eval_weights_m2lnp_vec) ( NcmStatsDist* sd, NcmVector* weights, GPtrArray* x_a, NcmVector* m2lnp )

No description available.

eval_weights_m2lnp_loo: void (* eval_weights_m2lnp_loo) ( NcmStatsDist* sd, NcmVector* weights, GPtrArray* x_a, NcmVector* m2lnp )

No description available.

reset: void (* reset) ( NcmStatsDist* sd )

No description available.

Virtual methods

NumCosmoMath.StatsDistClass.amise
No description available.

NumCosmoMath.StatsDistClass.bandwidth
No description available.

NumCosmoMath.StatsDistClass.compute_IM
No description available.

NumCosmoMath.StatsDistClass.eval_weights
No description available.

NumCosmoMath.StatsDistClass.get_lnnorm

Gets the logarithm of the normalization $N_i$ of kernel i at the applied bandwidth, $K_i(x) = \bar{K}(\chi^2) / N_i$, see NcmStatsDistKernel.

NumCosmoMath.StatsDistClass.peek_cov_decomp

Gets the upper-triangular Cholesky factor $U_i$ of the scale matrix $\Sigma_i = U_i^T U_i$ of kernel i, as applied in the last preparation. The kernel covariance is $\kappa h^2 \Sigma_i$, with $h$ from ncm_stats_dist_get_href().

NumCosmoMath.StatsDistClass.peek_full_cov

Gets the subclass’s global scale matrix: for NcmStatsDistKDE the scale matrix shared by all kernels (the sample, fixed or robust covariance).

NumCosmoMath.StatsDistClass.peek_full_cov_decomp

Gets the upper-triangular Cholesky factor of the subclass’s global scale matrix, ncm_stats_dist_peek_full_cov(), as applied in the last preparation.

NumCosmoMath.StatsDistClass.prepare_shapes

Runs the first stage of a prepare alone: the per-kernel covariance structures the subclass builds from the sample, before any bandwidth is applied, over the kernel count of the last ncm_stats_dist_prepare(). For inspecting those structures; ncm_stats_dist_prepare() runs it as part of the whole pipeline, and only after that is the object ready for ncm_stats_dist_eval().

NumCosmoMath.StatsDistClass.reset

Discards all sample points added so far.

NumCosmoMath.StatsDistClass.set_dim
No description available.