Indicators and collections

Indicators are the main tool xclim provides. In contrast to the functions defined in xclim.compute, Indicators add a layer of health checks and metadata handling and return xarray.Dataset objects by default. Indicator objects are split into submodules according to their “realm” : atmos, land and seaIce, with two additional submodules : generic (for indicators that don’t apply to a specific variable) and convert (for non-resampling indicators that transform between variables).

Control indicator behaviour

class xclim.core.options.set_options(**kwargs)[source]

Set options for xclim in a controlled context.

Parameters:
  • metadata_locales (list[Any]) – List of IETF language tags or tuples of language tags and a translation dict, or tuples of language tags and a path to a json file defining translation of attributes. Default: [].

  • data_validation ({“log”, “raise”, “error”}) – Whether to “log”, “raise” an error or ‘warn’ the user on inputs that fail the data checks in xclim.core.datachecks(). Default: "raise".

  • cf_compliance ({“log”, “raise”, “error”}) – Whether to “log”, “raise” an error or “warn” the user on inputs that fail the CF compliance checks in xclim.core.cfchecks(). Default: "warn".

  • check_missing ({“any”, “wmo”, “pct”, “at_least_n”, “skip”}) – How to check for missing data and flag computed indicators. Available methods are “any”, “wmo”, “pct”, “at_least_n” and “skip”. Missing method can be registered through the xclim.core.options.register_missing_method decorator. Default: "any"

  • missing_options (dict) – Dictionary of options to pass to the missing method. Keys must the name of missing method and values must be mappings from option names to values.

  • run_length_ufunc (str) – Whether to use the 1D ufunc version of run length algorithms or the dask-ready broadcasting version. Default is "auto", which means the latter is used for dask-backed and large arrays.

  • as_dataset (bool) – If True, indicators output datasets. If False, they output DataArrays or tuple of DataArrays. The output dataset inherits attributes from the input dataset (if any) according to xarray’s keep_attrs option, which defaults to preserving attributes. Default :True.

  • resample_map_blocks (bool) – If True, some indicators will wrap their resampling operations with xr.map_blocks, using xclim.compute.helpers.resample_map(). This requires flox to be installed in order to ensure the chunking is appropriate.

Examples

You can use set_options either as a context manager:

>>> import xclim
>>> ds = xr.open_dataset(path_to_tas_file).tas
>>> with xclim.set_options(metadata_locales=["fr"]):
...     out = xclim.atmos.tg_mean(ds)

Or to set global options:

import xclim

xclim.set_options(missing_options={"pct": {"tolerance": 0.04}})

The Indicator object

The Indicator class wraps computations with pre- and post-processing functionality. Prior to computations, the class runs data and metadata health checks. After computations, the class masks values that should be considered missing and adds metadata attributes to the output.

There are many ways to construct indicators. A good place to start is this notebook.

class xclim.core.indicator.Indicator(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]

Climate indicator base class.

Climate indicator object that, when called, computes an indicator and assigns its output a number of CF-compliant attributes. These attributes can be templated, allowing metadata to reflect the value of call arguments.

Instantiating a new indicator returns an instance but also registers is in xclim.core.indicator.registry.

Metadata and attributes in Indicator.outputs will be formatted and added to the output variable(s). This attribute is a list of Output objects. An history attribute is also created an added to the output, merging with a parent’s dataset history if possible. Finally a xclim_description attribute is added to each output variable, summarizing the xclim call done.

A lot of the Indicator’s metadata is parsed from the underlying compute function’s docstring and signature. Input variables and parameters are listed in xclim.core.indicator.Indicator.parameters, while parameters that will be injected in the compute function are in xclim.core.indicator.Indicator.injected_parameters.

Compared to their base compute function, indicators add the possibility of using a dataset or a xarray.DataTree as input, with the added argument ds in the call signature. All arguments that were indicated by the compute function to be variables (DataArrays) through annotations will be promoted to also accept strings that correspond to variable names in the ds Dataset (or on each DataTree nodes). Also, indicators return Datasets by default, while compute functions return one or multiple DataArrays.

__init__(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]

Create a new indicator.

Parameters:
  • identifier (str) – Unique ID for this indicator. Single-output indicator will use this as their output variable name if no var_name`is passed to the first element of `attrs. Unless register is False, indicators are registered to xclim.core.indicator.registry, using this ID. The registry is case-insensitive. When defining indicators in a python module, it can be helpful to use the same name in the code as the identifier, to avoid confusion between the two, especially for collections and translations.

  • compute (func) – The function computing the indicators. It should return one or more DataArray. Metadata will first be parsed from it as much as possible.

  • title (str, optional) – A succinct description of what is in the computed outputs. Parsed from compute docstring if None (first paragraph).

  • abstract (str, optional) – A long description of what is in the computed outputs. Parsed from compute docstring if None (second paragraph).

  • realm ({‘atmos’, ‘convert’, ‘seaIce’, ‘land’, ‘ocean’}, optional) – General domain of validity of the indicator.

  • keywords (list of strings, optional) – Keywords.

  • references (str, optional) – Published or web-based references that describe the data or methods used to produce it. Parsed from compute docstring if None (from the “References” section).

  • notes (str, optional) – Notes regarding computing function, for example the mathematical formulation. Parsed from compute docstring if None (form the “Notes” section).

  • input (dict, optional) – Mapping from input variable name in the compute function to known variable name. Useful for transforming generic compute functions into variable-specific indicator. The new variables names must be defined in xclim.core.VARIABLES.

  • parameters (dict, optional) – Overrides for the parameters. Either value to “inject”, removing that parameter from the call signature, or dictionaries of properties to override the ones parsed from the docstring. See Parameter for valid properties. Additionally, name can be passed to change the name of the argument in the call signature.

  • outputs (list of dict or list of Output) – Metadata for the computation’s output : name, units and attributes. Any attribute are accepted, but giving a var_name is required for multi-output indicators. The list must be the same length as the number of outputs of the compute function.

  • context (str) – A pint unit context enabled during the computation of this indicator. For example use ‘hydro’ to allow conversion from ‘kg m-2 s-1’ to ‘mm/day’ for all inputs an outputs.

  • src_freq (str or sequence of str, optional) – The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.

  • register (bool) – If True (default), the indicator is registered into the registry dictionary of indicators using its identifier as key.

  • **outputs_kwargs – For convenience, output attributes and metadata can also be passed by name to the constructor.

abstract: str

Description of the indicator.

cfcheck(**das)

Compare metadata attributes to CF-Convention standards.

Default cfchecks use the specifications in xclim.core.VARIABLES, assuming the indicator’s inputs are using the CMIP6/xclim variable names correctly. Variables absent from these default specs are silently ignored.

When subclassing this method, use functions decorated using xclim.core.options.cfcheck.

Parameters:

**das (dict) – A dictionary of DataArrays to check.

Return type:

None

static compute(*args, **kwargs)

The index-like compute function.

context: str = 'none'

Name of xclim.units.units context which will be enabled during computation.

classmethod copy(identifier=None, register=True, **kwargs)[source]

Create a new indicator by copying and modifying this indicator, similar to subclassing.

This accepts the same arguments as the indicator constructor, but parameters and attributes will default to this indicator’s data.

Parameters:
  • identifier (str) – Unique ID for this indicator. Identifier are never inherited from their parent. This must be set if register is True.

  • register (bool) – Whether to register this new indicator in the registry. Must be set to False is identifier is not given.

  • **kwargs – All other arguments that Indicator.__init__() accepts.

Return type:

Indicator

Returns:

Indicator – A new indicator instance derived from this one.

See also

Indicator.__init__

Indicator constructor.

datacheck(**das)

Verify that input data is valid.

For example, checks could include: * assert no precipitation is negative * assert no temperature has the same value 5 days in a row

This base datacheck checks that the input data has a valid sampling frequency, as given in self.src_freq. If there are multiple inputs, it also checks if they all have the same frequency and the same anchor.

Parameters:

**das (dict) – A dictionary of DataArrays to check.

Raises:

ValidationError –

  • if the frequency of any input can’t be inferred. - if inputs have different frequencies. - if inputs have a daily or hourly frequency, but they are not given at the same time of day.

Return type:

None

classmethod from_dict(data, identifier, module=None)

Deprecated method to create an indicator, please use Indicator.copy() directly on the base indicator instead.

Return type:

Indicator

classmethod get_parent_ids()

Return the list of indicator identifiers this indicator was derived from.

Returns:

list – All parent indicator classes of this indicator. Only classes defining an identifier are included.

identifier: str = None

Unique ID identifying this indicator. Mostly for registry purposes.

property injected_parameters: collections.abc.Mapping[str, Any]

Dictionary of all injected parameters (values).

Returns:

dict – A dictionary of all injected parameters’ values.

property is_generic: bool

If the indicator is “generic” returns True, meaning that it can accept variables with any units.

Returns:

bool – True if the indicator is generic.

json(args=None)

Return a serializable dictionary representation of the indicator.

Parameters:

args (mapping, optional) – Arguments as passed to the call method of the indicator. If not given, the default arguments will be used when formatting the attributes.

Return type:

dict

Returns:

dict – A dictionary representation of the indicator.

Notes

This is meant to be used by a third-party library wanting to wrap this indicator into another interface.

keywords: tuple[str] = ()

Keywords describing the indicator and its domains of application. Child classes append to the list when inheriting.

property n_outs: int

The number of outputs of this indicator.

Returns:

int – The number of outputs.

notes: str

Additional information about the indicator.

outputs: list[xclim.core.indicator._indicator.Output]

List of output metadata.

property parameters: collections.abc.Mapping[str, xclim.core.indicator._indicator.Parameter]

Dictionary of controllable (non-injected) parameters.

Similar to IndexWrapper._all_parameters, but doesn’t include injected parameters.

Returns:

dict – A dictionary of controllable parameters.

realm: str = None

General domain of validity of the indicator. Should use the same vocabulary as CMIP.

references: str

rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.

src_freq: str | list[str] | None = None

The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.

title: str

Short description of the indicator.

translate(locale, fill_missing=True)

Return a dictionary of metadata and (unformatted) output attributes for the requested locale.

Return type:

dict

class xclim.core.indicator.Parameter(kind, default, compute_name=<class 'xclim.core.indicator._indicator._empty'>, description='', units=<class 'xclim.core.indicator._indicator._empty'>, choices=<class 'xclim.core.indicator._indicator._empty'>, value=<class 'xclim.core.indicator._indicator._empty'>, annotation=<class 'xclim.core.indicator._indicator._empty'>)[source]

Object representing an indicator’s parameter.

For convenience, this class implements a special “contains”.

Examples

>>> p = Parameter(InputKind.NUMBER, default=2, description="A simple number")
>>> p.units is Parameter._empty  # has not been set
True
>>> "units" in p  # Easier/retrocompatible way to test if units are set
False
>>> p.description
'A simple number'
property injected: bool

Indicate whether values are injected.

Returns:

bool – Whether values are injected.

json()[source]

Return a json-serializable dictionary of the Parameter.

Return type:

dict

Returns:

dict – Dictionary representation of the object, ready for serialization into json.

update(other)[source]

Update a parameter’s values from a dict.

Parameters:

other (dict) – A dictionary of parameters to update the current.

Return type:

None

class xclim.core.indicator.Output(var_name=None, dimensionality=None, units=None, units_metadata=None, attrs=None, **attrs_kwargs)[source]

Metadata for the output of an indicator.

attrs: dict

Output variable attributes.

dimensionality: str | None

Dimensionality specification, similar but not necessarily compatible with pint.

get(key, default=None)[source]

Convenience method to access any metadata element.

This method acts as if the Output object was a single dictionary of all metadata elements and attributes.

Parameters:
  • key (str) – Name of the metadata element (or attribute) to return. If key is not one of var_name, dimensionality, units or units_metadata it is searched in self.attrs.

  • default (any) – If the key is not found, default value to return.

Returns:

any – The corresponding value, or default if the key isn’t found.

property meta: dict

A dictionary of the non-attribute metadata for this output.

Returns:

dict – The non-attribute metadata of this Output.

units: str | None

Units of the output.

units_metadata: str | None

Additional CF metadata for the units.

var_name: str | None

Output variable name.

class xclim.core.indicator.ReducingIndicator(**kwargs)[source]

Indicator that performs a time-reducing computation.

class xclim.core.indicator.IndexingIndicator(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]

Indicator that also adds the “indexer” kwargs to subset the inputs before computation.

class xclim.core.indicator.ResamplingIndicator(**kwargs)[source]

Indicator that performs a resampling computation.

Compared to the base Indicator, this adds the handling of missing data, and the check of allowed periods.

abstract: str

Description of the indicator.

allowed_periods: list[str] = None

A list of allowed periods, i.e. base parts of the freq parameter. For example, indicators meant to be computed annually only will have allowed_periods=[“Y”]. None means “any period” or that the indicator doesn’t take a freq argument.

notes: str

Additional information about the indicator.

outputs: list[xclim.core.indicator._indicator.Output]

List of output metadata.

references: str

rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.

title: str

Short description of the indicator.

class xclim.core.indicator.ResamplingIndicatorWithIndexing(**kwargs)[source]

Resampling indicator that also adds “indexer” kwargs to subset the inputs before computation.

class xclim.core.indicator.Hourly(**kwargs)[source]

Class for hourly inputs and resampling computes.

abstract: str

Description of the indicator.

notes: str

Additional information about the indicator.

outputs: list[xclim.core.indicator._indicator.Output]

List of output metadata.

references: str

rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.

src_freq: str | list[str] | None = 'h'

The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.

title: str

Short description of the indicator.

class xclim.core.indicator.Daily(**kwargs)[source]

Class for daily inputs and resampling computes.

abstract: str

Description of the indicator.

notes: str

Additional information about the indicator.

outputs: list[xclim.core.indicator._indicator.Output]

List of output metadata.

references: str

rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.

src_freq: str | list[str] | None = 'D'

The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.

title: str

Short description of the indicator.

Indicator collections

An indicator collection is a structure holding multiple indicators. It can be created through a yaml configuration file.

YAML file structure

Indicator-defining yaml files are structured in the following way. Most entries of the indicators section are mirroring attributes of the xclim.core.indicator.Indicator, please refer to its documentation for more details on each.

module: <module name>  # Defaults to the file name
realm: <realm>  # If given here, applies to all indicators that do not already provide it.
keywords:
  - <keyword>  # Merged with indicator-specific keywords (appended to the list)
references: <references> # Merged with indicator-specific references (joined with a new line)
base: <base indicator class>  # Defaults to "Daily" and applies to all indicators that do not give it.
doc: <module docstring>  # Defaults to a minimal header, only valid if the module doesn't already exist.
variables:  # Optional section if indicators declared below rely on variables unknown to xclim
            # (not in `xclim.core.VARIABLES`)
            # The variables are not module-dependent and will overwrite any already existing with the same name.
  <varname>:
    canonical_units: <units> # required
    description: <description> # required
    standard_name: <expected standard_name> # optional
    cell_methods: <expected cell_methods> # optional
# The `bases` and `indicators` sections have the same syntax. Indicators defined in the `bases` section
# will only be created as classes and not instances. They will not be included in the IndicatorCollection's items,
# but rather in its `bases` property. This is useful for creating a base from which multiple indicators are declared
# in the `indicators`section.
bases:
indicators:
  <identifier>:  # The actual indicator identifier will be prepended by the module name.
    # From which Indicator to inherit
    base: <base indicator class>  # Defaults to module-wide base class
                                  # See :ref:`Base class specification` below.

    # General metadata, usually parsed from the `compute`s docstring when possible.
    realm: <realm>  # defaults to module-wide realm. One of "atmos", "land", "seaIce", "ocean".
    title: <title>
    abstract: <abstract>
    keywords:
      - <keyword>  # merged to module-wide keywords.
    references: <references>  # newline-seperated, merged to module-wide references.
    notes: <notes>

    # Other options (not all indicator classes support them)
    missing: <missing method name>
    missing_options: <missing options mapping>
    allowed_periods: [<list>, <of>, <allowed>, <periods>]
    context: <context> # A unit context enabled during the conversion of the compute's output to the requested units

    # Compute function
    compute: <function name>  # See :ref:`Compute function specification`, below.

    input:  # When "compute" is a generic function, this is a mapping from argument name to the expected variable.
            # It will change the expected name of the variable as well as its units/dimensionality.
            # Can refer to a variable declared in the `variables` section above or in `xclim.core.VARIABLES`.
            # See also :ref:`Inputs` below.
      <var name in compute> : <variable official name>
      ...

    # Parameters
    parameters:
      <param name>: <param data>  # Simplest case, to inject parameters in the compute function.
                                  # Kwargs-like parameters like ``indexer`` must be injected as a dictionary here.
      <param name>:  # To change parameters metadata or to declare units when "compute" is a generic function.
        default: <param default>
        description: <param description>
        name: <param name>  # Change the name of the parameter (similar to what `input` does for variables)
        kind: <param kind> # Override the parameter kind. This is mostly useful for transforming an
                           # optional variable into a required one by passing ``kind: 0``.
      ...

    # Output metadata
    outputs:   # List of mappings, one for each output.
        - var_name: <var name>  # Name to give to the output
          units: <units>        # Units to convert the output to and assign as attribute
          attrs:                # Mapping of attributes to assign to the output, can be templated strings.
              long_name: <...>
              description: <...>
        - ...

  ...  # and so on.

All fields are optional. Other fields found in the yaml file will trigger errors when validation is activated.

When a module is built from a yaml file, the yaml is first validated against the schema (see xclim/data/schema.yml) using the YAMALE library ([Lopker, 2022]). See the “Extending xclim” notebook for more info.

Base class specification

There are multiple ways to specify a base class when defining an indicator. In priority order:

  • If base starts with a ‘.’ (ex: .RXXp), the base class is taken from the current module.
    • It is first searched in bases section.

    • If not found, it is searched as another indicator declared _above_ the current definition.

  • The name is searched in the base class registry, xclim.core.indicator.base_registry (example: Daily).

  • The name is searched in the indicator registry, xclim.core.indicator.registry (example: prcptot).

  • If base contains a ‘.’ :
    • If the first element is one of xclim’s indicators submodules (ex: atmos.precip_accumulation), that indicator is used as a base. Any submodule of xclim.indicators are possible.

    • Otherwise, that path is loaded with python’s normal import mechanism. (ex: mymodule.submod.MyIndicator isloaded as from mymodule.submod import MyIndicator).

If the field base is not given, it defaults to the module-wide base, which itself defaults to xclim.core.indicator.Daily`.

Compute function specification

Similar to the base field, there are multiple ways to refer to a compute function in the compute field. In priority order:

  • If a module or mapping of compute functions was passed to IndicatorCollection.from_yaml(), the name is searched there (ex: extreme_precip_accumulation_and_days).

  • The name is searched in xclim.compute.generic (ex: statistics).

  • The name is searched in :py:mod:`xclim.compute (ex: corn_heat_units). It may contain a ‘.’ to denote a submodule (ex: generic.statistics).

  • Otherwise, it is loaded with python’s normal import mechanism (ex: mymodule.submod.my_function is loaded as from mymodule.submod import my_function).

Inputs

As xclim has strict definitions of possible input variables (see xclim.core.VARIABLES), the mapping of indicators.<identifier>.input simply links an argument name from the function given in “compute” to one of those official variables.

class xclim.core.collection.IndicatorCollection(indicators, name=None, bases=None, doc=None)[source]

A collection of indicators.

classmethod from_yaml(filename, name=None, computes=None, translations=None, mode='raise', encoding='UTF8', validate=True, register=False)[source]

Build an indicator collection from a YAML file.

When given only a base filename (no ‘yml’ extension), this tries to find custom indicators in a module of the same name (.py) and translations in json files (.<lang>.json), see Notes.

Indicator created here will have the name of the module prepended to their identifier (ex: {mod}.{baseId}). The base identifier being the key name within the indicators mapping in the yaml.

Parameters:
  • filename (PathLike) – Path to a YAML file or to the stem of all module files. See Notes for behaviour when passing a basename only.

  • name (str, optional) – The name of the new or existing module, defaults to the basename of the file (e.g: atmos.yml -> atmos).

  • computes (Mapping of callables or module or path, optional) – A mapping or module of compute functions or a python file declaring such a module. When creating the indicator, the name in the compute field is first sought here, then the indicator class will search in xclim.compute.generic and finally in xclim.compute.

  • translations (Mapping of dicts or path, optional) – Translated metadata for the new indicators. Keys of the mapping must be two-character language tags. Values can be translations dictionaries as defined in xclim.core.locales. They can also be a path to a JSON file defining the translations.

  • mode ({‘raise’, ‘warn’, ‘ignore’}) – How to deal with broken indicator definitions.

  • encoding (str) – The encoding used to open the .yaml and .json files. It defaults to UTF-8, overriding python’s mechanism which is machine dependent.

  • validate (bool or PathLike) – If True (default), the yaml module is validated against the xclim schema. Can also be the path to a YAML schema against which to validate; Or False, in which case validation is simply skipped.

  • register (bool) – If True, the indicators created here are registered in xclim’s indicators registry registry upon creation, using the collection’s name prepended to their identifier as key, as explained above. Defaults to False, making collections independent from xclim’s registry. This does not change the behaviour of registering new variables, which are always added to xclim’s central xclim.core.VARIABLES.

Returns:

IndicatorCollection – A collection of indicators.

See also

xclim.core.indicator.Indicator

Indicator build logic.

Notes

When the given filename has no suffix (usually ‘.yaml’ or ‘.yml’), the function will try to load custom compute functions definitions from a file with the same name but with a .py extension. Similarly, it will try to load translations in *.<lang>.json files, where <lang> is the IETF language tag. Note that the file name can not contain a dot (.) for this logic to work.

For example. a set of custom indicators could be fully described by the following files:

  • example.yml : defining the indicator’s metadata.

  • example.py : defining a few compute functions.

  • example.fr.json : French translations

iter_indicators()[source]

Iterate over the (name, indicator) pairs in this collection.

xclim.core.indicator.registry

Registry of indicator instances.

This case-insensitive dictionary holds all the indicators defined in xclim, mapped by their identifier. Indicators are added here by default, but it can be avoided by passing register=False to the constructor. Indicators here can be referred to in the base field of the YAML definitions when creating a IndicatorCollection.

xclim.core.indicator.base_registry

Registry of standard base indicator classes.

This dictionary holds some useful base indicator classes to construct new indicators. It is mostly useful when defining indicators within a IndicatorCollection, keys of this dictionary can be values of the base field of the YAML definitions.