Indicators and collections¶
Indicators are the main tool xclim provides. In contrast to the functions defined in xclim.compute, Indicators add a layer of health checks and metadata handling and return xarray.Dataset objects by default. Indicator objects are split into submodules according to their “realm” : atmos, land and seaIce, with two additional submodules : generic (for indicators that don’t apply to a specific variable) and convert (for non-resampling indicators that transform between variables).
Control indicator behaviour¶
- class xclim.core.options.set_options(**kwargs)[source]¶
Set options for xclim in a controlled context.
- Parameters:
metadata_locales (list[Any]) – List of IETF language tags or tuples of language tags and a translation dict, or tuples of language tags and a path to a json file defining translation of attributes. Default:
[].data_validation ({“log”, “raise”, “error”}) – Whether to “log”, “raise” an error or ‘warn’ the user on inputs that fail the data checks in
xclim.core.datachecks(). Default:"raise".cf_compliance ({“log”, “raise”, “error”}) – Whether to “log”, “raise” an error or “warn” the user on inputs that fail the CF compliance checks in
xclim.core.cfchecks(). Default:"warn".check_missing ({“any”, “wmo”, “pct”, “at_least_n”, “skip”}) – How to check for missing data and flag computed indicators. Available methods are “any”, “wmo”, “pct”, “at_least_n” and “skip”. Missing method can be registered through the xclim.core.options.register_missing_method decorator. Default:
"any"missing_options (dict) – Dictionary of options to pass to the missing method. Keys must the name of missing method and values must be mappings from option names to values.
run_length_ufunc (str) – Whether to use the 1D ufunc version of run length algorithms or the dask-ready broadcasting version. Default is
"auto", which means the latter is used for dask-backed and large arrays.as_dataset (bool) – If True, indicators output datasets. If False, they output DataArrays or tuple of DataArrays. The output dataset inherits attributes from the input dataset (if any) according to xarray’s
keep_attrsoption, which defaults to preserving attributes. Default :True.resample_map_blocks (bool) – If True, some indicators will wrap their resampling operations with xr.map_blocks, using
xclim.compute.helpers.resample_map(). This requires flox to be installed in order to ensure the chunking is appropriate.
Examples
You can use
set_optionseither as a context manager:>>> import xclim >>> ds = xr.open_dataset(path_to_tas_file).tas >>> with xclim.set_options(metadata_locales=["fr"]): ... out = xclim.atmos.tg_mean(ds)
Or to set global options:
import xclim xclim.set_options(missing_options={"pct": {"tolerance": 0.04}})
The Indicator object¶
The Indicator class wraps computations with pre- and post-processing functionality. Prior to computations, the class runs data and metadata health checks. After computations, the class masks values that should be considered missing and adds metadata attributes to the output.
There are many ways to construct indicators. A good place to start is this notebook.
- class xclim.core.indicator.Indicator(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]¶
Climate indicator base class.
Climate indicator object that, when called, computes an indicator and assigns its output a number of CF-compliant attributes. These attributes can be templated, allowing metadata to reflect the value of call arguments.
Instantiating a new indicator returns an instance but also registers is in
xclim.core.indicator.registry.Metadata and attributes in Indicator.outputs will be formatted and added to the output variable(s). This attribute is a list of
Outputobjects. An history attribute is also created an added to the output, merging with a parent’s dataset history if possible. Finally axclim_descriptionattribute is added to each output variable, summarizing the xclim call done.A lot of the Indicator’s metadata is parsed from the underlying compute function’s docstring and signature. Input variables and parameters are listed in
xclim.core.indicator.Indicator.parameters, while parameters that will be injected in the compute function are inxclim.core.indicator.Indicator.injected_parameters.Compared to their base compute function, indicators add the possibility of using a dataset or a
xarray.DataTreeas input, with the added argument ds in the call signature. All arguments that were indicated by the compute function to be variables (DataArrays) through annotations will be promoted to also accept strings that correspond to variable names in the ds Dataset (or on each DataTree nodes). Also, indicators return Datasets by default, while compute functions return one or multiple DataArrays.- __init__(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]¶
Create a new indicator.
- Parameters:
identifier (str) – Unique ID for this indicator. Single-output indicator will use this as their output variable name if no var_name`is passed to the first element of `attrs. Unless
registeris False, indicators are registered toxclim.core.indicator.registry, using this ID. The registry is case-insensitive. When defining indicators in a python module, it can be helpful to use the same name in the code as the identifier, to avoid confusion between the two, especially for collections and translations.compute (func) – The function computing the indicators. It should return one or more DataArray. Metadata will first be parsed from it as much as possible.
title (str, optional) – A succinct description of what is in the computed outputs. Parsed from compute docstring if None (first paragraph).
abstract (str, optional) – A long description of what is in the computed outputs. Parsed from compute docstring if None (second paragraph).
realm ({‘atmos’, ‘convert’, ‘seaIce’, ‘land’, ‘ocean’}, optional) – General domain of validity of the indicator.
keywords (list of strings, optional) – Keywords.
references (str, optional) – Published or web-based references that describe the data or methods used to produce it. Parsed from compute docstring if None (from the “References” section).
notes (str, optional) – Notes regarding computing function, for example the mathematical formulation. Parsed from compute docstring if None (form the “Notes” section).
input (dict, optional) – Mapping from input variable name in the compute function to known variable name. Useful for transforming generic compute functions into variable-specific indicator. The new variables names must be defined in
xclim.core.VARIABLES.parameters (dict, optional) – Overrides for the parameters. Either value to “inject”, removing that parameter from the call signature, or dictionaries of properties to override the ones parsed from the docstring. See
Parameterfor valid properties. Additionally, name can be passed to change the name of the argument in the call signature.outputs (list of dict or list of Output) – Metadata for the computation’s output : name, units and attributes. Any attribute are accepted, but giving a var_name is required for multi-output indicators. The list must be the same length as the number of outputs of the compute function.
context (str) – A pint unit context enabled during the computation of this indicator. For example use ‘hydro’ to allow conversion from ‘kg m-2 s-1’ to ‘mm/day’ for all inputs an outputs.
src_freq (str or sequence of str, optional) – The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.
register (bool) – If True (default), the indicator is registered into the
registrydictionary of indicators using its identifier as key.**outputs_kwargs – For convenience, output attributes and metadata can also be passed by name to the constructor.
- abstract: str¶
Description of the indicator.
- cfcheck(**das)¶
Compare metadata attributes to CF-Convention standards.
Default cfchecks use the specifications in xclim.core.VARIABLES, assuming the indicator’s inputs are using the CMIP6/xclim variable names correctly. Variables absent from these default specs are silently ignored.
When subclassing this method, use functions decorated using xclim.core.options.cfcheck.
- Parameters:
**das (dict) – A dictionary of DataArrays to check.
- Return type:
None
- static compute(*args, **kwargs)¶
The index-like compute function.
- context: str = 'none'¶
Name of xclim.units.units context which will be enabled during computation.
- classmethod copy(identifier=None, register=True, **kwargs)[source]¶
Create a new indicator by copying and modifying this indicator, similar to subclassing.
This accepts the same arguments as the indicator constructor, but parameters and attributes will default to this indicator’s data.
- Parameters:
identifier (str) – Unique ID for this indicator. Identifier are never inherited from their parent. This must be set if register is True.
register (bool) – Whether to register this new indicator in the
registry. Must be set to False is identifier is not given.**kwargs – All other arguments that
Indicator.__init__()accepts.
- Return type:
- Returns:
Indicator – A new indicator instance derived from this one.
See also
Indicator.__init__Indicator constructor.
- datacheck(**das)¶
Verify that input data is valid.
For example, checks could include: * assert no precipitation is negative * assert no temperature has the same value 5 days in a row
This base datacheck checks that the input data has a valid sampling frequency, as given in self.src_freq. If there are multiple inputs, it also checks if they all have the same frequency and the same anchor.
- Parameters:
**das (dict) – A dictionary of DataArrays to check.
- Raises:
if the frequency of any input can’t be inferred. - if inputs have different frequencies. - if inputs have a daily or hourly frequency, but they are not given at the same time of day.
- Return type:
None
- classmethod from_dict(data, identifier, module=None)¶
Deprecated method to create an indicator, please use
Indicator.copy()directly on the base indicator instead.- Return type:
- classmethod get_parent_ids()¶
Return the list of indicator identifiers this indicator was derived from.
- Returns:
list – All parent indicator classes of this indicator. Only classes defining an identifier are included.
- identifier: str = None¶
Unique ID identifying this indicator. Mostly for registry purposes.
- property injected_parameters: collections.abc.Mapping[str, Any]¶
Dictionary of all injected parameters (values).
- Returns:
dict – A dictionary of all injected parameters’ values.
- property is_generic: bool¶
If the indicator is “generic” returns True, meaning that it can accept variables with any units.
- Returns:
bool – True if the indicator is generic.
- json(args=None)¶
Return a serializable dictionary representation of the indicator.
- Parameters:
args (mapping, optional) – Arguments as passed to the call method of the indicator. If not given, the default arguments will be used when formatting the attributes.
- Return type:
dict- Returns:
dict – A dictionary representation of the indicator.
Notes
This is meant to be used by a third-party library wanting to wrap this indicator into another interface.
- keywords: tuple[str] = ()¶
Keywords describing the indicator and its domains of application. Child classes append to the list when inheriting.
- property n_outs: int¶
The number of outputs of this indicator.
- Returns:
int – The number of outputs.
- notes: str¶
Additional information about the indicator.
- outputs: list[xclim.core.indicator._indicator.Output]¶
List of output metadata.
- property parameters: collections.abc.Mapping[str, xclim.core.indicator._indicator.Parameter]¶
Dictionary of controllable (non-injected) parameters.
Similar to
IndexWrapper._all_parameters, but doesn’t include injected parameters.- Returns:
dict – A dictionary of controllable parameters.
- realm: str = None¶
General domain of validity of the indicator. Should use the same vocabulary as CMIP.
- references: str¶
rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.
- src_freq: str | list[str] | None = None¶
The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.
- title: str¶
Short description of the indicator.
- translate(locale, fill_missing=True)¶
Return a dictionary of metadata and (unformatted) output attributes for the requested locale.
- Return type:
dict
- class xclim.core.indicator.Parameter(kind, default, compute_name=<class 'xclim.core.indicator._indicator._empty'>, description='', units=<class 'xclim.core.indicator._indicator._empty'>, choices=<class 'xclim.core.indicator._indicator._empty'>, value=<class 'xclim.core.indicator._indicator._empty'>, annotation=<class 'xclim.core.indicator._indicator._empty'>)[source]¶
Object representing an indicator’s parameter.
For convenience, this class implements a special “contains”.
Examples
>>> p = Parameter(InputKind.NUMBER, default=2, description="A simple number") >>> p.units is Parameter._empty # has not been set True >>> "units" in p # Easier/retrocompatible way to test if units are set False >>> p.description 'A simple number'
- property injected: bool¶
Indicate whether values are injected.
- Returns:
bool – Whether values are injected.
- class xclim.core.indicator.Output(var_name=None, dimensionality=None, units=None, units_metadata=None, attrs=None, **attrs_kwargs)[source]¶
Metadata for the output of an indicator.
- attrs: dict¶
Output variable attributes.
- dimensionality: str | None¶
Dimensionality specification, similar but not necessarily compatible with pint.
- get(key, default=None)[source]¶
Convenience method to access any metadata element.
This method acts as if the
Outputobject was a single dictionary of all metadata elements and attributes.- Parameters:
key (str) – Name of the metadata element (or attribute) to return. If
keyis not one ofvar_name,dimensionality,unitsorunits_metadatait is searched inself.attrs.default (any) – If the key is not found, default value to return.
- Returns:
any – The corresponding value, or
defaultif the key isn’t found.
- property meta: dict¶
A dictionary of the non-attribute metadata for this output.
- Returns:
dict – The non-attribute metadata of this Output.
- units: str | None¶
Units of the output.
- units_metadata: str | None¶
Additional CF metadata for the units.
- var_name: str | None¶
Output variable name.
- class xclim.core.indicator.ReducingIndicator(**kwargs)[source]¶
Indicator that performs a time-reducing computation.
- class xclim.core.indicator.IndexingIndicator(identifier=None, compute=None, title=None, abstract=None, realm=None, keywords=None, references=None, notes=None, input=None, parameters=None, outputs=None, context='none', src_freq=None, register=True, **outputs_kwargs)[source]¶
Indicator that also adds the “indexer” kwargs to subset the inputs before computation.
- class xclim.core.indicator.ResamplingIndicator(**kwargs)[source]¶
Indicator that performs a resampling computation.
Compared to the base Indicator, this adds the handling of missing data, and the check of allowed periods.
- abstract: str¶
Description of the indicator.
- allowed_periods: list[str] = None¶
A list of allowed periods, i.e. base parts of the freq parameter. For example, indicators meant to be computed annually only will have allowed_periods=[“Y”]. None means “any period” or that the indicator doesn’t take a freq argument.
- notes: str¶
Additional information about the indicator.
- outputs: list[xclim.core.indicator._indicator.Output]¶
List of output metadata.
- references: str¶
rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.
- title: str¶
Short description of the indicator.
- class xclim.core.indicator.ResamplingIndicatorWithIndexing(**kwargs)[source]¶
Resampling indicator that also adds “indexer” kwargs to subset the inputs before computation.
- class xclim.core.indicator.Hourly(**kwargs)[source]¶
Class for hourly inputs and resampling computes.
- abstract: str¶
Description of the indicator.
- notes: str¶
Additional information about the indicator.
- outputs: list[xclim.core.indicator._indicator.Output]¶
List of output metadata.
- references: str¶
rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.
- src_freq: str | list[str] | None = 'h'¶
The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.
- title: str¶
Short description of the indicator.
- class xclim.core.indicator.Daily(**kwargs)[source]¶
Class for daily inputs and resampling computes.
- abstract: str¶
Description of the indicator.
- notes: str¶
Additional information about the indicator.
- outputs: list[xclim.core.indicator._indicator.Output]¶
List of output metadata.
- references: str¶
rst cite directives for literature about this indicator. Child classes append their references as a new line when inheriting.
- src_freq: str | list[str] | None = 'D'¶
The expected frequency of the input data. Can be a list for multiple frequencies, or None if irrelevant.
- title: str¶
Short description of the indicator.
Indicator collections¶
An indicator collection is a structure holding multiple indicators. It can be created through a yaml configuration file.
YAML file structure¶
Indicator-defining yaml files are structured in the following way. Most entries of the indicators section are
mirroring attributes of the xclim.core.indicator.Indicator, please refer to its documentation for more
details on each.
module: <module name> # Defaults to the file name
realm: <realm> # If given here, applies to all indicators that do not already provide it.
keywords:
- <keyword> # Merged with indicator-specific keywords (appended to the list)
references: <references> # Merged with indicator-specific references (joined with a new line)
base: <base indicator class> # Defaults to "Daily" and applies to all indicators that do not give it.
doc: <module docstring> # Defaults to a minimal header, only valid if the module doesn't already exist.
variables: # Optional section if indicators declared below rely on variables unknown to xclim
# (not in `xclim.core.VARIABLES`)
# The variables are not module-dependent and will overwrite any already existing with the same name.
<varname>:
canonical_units: <units> # required
description: <description> # required
standard_name: <expected standard_name> # optional
cell_methods: <expected cell_methods> # optional
# The `bases` and `indicators` sections have the same syntax. Indicators defined in the `bases` section
# will only be created as classes and not instances. They will not be included in the IndicatorCollection's items,
# but rather in its `bases` property. This is useful for creating a base from which multiple indicators are declared
# in the `indicators`section.
bases:
indicators:
<identifier>: # The actual indicator identifier will be prepended by the module name.
# From which Indicator to inherit
base: <base indicator class> # Defaults to module-wide base class
# See :ref:`Base class specification` below.
# General metadata, usually parsed from the `compute`s docstring when possible.
realm: <realm> # defaults to module-wide realm. One of "atmos", "land", "seaIce", "ocean".
title: <title>
abstract: <abstract>
keywords:
- <keyword> # merged to module-wide keywords.
references: <references> # newline-seperated, merged to module-wide references.
notes: <notes>
# Other options (not all indicator classes support them)
missing: <missing method name>
missing_options: <missing options mapping>
allowed_periods: [<list>, <of>, <allowed>, <periods>]
context: <context> # A unit context enabled during the conversion of the compute's output to the requested units
# Compute function
compute: <function name> # See :ref:`Compute function specification`, below.
input: # When "compute" is a generic function, this is a mapping from argument name to the expected variable.
# It will change the expected name of the variable as well as its units/dimensionality.
# Can refer to a variable declared in the `variables` section above or in `xclim.core.VARIABLES`.
# See also :ref:`Inputs` below.
<var name in compute> : <variable official name>
...
# Parameters
parameters:
<param name>: <param data> # Simplest case, to inject parameters in the compute function.
# Kwargs-like parameters like ``indexer`` must be injected as a dictionary here.
<param name>: # To change parameters metadata or to declare units when "compute" is a generic function.
default: <param default>
description: <param description>
name: <param name> # Change the name of the parameter (similar to what `input` does for variables)
kind: <param kind> # Override the parameter kind. This is mostly useful for transforming an
# optional variable into a required one by passing ``kind: 0``.
...
# Output metadata
outputs: # List of mappings, one for each output.
- var_name: <var name> # Name to give to the output
units: <units> # Units to convert the output to and assign as attribute
attrs: # Mapping of attributes to assign to the output, can be templated strings.
long_name: <...>
description: <...>
- ...
... # and so on.
All fields are optional. Other fields found in the yaml file will trigger errors when validation is activated.
When a module is built from a yaml file, the yaml is first validated against the schema (see xclim/data/schema.yml) using the YAMALE library ([Lopker, 2022]). See the “Extending xclim” notebook for more info.
Base class specification¶
There are multiple ways to specify a base class when defining an indicator. In priority order:
- If
basestarts with a ‘.’ (ex: .RXXp), the base class is taken from the current module. It is first searched in bases section.
If not found, it is searched as another indicator declared _above_ the current definition.
- If
The name is searched in the base class registry,
xclim.core.indicator.base_registry(example:Daily).The name is searched in the indicator registry,
xclim.core.indicator.registry(example:prcptot).- If
basecontains a ‘.’ : If the first element is one of xclim’s indicators submodules (ex:
atmos.precip_accumulation), that indicator is used as a base. Any submodule ofxclim.indicatorsare possible.Otherwise, that path is loaded with python’s normal import mechanism. (ex:
mymodule.submod.MyIndicatorisloaded asfrom mymodule.submod import MyIndicator).
- If
If the field base is not given, it defaults to the module-wide base, which itself defaults to
xclim.core.indicator.Daily`.
Compute function specification¶
Similar to the base field, there are multiple ways to refer to a compute function in the compute field.
In priority order:
If a module or mapping of compute functions was passed to
IndicatorCollection.from_yaml(), the name is searched there (ex: extreme_precip_accumulation_and_days).The name is searched in
xclim.compute.generic(ex:statistics).The name is searched in :py:mod:`xclim.compute (ex:
corn_heat_units). It may contain a ‘.’ to denote a submodule (ex:generic.statistics).Otherwise, it is loaded with python’s normal import mechanism (ex:
mymodule.submod.my_functionis loaded asfrom mymodule.submod import my_function).
Inputs¶
As xclim has strict definitions of possible input variables (see xclim.core.VARIABLES),
the mapping of indicators.<identifier>.input simply links an argument name from the function given in “compute”
to one of those official variables.
- class xclim.core.collection.IndicatorCollection(indicators, name=None, bases=None, doc=None)[source]¶
A collection of indicators.
- classmethod from_yaml(filename, name=None, computes=None, translations=None, mode='raise', encoding='UTF8', validate=True, register=False)[source]¶
Build an indicator collection from a YAML file.
When given only a base filename (no ‘yml’ extension), this tries to find custom indicators in a module of the same name (.py) and translations in json files (.<lang>.json), see Notes.
Indicator created here will have the name of the module prepended to their identifier (ex: {mod}.{baseId}). The base identifier being the key name within the indicators mapping in the yaml.
- Parameters:
filename (PathLike) – Path to a YAML file or to the stem of all module files. See Notes for behaviour when passing a basename only.
name (str, optional) – The name of the new or existing module, defaults to the basename of the file (e.g: atmos.yml -> atmos).
computes (Mapping of callables or module or path, optional) – A mapping or module of compute functions or a python file declaring such a module. When creating the indicator, the name in the compute field is first sought here, then the indicator class will search in
xclim.compute.genericand finally inxclim.compute.translations (Mapping of dicts or path, optional) – Translated metadata for the new indicators. Keys of the mapping must be two-character language tags. Values can be translations dictionaries as defined in
xclim.core.locales. They can also be a path to a JSON file defining the translations.mode ({‘raise’, ‘warn’, ‘ignore’}) – How to deal with broken indicator definitions.
encoding (str) – The encoding used to open the .yaml and .json files. It defaults to UTF-8, overriding python’s mechanism which is machine dependent.
validate (bool or PathLike) – If True (default), the yaml module is validated against the xclim schema. Can also be the path to a YAML schema against which to validate; Or False, in which case validation is simply skipped.
register (bool) – If True, the indicators created here are registered in xclim’s indicators registry
registryupon creation, using the collection’s name prepended to their identifier as key, as explained above. Defaults to False, making collections independent from xclim’s registry. This does not change the behaviour of registering new variables, which are always added to xclim’s centralxclim.core.VARIABLES.
- Returns:
IndicatorCollection – A collection of indicators.
See also
xclim.core.indicator.IndicatorIndicator build logic.
Notes
When the given filename has no suffix (usually ‘.yaml’ or ‘.yml’), the function will try to load custom compute functions definitions from a file with the same name but with a .py extension. Similarly, it will try to load translations in *.<lang>.json files, where <lang> is the IETF language tag. Note that the file name can not contain a dot (
.) for this logic to work.For example. a set of custom indicators could be fully described by the following files:
example.yml : defining the indicator’s metadata.
example.py : defining a few compute functions.
example.fr.json : French translations
- xclim.core.indicator.registry¶
Registry of indicator instances.
This case-insensitive dictionary holds all the indicators defined in xclim, mapped by their identifier. Indicators are added here by default, but it can be avoided by passing register=False to the constructor. Indicators here can be referred to in the base field of the YAML definitions when creating a
IndicatorCollection.
- xclim.core.indicator.base_registry¶
Registry of standard base indicator classes.
This dictionary holds some useful base indicator classes to construct new indicators. It is mostly useful when defining indicators within a
IndicatorCollection, keys of this dictionary can be values of the base field of the YAML definitions.