Units and calendar tools¶
xclim implements multiples methods that can be useful to handle climate data with xarray.
Units Handling Submodule¶
xclim’s pint-based unit registry is an extension of the registry defined in cf-xarray. This module defines most unit handling methods.
- xclim.core.units.amount2lwethickness(amount, out_units=None)[source]¶
Convert a liquid water amount (mass over area) to its equivalent area-averaged thickness (length).
This will simply divide the amount by the density of liquid water, 1000 kg/m³. This is equivalent to using the “hydro” context of
xclim.core.units.units.- Parameters:
amount (xr.DataArray) – A DataArray storing a liquid water amount quantity.
out_units (str, optional) – Specific output units, if needed.
- Return type:
Union[DataArray,TypeVar(Quantified,DataArray,str,Quantity)]- Returns:
xr.DataArray or Quantified – The standard_name of amount is modified if a conversion is found (see
xclim.core.units.cf_conversion()), it is removed otherwise. Other attributes are left untouched.
See also
lwethickness2amountConvert a liquid water equivalent thickness to an amount.
- xclim.core.units.amount2rate(amount, dim='time', sampling_rate_from_coord=False, out_units=None)[source]¶
Convert an amount variable to a rate by dividing by the sampling period length.
If the sampling period length cannot be inferred, the amount values are divided by the duration between their time coordinate and the next one. The last period is estimated with the duration of the one just before.
This is the inverse operation of
xclim.core.units.rate2amount().- Parameters:
amount (xr.DataArray or pint.Quantity or str) – “amount” variable. Ex: Precipitation amount in “mm”.
dim (str or xr.DataArray) – The name of the time dimension or the time coordinate itself.
sampling_rate_from_coord (bool) – For data with irregular time coordinates. If True, the diff of the time coordinate will be used as the sampling rate, meaning each data point will be assumed to span the interval ending at the next point. See notes of
xclim.core.units.rate2amount(). Defaults to False, which raises an error if the time coordinate is irregular.out_units (str, optional) – Specific output units, if needed.
- Return type:
DataArray- Returns:
xr.DataArray or Quantity – The converted variable. The standard_name of amount is modified if a conversion is found.
- Raises:
ValueError – If the time coordinate is irregular and sampling_rate_from_coord is False (default).
See also
rate2amountConvert a rate to an amount.
is_temporal_rateDetermine if a variable is a rate based on its CF attributes.
- xclim.core.units.cf_conversion(standard_name, conversion, direction)[source]¶
Get the standard name of the specific conversion for the given standard name.
- Parameters:
standard_name (str) – Standard name of the input.
conversion ({‘amount2rate’, ‘amount2lwethickness’}) – Type of conversion. Available conversions are the keys of the conversions entry in xclim/data/variables.yml. See
xclim.core.units.CF_CONVERSIONS. They also correspond to functions in this module.direction ({‘to’, ‘from’}) – The direction of the requested conversion. “to” means the conversion as given by the conversion name, while “from” means the reverse operation. For example conversion=”amount2rate” and direction=”from” will search for a conversion from a rate or flux to an amount or thickness for the given standard name.
- Return type:
str|None- Returns:
str or None – If a string, this means the conversion is possible and the result should have this standard name. If None, the conversion is not possible within the CF standards.
- xclim.core.units.check_units(val, dim=None)[source]¶
Check that units are compatible with dimensions, otherwise raise a ValidationError.
- Parameters:
val (str or xr.DataArray, optional) – Value to check.
dim (str or xr.DataArray, optional) – Expected dimension, e.g. [temperature]. If a quantity or DataArray is given, the dimensionality is extracted.
- Return type:
None
- xclim.core.units.convert_units_to(source, target, context=None)[source]¶
Convert a mathematical expression into a value with the same units as a DataArray.
If the dimensionalities of source and target units differ, automatic CF conversions will be applied when possible. See
xclim.core.units.cf_conversion().- Parameters:
source (str or xr.DataArray or units.Quantity or xr.Dataset or xr.DataTree) – The value to be converted, e.g. ‘4C’ or ‘1 mm/d’. If a Dataset, target must also be a mapping from variable name to target units. If a DataTree, this function will be applied over nodes with
xarray.DataTree.map_over_datasets().target (str or xr.DataArray or units.Quantity or units.Unit or dict) – Target array of values to which units must conform. If source is a Dataset, it must be mapping from variable name to target units.
context ({“infer”, “hydro”, “none”}, optional) – The unit definition context. Default: None. If “infer”, it will be inferred with
xclim.core.units.infer_context()using the standard name from the source or, if none is found, from the target. This means that the “hydro” context could be activated if any one of the standard names allows it.
- Return type:
DataArray|float|Dataset- Returns:
xr.DataArray or float or xr.Dataset – The source value converted to target’s units. The outputted type is always similar to source initial type. Attributes are preserved unless an automatic CF conversion is performed, in which case only the new standard_name appears in the result.
See also
cf_conversionGet the standard name of the specific conversion for the given standard name.
amount2rateConvert an amount to a rate.
rate2amountConvert a rate to an amount.
amount2lwethicknessConvert an amount to a liquid water equivalent thickness.
lwethickness2amountConvert a liquid water equivalent thickness to an amount.
- xclim.core.units.declare_relative_units(**units_by_name)[source]¶
Function decorator checking the units of arguments.
The decorator checks that input values have units that are compatible with each other. It also stores the input units as a ‘relative_units’ attribute.
- Parameters:
**units_by_name (str) – Mapping from the input parameter names to dimensions relative to other parameters. The dimensions can be a single parameter name as <other_var> or more complex expressions, such as <other_var> * [time].
- Return type:
Callable- Returns:
Callable – The decorated function.
See also
declare_unitsA decorator to check units of function arguments.
Examples
In the following function definition:
@declare_relative_units(thresh="<da>", thresh2="<da> / [time]") def func(da, thresh, thresh2): ...
The decorator will check that thresh has units compatible with those of da and that thresh2 has units compatible with the time derivative of da.
Usually, the function would be decorated further by
declare_units()to create a unit-aware index:temperature_func = declare_units(da="[temperature]")(func)
This call will replace the “<da>” by “[temperature]” everywhere needed.
- xclim.core.units.declare_units(**units_by_name)[source]¶
Create a decorator to check units of function arguments.
The decorator checks that input and output values have units that are compatible with expected dimensions. It also stores the input units as an ‘in_units’ attribute.
- Parameters:
**units_by_name (str) – Mapping from the input parameter names to their units or dimensionality (“[…]”). If this decorates a function previously decorated with
declare_relative_units(), the relative unit declarations are made absolute with the information passed here.- Return type:
Callable- Returns:
Callable – The decorated function.
See also
declare_relative_unitsA decorator to check for relative units of function arguments.
Examples
In the following function definition:
@declare_units(tas="[temperature]") def func(tas): ...
The decorator will check that tas has units of temperature (C, K, F).
- xclim.core.units.ensure_absolute_temperature(units)[source]¶
Convert temperature units to their absolute counterpart, assuming they represented a difference (delta).
Celsius becomes Kelvin, Fahrenheit becomes Rankine. Does nothing for other units.
- Parameters:
units (str) – Units to transform.
- Return type:
str- Returns:
str – The transformed units.
See also
ensure_deltaEnsure a unit is a delta unit.
- xclim.core.units.ensure_cf_units(ustr)[source]¶
Ensure the passed unit string is CF-compliant.
The string will be parsed to pint then recast to a string by
xclim.core.units.pint2cfunits().- Parameters:
ustr (str) – A unit string.
- Return type:
str- Returns:
str – The unit string in CF-compliant form.
- xclim.core.units.ensure_delta(unit)[source]¶
Return delta units for temperature.
For dimensions where delta exist in pint (Temperature), it replaces the temperature unit by delta_degC or delta_degF based on the input unit. For other dimensionality, it just gives back the input units.
- Parameters:
unit (str) – Unit to transform in delta (or not).
- Return type:
str- Returns:
str – The transformed units.
- xclim.core.units.flux2rate(flux, density, out_units=None)[source]¶
Convert a flux variable to a rate by dividing with a density.
This is the inverse operation of
xclim.core.units.rate2flux().- Parameters:
flux (xr.DataArray) – “flux” variable, e.g. Snowfall flux in “kg m-2 s-1”.
density (Quantified) – Density used to convert from a flux to a rate, e.g. Snowfall density “312 kg m-3”. Density can also be an array with the same shape as flux.
out_units (str, optional) – Specific output units, if needed.
- Return type:
DataArray- Returns:
xr.DataArray – The converted rate value.
See also
rate2fluxConvert a rate to a flux.
Examples
The following converts an array of snowfall flux in kg m-2 s-1 to snowfall flux in mm/s, assuming a density of 100 kg m-3:
>>> time = xr.date_range("2001-01-01", freq="D", periods=365) >>> prsn = xr.DataArray( ... [0.1] * 365, ... dims=("time",), ... coords={"time": time}, ... attrs={"units": "kg m-2 s-1"}, ... ) >>> prsnd = flux2rate(prsn, density="100 kg m-3", out_units="mm/s") >>> prsnd.units 'mm s-1' >>> float(prsnd[0]) 1.0
- xclim.core.units.infer_context(standard_name=None, dimension=None)[source]¶
Return units context based on either the variable’s standard name or the pint dimension.
Valid standard names for the hydro context are those including the terms “rainfall”, “lwe” (liquid water equivalent) and “precipitation”. The latter is technically incorrect, as any phase of precipitation could be referenced. Standard names for evapotranspiration, evaporation and canopy water amounts are also associated with the hydro context.
- Parameters:
standard_name (str, optional) – CF-Convention standard name.
dimension (str, optional) – Pint dimension, e.g. ‘[time]’.
- Return type:
str- Returns:
str – “hydro” if variable refers to liquid water or to a mass of water in any phase, otherwise “none”.
- xclim.core.units.infer_sampling_units(da, deffreq=None, dim='time')[source]¶
Infer a multiplier and the units corresponding to one sampling period.
- Parameters:
da (xr.DataArray) – A DataArray from which to take coordinate dim.
deffreq (str, optional) – If no frequency is inferred from da[dim], take this one.
dim (str) – Dimension from which to infer the frequency.
- Return type:
tuple[int,str]- Returns:
int – The magnitude (number of base periods per period).
str – Units as a string, understandable by pint.
- Raises:
ValueError – If the frequency has no corresponding units.
- xclim.core.units.lwethickness2amount(thickness, out_units=None)[source]¶
Convert a liquid water thickness (length) to its equivalent amount (mass over area).
This will simply multiply the thickness by the density of liquid water, 1000 kg/m³. This is equivalent to using the “hydro” context of
xclim.core.units.units.- Parameters:
thickness (xr.DataArray) – A DataArray storing a liquid water thickness quantity.
out_units (str, optional) – Specific output units, if needed.
- Return type:
Union[DataArray,TypeVar(Quantified,DataArray,str,Quantity)]- Returns:
xr.DataArray or Quantified – The standard_name of amount is modified if a conversion is found (see
xclim.core.units.cf_conversion()), it is removed otherwise. Other attributes are left untouched.
See also
amount2lwethicknessConvert an amount to a liquid water equivalent thickness.
- xclim.core.units.pint2cfattrs(value, is_difference=None)[source]¶
Return CF-compliant units attributes from a pint unit.
- Parameters:
value (pint.Unit) – Input unit.
is_difference (bool, optional) – Whether the value represent a difference in temperature, which is ambiguous in the case of absolute temperature scales like Kelvin or Rankine. Default is to guess, see notes.
- Return type:
dict[str,str]- Returns:
dict – Units following CF-Convention, using symbols.
Notes
Temperatures are understood as differences if the input unit contains any of the known difference units (see
TEMPERATURE_DELTA_UNITS). They are understood as on-scale if the input unit is exactly one of the known on-scale units (°C, °F or °Re). Otherwise, “unknown” is given as the “units metadata” (this usually happens with K or °R, or with composed units like °C d). Of course,units_metadatais not added if the input has no temperature dimension.
- xclim.core.units.pint2cfunits(value)[source]¶
Return a CF-compliant unit string from a pint unit.
- Parameters:
value (pint.Unit) – Input unit.
- Return type:
str- Returns:
str – Units following CF-Convention, using symbols.
- xclim.core.units.pint_multiply(da, q, out_units=None)[source]¶
Multiply xarray.DataArray by pint.Quantity.
- Parameters:
da (xr.DataArray) – Input array.
q (pint.Quantity) – Multiplicative factor.
out_units (str, optional) – Units the output array should be converted into.
- Return type:
DataArray- Returns:
xr.DataArray – The product DataArray.
- xclim.core.units.rate2amount(rate, dim='time', sampling_rate_from_coord=False, out_units=None)[source]¶
Convert a rate variable to an amount by multiplying by the sampling period length.
If the sampling period length cannot be inferred, the rate values are multiplied by the duration between their time coordinate and the next one. The last period is estimated with the duration of the one just before.
This is the inverse operation of
xclim.core.units.amount2rate().- Parameters:
rate (xr.DataArray or pint.Quantity or str) – “Rate” variable, with units of “amount” per time. Ex: Precipitation in “mm / d”.
dim (str or DataArray) – The name of time dimension or the coordinate itself.
sampling_rate_from_coord (bool) – For data with irregular time coordinates. If True, the diff of the time coordinate will be used as the sampling rate, meaning each data point will be assumed to apply for the interval ending at the next point. See notes. Defaults to False, which raises an error if the time coordinate is irregular.
out_units (str, optional) – Specific output units, if needed.
- Return type:
DataArray- Returns:
xr.DataArray or Quantity – The converted variable. The standard_name of rate is modified if a conversion is found.
- Raises:
ValueError – If the time coordinate is irregular and sampling_rate_from_coord is False (default).
See also
amount2rateConvert an amount to a rate.
is_temporal_rateDetermine if a variable is a rate based on its CF attributes.
Notes
Floating-point precision can have surprising results. For example, a daily series of 1 mm/d precipitation rates might not convert to exactly 1 mm daily amounts. This is because a float multiplication is still happening in the background and the time step duration might have been stored in [nano]seconds at one point.
Examples
The following converts a daily array of precipitation in mm/h to the daily amounts in mm:
>>> time = xr.date_range("2001-01-01", freq="D", periods=365) >>> pr = xr.DataArray([1] * 365, dims=("time",), coords={"time": time}, attrs={"units": "mm/h"}) >>> pram = rate2amount(pr) >>> pram.units 'mm' >>> float(pram[0]) 24.0
Also works if the time axis is irregular : the rates are assumed constant for the whole period starting on the values timestamp to the next timestamp. This option is activated with sampling_rate_from_coord=True.
>>> time = time[[0, 9, 30]] # The time axis is Jan 1st, Jan 10th, Jan 31st >>> pr = xr.DataArray([1] * 3, dims=("time",), coords={"time": time}, attrs={"units": "mm/h"}) >>> pram = rate2amount(pr, sampling_rate_from_coord=True) >>> pram.values array([216., 504., 504.])
Finally, we can force output units:
>>> pram = rate2amount(pr, out_units="pc") # Get rain amount in parsecs. Why not. >>> pram.values array([7.00008327e-18, 1.63335276e-17, 1.63335276e-17])
- xclim.core.units.rate2flux(rate, density, out_units=None)[source]¶
Convert a rate variable to a flux by multiplying with a density.
This is the inverse operation of
xclim.core.units.flux2rate().- Parameters:
rate (xr.DataArray) – “Rate” variable, e.g. Snowfall rate in “mm / d”.
density (Quantified) – Density used to convert from a rate to a flux, e.g. Snowfall density “312 kg m-3”. Density can also be an array with the same shape as rate.
out_units (str, optional) – Specific output units, if needed.
- Return type:
DataArray- Returns:
xr.DataArray – The converted flux value.
See also
flux2rateConvert a flux to a rate.
Examples
The following converts an array of snowfall rate in mm/s to snowfall flux in kg m-2 s-1, assuming a density of 100 kg m-3:
>>> time = xr.date_range("2001-01-01", freq="D", periods=365) >>> prsnd = xr.DataArray([1] * 365, dims=("time",), coords={"time": time}, attrs={"units": "mm/s"}) >>> prsn = rate2flux(prsnd, density="100 kg m-3", out_units="kg m-2 s-1") >>> prsn.units 'kg m-2 s-1' >>> float(prsn[0]) 0.1
- xclim.core.units.str2pint(val)[source]¶
Convert a string to a pint.Quantity, splitting the magnitude and the units.
- Parameters:
val (str) – A quantity in the form “[{magnitude} ]{units}”, where magnitude can be cast to a float and units is understood by
xclim.core.units.units2pint().- Return type:
Quantity- Returns:
pint.Quantity – Magnitude is 1 if no magnitude was present in the string.
- xclim.core.units.to_agg_units(out, orig, statistic, dim='time', deffreq='D')[source]¶
Set and convert units of an array after an aggregation operation along the sampling dimension (time).
- Parameters:
out (xr.DataArray) – The output array of the aggregation operation, no units operation done yet.
orig (xr.DataArray) – The original array before the aggregation operation, used to infer the sampling units and get the variable units.
statistic ({‘min’, ‘max’, ‘mean’, ‘std’, ‘var’, ‘doymin’, ‘doymax’, ‘count’, ‘integral’, ‘sum’} or Callable) – The type of aggregation operation performed. “integral” is mathematically equivalent to “sum”, but the units are multiplied by the timestep of the data (requires an inferrable frequency).
dim (str) – The time dimension along which the aggregation was performed.
deffreq (str, optional) – For operations count and integral, this gives the default source frequency to assume, if it can’t be inferred from
out[dim].
- Return type:
DataArray- Returns:
xr.DataArray – The DataArray with aggregated values. Depending on configurations, units may also be converted or simplified.
Examples
Take a daily array of temperature and count number of days above a threshold. to_agg_units will infer the units from the sampling rate along “time”, so we ensure the final units are correct:
>>> time = xr.date_range("2001-01-01", freq="D", periods=365) >>> tas = xr.DataArray( ... np.arange(365), ... dims=("time",), ... coords={"time": time}, ... attrs={"units": "degC"}, ... ) >>> cond = tas > 100 # Which days are boiling >>> Ndays = cond.sum("time") # Number of boiling days
# Note: older xarray drops units while modern xarray preserves them >>> Ndays.attrs.get(“units”) # doctest: +SKIP ‘degC’ >>> Ndays = to_agg_units(Ndays, tas, “count”) >>> Ndays.units ‘d’
Similarly, here we compute the total heating degree-days, but we have weekly data:
>>> time = xr.date_range("2001-01-01", freq="7D", periods=52) >>> tas = xr.DataArray( ... np.arange(52) + 10, ... dims=("time",), ... coords={"time": time}, ... ) >>> dt = (tas - 16).assign_attrs(units="degC", units_metadata="temperature: difference") >>> degdays = dt.clip(0).sum("time") # Integral of temperature above a threshold >>> degdays = to_agg_units(degdays, dt, statistic="integral") >>> degdays.units '°C week'
Which we can always convert to the more common “K days”:
>>> degdays = convert_units_to(degdays, "K days") >>> degdays.units 'd K'
- xclim.core.units.units2pint(value)[source]¶
Return the pint Unit for the DataArray units.
- Parameters:
value (xr.DataArray or pint.Unit or pint.Quantity or dict or str) – Input data array or string representing a unit (with no magnitude).
- Return type:
Unit- Returns:
pint.Unit – Units of the data array.
Notes
To avoid ambiguity related to differences in temperature vs absolute temperatures, set the units_metadata attribute to “temperature: difference” or “temperature: on_scale” on the DataArray.
Calendar Handling Utilities¶
Helper function to handle dates, times and different calendars with xarray.
- xclim.core.calendar.add_season_coord(ds, freq)[source]¶
Add a season coordinates on a resampled dataset.
- Parameters:
ds (xr.Dataset or xr.DataArray) – The xarray object with a “time” coordinate. Only supports daily or coarser frequencies (excluding weekly). The time axis must be complete and regular (xr.infer_freq(ds.time) doesn’t fail).
freq (str) – Resampling frequency. Must be between “MS” and “YS” and divide a year evenly.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – Input dataset with season coordinate.
- xclim.core.calendar.adjust_doy_calendar(source, target)[source]¶
Interpolate from one set of dayofyear range to another calendar.
Interpolate an array defined over a dayofyear range (say 1 to 360) to another dayofyear range (say 1 to 365).
- Parameters:
source (xr.DataArray or xr.Dataset) – Array with dayofyear coordinate.
target (xr.DataArray or xr.Dataset) – Array with time coordinate.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – Interpolated source array over coordinates spanning the target dayofyear range.
- xclim.core.calendar.build_climatology_bounds(da)[source]¶
Build the climatology_bounds property with the start and end dates of input data.
- Parameters:
da (xr.DataArray) – The input data. Must have a time dimension.
- Return type:
list[str]- Returns:
list of str – The climatology bounds.
- xclim.core.calendar.climatological_mean_doy(arr, window=5)[source]¶
Calculate the climatological mean and standard deviation for each day of the year.
- Parameters:
arr (xarray.DataArray) – Input array.
window (int) – Window size in days.
- Return type:
tuple[DataArray,DataArray]- Returns:
xarray.DataArray, xarray.DataArray – Mean and standard deviation.
- xclim.core.calendar.common_calendar(calendars, join='outer')[source]¶
Return a calendar common to all calendars from a list.
Uses the hierarchy: 360_day < noleap < standard < all_leap.
- Parameters:
calendars (Sequence of str) – List of calendar names.
join ({‘inner’, ‘outer’}) –
- The criterion for the common calendar.
- ‘outer’: the common calendar is the biggest calendar (in number of days by year) that will include all the
dates of the other calendars. When converting the data to this calendar, no timeseries will lose elements, but some might be missing (gaps or NaNs in the series).
- ‘inner’: the common calendar is the smallest calendar of the list.
When converting the data to this calendar, no timeseries will have missing elements (no gaps or NaNs), but some might be dropped.
- Return type:
str- Returns:
str – Returns “default” only if all calendars are “default”.
Examples
>>> common_calendar(["360_day", "noleap", "default"], join="outer") 'standard' >>> common_calendar(["360_day", "noleap", "default"], join="inner") '360_day'
- xclim.core.calendar.compare_offsets(freqA, op, freqB)[source]¶
Compare offsets string based on their approximate length, according to a given operator.
Offsets are compared based on their length approximated for a period starting after 1970-01-01 00:00:00. If the offsets are from the same category (same first letter), only the multiplier prefix is compared (QS-DEC == QS-JAN, MS < 2MS). “Business” offsets are not implemented.
- Parameters:
freqA (str) – RHS Date offset string (‘YS’, ‘1D’, ‘QS-DEC’, …).
op ({“>”, “gt”, “<”, “lt”, “>=”, “ge”, “<=”, “le”, “==”, “eq”, “!=”, “ne”}) – Operator to use.
freqB (str) – LHS Date offset string (‘YS’, ‘1D’, ‘QS-DEC’, …).
- Return type:
bool- Returns:
bool – The result of freqA op freqB.
- xclim.core.calendar.construct_offset(mult, base, start_anchored, anchor)[source]¶
Reconstruct an offset string from its parts.
- Parameters:
mult (int) – The period multiplier (>= 1).
base (str) – The base period string (one char).
start_anchored (bool) – If True and base in [Y, Q, M], adds the “S” flag, False add “E”.
anchor (str, optional) – The month anchor of the offset. Defaults to JAN for bases YS and QS and to DEC for bases YE and QE.
- Returns:
str – An offset string, conformant to pandas-like naming conventions.
Notes
This provides the mirror opposite functionality of
parse_offset().
- xclim.core.calendar.convert_doy(source, target_cal, source_cal=None, align_on='year', missing=nan, dim='time')[source]¶
Convert the calendar of day of year (doy) data.
- Parameters:
source (xr.DataArray or xr.Dataset) – Day of year data (range [1, 366], max depending on the calendar). If a Dataset, the function is mapped to each variable with attribute is_day_of_year == 1.
target_cal (str) – Name of the calendar to convert to.
source_cal (str, optional) – Calendar the doys are in. If not given, will use the “calendar” attribute of source or, if absent, the calendar of its dim axis.
align_on ({‘date’, ‘year’}) – If ‘year’ (default), the doy is seen as a “percentage” of the year and is simply rescaled onto the new doy range. This always results in floating point data, changing the decimal part of the value. If ‘date’, the doy is seen as a specific date. See notes. This never changes the decimal part of the value.
missing (Any) – If align_on is “date” and the new doy doesn’t exist in the new calendar, this value is used.
dim (str) – Name of the temporal dimension.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – The converted doy data.
- xclim.core.calendar.days_since_to_doy(da, start=None, calendar=None)[source]¶
Reverse the conversion made by
doy_to_days_since().Converts data given in days since a specific date to day-of-year.
- Parameters:
da (xr.DataArray) – The result of
doy_to_days_since().start (DateOfYearStr, optional) – da is considered as days since that start date (in the year of the time index). If None (default), it is read from the attributes.
calendar (str, optional) – Calendar the “days since” were computed in. If None (default), it is read from the attributes.
- Return type:
DataArray- Returns:
xr.DataArray – Same shape as da, values as day of year.
Examples
>>> time = xr.date_range("2020-07-01", "2021-07-01", freq="YS-JUL") >>> da = xr.DataArray( ... [-86, 92], ... dims=("time",), ... coords={"time": time}, ... attrs={"units": "days since 10-02"}, ... ) >>> days_since_to_doy(da).values array([190, 2])
- xclim.core.calendar.doy_from_string(doy, year, calendar)[source]¶
Return the day-of-year corresponding to an “MM-DD” string for a given year and calendar.
- Parameters:
doy (str) – The day of year in the format “MM-DD”.
year (int) – The year.
calendar (str) – The calendar name.
- Return type:
int- Returns:
int – The day of year.
- xclim.core.calendar.doy_to_days_since(da, start=None, calendar=None)[source]¶
Convert day-of-year data to days since a given date.
This is useful for computing meaningful statistics on doy data.
- Parameters:
da (xr.DataArray) – Array of “day-of-year”, usually int dtype, must have a time dimension. Sampling frequency should be finer or similar to yearly and coarser than daily.
start (date of year str, optional) – A date in “MM-DD” format, the base day of the new array. If None (default), the time axis is used. Passing start only makes sense if da has a yearly sampling frequency.
calendar (str, optional) – The calendar to use when computing the new interval. If None (default), the calendar attribute of the data or of its time axis is used. All time coordinates of da must exist in this calendar. No check is done to ensure doy values exist in this calendar.
- Return type:
DataArray- Returns:
xr.DataArray – Same shape as da, int dtype, day-of-year data translated to a number of days since a given date. If start is not None, there might be negative values.
Notes
The time coordinates of da are considered as the START of the period. For example, a doy value of 350 with a timestamp of ‘2020-12-31’ is understood as ‘2021-12-16’ (the 350th day of 2021). Passing start=None, will use the time coordinate as the base, so in this case the converted value will be 350 “days since time coordinate”.
Examples
>>> time = xr.date_range("2020-07-01", "2021-07-01", freq="YS-JUL") >>> # July 8th 2020 and Jan 2nd 2022 >>> da = xr.DataArray([190, 2], dims=("time",), coords={"time": time}) >>> # Convert to days since Oct. 2nd, of the data's year. >>> doy_to_days_since(da, start="10-02").values array([-86, 92])
- xclim.core.calendar.ensure_cftime_array(time)[source]¶
Convert an input 1D array to a numpy array of cftime objects.
Python datetimes are converted to cftime.DatetimeGregorian (“standard” calendar).
- Parameters:
time (sequence) – A 1D array of datetime-like objects.
- Return type:
ndarray|Sequence[datetime]- Returns:
np.ndarray – An array of cftime.datetime objects.
- Raises:
ValueError – When unable to cast the input.:
- xclim.core.calendar.get_calendar(obj, dim='time')[source]¶
Return the calendar of an object.
- Parameters:
obj (Any) – An object defining some date. If obj is an array/dataset with a datetime coordinate, use dim to specify its name. Values must have either a datetime64 dtype or a cftime dtype. obj can also be a python datetime.datetime, a cftime object or a pandas Timestamp or an iterable of those, in which case the calendar is inferred from the first value.
dim (str) – Name of the coordinate to check (if obj is a DataArray or Dataset).
- Return type:
str- Returns:
str – The Climate and Forecasting (CF) calendar name. Will always return “standard” instead of “gregorian”, following CF-Conventions v1.9.
- Raises:
ValueError – If no calendar could be inferred.
- xclim.core.calendar.is_offset_divisor(divisor, offset)[source]¶
Check that divisor is a divisor of offset.
A frequency is a “divisor” of another if a whole number of periods of the former fit within a single period of the latter.
- Parameters:
divisor (str) – The divisor frequency.
offset (str) – The large frequency.
- Returns:
bool – Whether divisor is a divisor of offset.
Examples
>>> is_offset_divisor("QS-JAN", "YS") True >>> is_offset_divisor("QS-DEC", "YS-JUL") False >>> is_offset_divisor("D", "ME") True
- xclim.core.calendar.parse_offset(freq)[source]¶
Parse an offset string.
Parse a frequency offset and, if needed, convert to cftime-compatible components.
- Parameters:
freq (str) – Frequency offset.
- Return type:
tuple[int,str,bool,str|None]- Returns:
multiplier (int) – Multiplier of the base frequency. “[n]W” is always replaced with “[7n]D”, as xarray doesn’t support “W” for cftime indexes.
offset_base (str) – Base frequency.
is_start_anchored (bool) – Whether coordinates of this frequency should correspond to the beginning of the period (True) or its end (False). Can only be False when base is Y, Q or M; in other words, xclim assumes frequencies finer than monthly are all start-anchored.
anchor (str, optional) – Anchor date for bases Y or Q. As xarray doesn’t support “W”, neither does xclim (anchor information is lost when given).
- xclim.core.calendar.percentile_doy(arr, window=5, per=10.0, alpha=0.3333333333333333, beta=0.3333333333333333, copy=True)[source]¶
Percentile value for each day of the year.
Return the climatological percentile over a moving window around each day of the year. Different quantile estimators can be used by specifying alpha and beta according to specifications given by Hyndman and Fan [1996]. The default definition corresponds to method 8, which meets multiple desirable statistical properties for sample quantiles. Note that numpy.percentile corresponds to method 7, with alpha and beta set to 1.
- Parameters:
arr (xr.DataArray) – Input data, a daily frequency (or coarser) is required.
window (int) – Number of time-steps around each day of the year to include in the calculation.
per (float or sequence of float) – Percentile(s) between [0, 100].
alpha (float) – Plotting position parameter.
beta (float) – Plotting position parameter.
copy (bool) – If True (default) the input array will be deep-copied. It’s a necessary step to keep the data integrity, but it can be costly. If False, no copy is made of the input array. It will be mutated and rendered unusable, but performances may significantly improve. Put this flag to False only if you understand the consequences.
- Return type:
DataArray- Returns:
xr.DataArray – The percentiles indexed by the day of the year. For calendars with 366 days, percentiles of doys 1-365 are interpolated to the 1-366 range.
References
Hyndman and Fan [1996]
- xclim.core.calendar.resample_doy(doy, arr)[source]¶
Create a temporal DataArray where each day takes the value defined by the day-of-year.
- Parameters:
doy (xr.DataArray or xr.Dataset) – Array with dayofyear coordinate.
arr (xr.DataArray or xr.Dataset) – Array with time coordinate.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – An array with the same dimensions as doy, except for dayofyear, which is replaced by the time dimension of arr. Values are filled according to the day of year value in doy.
- xclim.core.calendar.select_time(da, drop=False, season=None, month=None, doy_bounds=None, date_bounds=None, include_bounds=True, include_doy_bounds_nans=True, bounds_freq=None)[source]¶
Select entries according to a time period.
This conveniently improves xarray’s
xarray.DataArray.where()andxarray.DataArray.sel()with fancier ways of indexing over time elements. In addition to the data da and argument drop, only one of season, month, doy_bounds or date_bounds may be passed.- Parameters:
da (xr.DataArray or xr.Dataset) – Input data.
drop (bool) – Whether to drop elements outside the period of interest (True) or to simply mask them (False, default). This option is incompatible with passing date_bounds or array-like doy_bounds.
season (str or sequence of str, optional) – One or more of ‘DJF’, ‘MAM’, ‘JJA’ and ‘SON’.
month (int or sequence of int, optional) – Sequence of month numbers (January = 1 … December = 12).
doy_bounds (2-tuple of optional integers or DataArray, optional) – The bounds as (start, end) of the period of interest expressed in day-of-year, integers going from 1 (January 1st) to 365 or 366 (December 31st). If DataArrays are passed, they must have the same coordinates on the dimensions they share. They may have a time dimension, in which case the selection is done independently for each period defined by the coordinate, which means the time coordinate must have an inferable frequency (see
xr.infer_freq()) or the frequency must be passed explicitly with the bounds_freq argument. If None is passed as a bound, it is replaced by the start or end of the year (1 or 366) if the other bound is an integer, or by the start or end of the period defined by the inferred or passed frequency of DataArrays. Timesteps of the input not appearing in the time coordinate of the bounds are considered as “outside the bounds”.date_bounds (2-tuple of optional strings, optional) – The bounds as (start, end) of the period of interest expressed as dates in the month-day (%m-%d) format. If None is passed as a bounds, it is replaced by the start or end of the period defined by the bounds_freq argument, corresponding to 1st January or 31st December for default “YS” bounds frequency.
include_bounds (bool or 2-tuple of bool, optional) – Whether the bounds of doy_bounds or date_bounds should be inclusive or not. Either one value for both or a tuple. Default is True, meaning bounds are inclusive.
include_doy_bounds_nans (bool, optional) – Whether to include values associated with NaN in doy_bounds. If True (default), missing values (NaN) in the start and end bounds are replaced by the start and end of the period, respectively.
bounds_freq (str, optional) – Needed with array-like doy_bounds without a time dimension or date_bounds, and corresponding to the frequency used to determine the start and end of the period (default “YS”). If doy_bounds have a time dimension, the frequency is first tried to be inferred from the time coordinate of the bounds; if it cannot be inferred, the frequency must be passed explicitly.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – Selected input values. If
drop=False, this has the same length asda(along dimension ‘time’), but with masked (NaN) values outside the period of interest.
Examples
Keep only the values of fall and spring.
>>> ds = xr.open_dataset("ERA5/daily_surface_cancities_1990-1993.nc") >>> ds.time.size 1461 >>> out = select_time(ds, drop=True, season=["MAM", "SON"]) >>> out.time.size 732
Or all values between two dates (included).
>>> out = select_time(ds, drop=True, date_bounds=("02-29", "03-02")) >>> out.time.values array(['1990-03-01T00:00:00.000000000', '1990-03-02T00:00:00.000000000', '1991-03-01T00:00:00.000000000', '1991-03-02T00:00:00.000000000', '1992-02-29T00:00:00.000000000', '1992-03-01T00:00:00.000000000', '1992-03-02T00:00:00.000000000', '1993-03-01T00:00:00.000000000', '1993-03-02T00:00:00.000000000'], dtype='datetime64[ns]')
- xclim.core.calendar.split_time_to_season_year(ds, freq)[source]¶
Split a resampled dataset into a yearly time and a season coordinate.
- Parameters:
ds (xr.Dataset or xr.DataArray) – The xarray object with a “time” coordinate. Only supports daily or coarser frequencies (excluding weekly). The time axis must be complete and regular (xr.infer_freq(ds.time) doesn’t fail).
freq (str) – Resampling frequency. Must be between “MS” and “YS” and divide a year evenly.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray or xr.Dataset – Input dataset with season coordinate and yearly time.
- xclim.core.calendar.stack_periods(da, window=30, stride=None, min_length=None, freq='YS', dim='period', start='1970-01-01', align_days=True, pad_value='<NA>')[source]¶
Construct a multi-period array.
Stack different equal-length periods of da into a new ‘period’ dimension.
This is similar to
da.rolling(time=window).construct(dim, stride=stride), but adapted for arguments in terms of a base temporal frequency that might be non-uniform (years, months, etc.). It is reversible for some cases (see stride). A rolling-construct method will be much more performant for uniform periods (days, weeks).- Parameters:
da (xr.Dataset or xr.DataArray) – An xarray object with a time dimension. Must have a uniform timestep length. Output might be strange if this does not use a uniform calendar (noleap, 360_day, all_leap).
window (int) – The length of the moving window as a multiple of
freq.stride (int, optional) – At which interval to take the windows, as a multiple of
freq. For the operation to be reversible withunstack_periods(), it must divide window into an odd number of parts. Default is window (no overlap between periods).min_length (int, optional) – Windows shorter than this are not included in the output. Given as a multiple of
freq. Default iswindow(every window must be complete). Similar to themin_periodsargument ofda.rolling. Iffreqis annual or quarterly andmin_length == ``window, the first period is considered complete if the first timestep is in the first month of the period.freq (str) – Units of
window,strideandmin_length, as a frequency string. Must be larger or equal to the data’s sampling frequency. Note that this function offers an easier interface for non-uniform period (like years or months) but is much slower than a rolling-construct method.dim (str) – The new dimension name.
start (str) – The start argument passed to
xarray.date_range()to generate the new placeholder time coordinate.align_days (bool) – When True (default), an error is raised if the output would have unaligned days across periods. If freq = ‘YS’, day-of-year alignment is checked and if freq is “MS” or “QS”, we check day-in-month. Only uniform-calendar will pass the test for freq=’YS’. For other frequencies, only the 360_day calendar will work. This check is ignored if the sampling rate of the data is coarser than “D”.
pad_value (Any) – When some periods are shorter than others, this value is used to pad them at the end. Passed directly as argument
fill_valuetoxarray.concat(), the default is the same as on that function.
- Return type:
TypeVar(DataType,DataArray,Dataset)- Returns:
xr.DataArray – A DataArray with a new period dimension and a time dimension with the length of the longest window. The new time coordinate has the same frequency as the input data but is generated using
xarray.date_range()with the given start value. That coordinate is the same for all periods, depending on the choice ofwindowandfreq, it might make sense. But for unequal periods or non-uniform calendars, it will certainly not. Ifstrideis a divisor ofwindow, the correct timeseries can be reconstructed withunstack_periods(). The coordinate of period is the first timestep of each window.
- xclim.core.calendar.time_bnds(time, freq=None)[source]¶
Find the time bounds for a datetime index by assuming an uniform sampling frequency.
As we are using datetime indices to stand in for period indices, assumptions regarding the period are made based on the given freq. This function does not implement finding bounds for an irregular time index.
- Parameters:
time (DataArray, Dataset, CFTimeIndex, DatetimeIndex, DataArrayResample or DatasetResample) – Object which contains a time index as a proxy representation for a period index.
freq (str, optional) – String specifying the frequency/offset such as ‘MS’, ‘2D’, or ‘3min’ If not given, it is inferred from the time index, which means that index must have at least three elements.
- Returns:
DataArray – The time bounds: start and end times of the periods inferred from the time index and a frequency. It has the original time index along it’s time coordinate and a new bnds coordinate. The dtype and calendar of the array are the same as the index. If a period follows another, its start is the same as the other’s end.
Notes
xclim assumes that indexes for greater-than-day frequencies are “floored” down to a daily resolution. For example, the coordinate “2000-01-31 00:00:00” with a “ME” frequency is assumed to mean a period going from “2000-01-01 00:00:00” to “2000-02-01 00:00:00”.
Similarly, it assumes that daily and finer frequencies yield indexes pointing to the period’s start. So “2000-01-31 00:00:00” with a “3h” frequency, means a period going from “2000-01-31 00:00:00” to “2000-01-31 03:00:00”.
See the relevant CF convention <https://cfconventions.org/Data/cf-conventions/cf-conventions-1.13/cf-conventions.html#bounds-one-d>.
- xclim.core.calendar.within_bnds_doy(arr, *, low, high)[source]¶
Return whether array values are within bounds for each day of the year.
- Parameters:
arr (xarray.DataArray) – Input array.
low (xarray.DataArray) – Low bound with dayofyear coordinate.
high (xarray.DataArray) – High bound with dayofyear coordinate.
- Return type:
DataArray- Returns:
xarray.DataArray – Boolean array of values within doy.