# Xarray MemoryError: with groupby workloads

**URL:** <https://discourse.pangeo.io/t/xarray-memoryerror-with-groupby-workloads/4273>\
**Category:** Uncategorized\
**Created:** [June 11, 2024, 9:07am UTC](https://discourse.pangeo.io/t/xarray-memoryerror-with-groupby-workloads/4273 "2024-06-11T09:07:11Z")\
**Posts on this page:** 1\
**Showing post:** 9

<div class="post-metadata">

**Author:** ![dcherian](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/dcherian/32/2235_2.png) [@dcherian](https://discourse.pangeo.io/u/dcherian)\
**Post date:** [June 11, 2024, 9:42pm UTC](https://discourse.pangeo.io/t/xarray-memoryerror-with-groupby-workloads/4273/9 "2024-06-11T21:42:41Z")

</div>

Julius makes a key observation. If the quantile is indeed calculated on a per year basis, then this is a much simpler problem.

### First construct a dummy dataset

I’ve assumed that a month of daily data is in a single netCDF file:

```python
import dask.array
import numpy as np
import xarray as xr

ds = xr.Dataset(
    {
        "pr": (
            ("time", "lat", "lon"),
            dask.array.random.random(
                (730487, 180, 360),
                chunks=(30, -1, -1), # assuming one month per file
            ),
        )
    },
    coords={"time": xr.date_range("0001-01-01", "2000-12-31", freq="D")},
)
ds

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/2/20af816652116f38b51c0d5d4fde944c4ff81aad.png)

### Now rechunk to a yearly frequency

This is the key observation. Since you’re operating on a year of data at a time, you don’t need to rechunk to `time=-1` instead you can rechunk so that a year of data is in a single chunk. This rechunking is **required** for computing the quantile.

```python
# rechunk to yearly frequency
from xarray.groupers import TimeResampler

ds = ds.chunk({"time": TimeResampler(freq="AS")})
ds

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/c/c32399b76485231be341bdbb12d9330970bfd1d3.png)

^ EDIT: Use rechunking with a TimeResampler object.

### Now do your groupby.apply

```python
ds = ds.where(ds >= 1)
result = ds.groupby("time.year").apply(
    lambda group: group.where(group > group.quantile(q=0.95)).sum("time")
)
result

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/5/5ade2068299ff52c05c15e6f940925677b6f8027.png)

### This works quite well!

```python
result.compute()

```

![dask](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/e/ee1455d29dfbabd86124e52a97d979a8289fb4e8.gif)

---

_[View the full topic](https://discourse.pangeo.io/t/xarray-memoryerror-with-groupby-workloads/4273)._
