# Extremely slow xarray/zarr writes

**URL:** <https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434>\
**Category:** Data\
**Created:** [August 20, 2024, 9:26pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434 "2024-08-20T21:26:08Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![vbalza](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/vbalza/32/2884_2.png) [@vbalza](https://discourse.pangeo.io/u/vbalza)\
**Post date:** [August 20, 2024, 9:26pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/1 "2024-08-20T21:26:08Z")

</div>

I’m trying to store a 64MiB `xarray` `DataSet` to a zarr store via `Dataset.to_zarr` and am seeing speeds of about ~10minutes with the following chunks for a DataSet that includes latitude, longitude, forecast day (`f_day`) and a mean abs error measuring diff between actual temperature and forecasted temperature on the given `f_day`.

I expected much faster write times but am not sure if I’m doing something wrong and am somewhat new to `dask`. I tried switching up the chunks to no avail. Any thoughts?

 ![Screenshot 2024-08-20 at 4.22.54 PM](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/7/72b3c9a91187ca75aeb938af0878c367d31503d5.png)

For reference, here is how I’m creating the data:

```auto
        monthly_means = []
        for f_day in forecast_days:
            logger.info("Processing forecast day %s", f_day)
            # Retrieve data from zarr store in AWS
            month_data = self.get_single_month_data(year, month, forecast_day=f_day)
            # Compare to in-memory ERA5 data also retrieved from zarr store in AWS and compute mean
            # Select only overlapping times between monthly and ERA5 data
            monthly_mean = (
                month_data - era5_data.sel(time=month_data.time.values)
            ).mean(dim="time")

            monthly_mean["f_day"] = f_day
            monthly_means.append(monthly_mean)
        # Concatenate means from all forecast days
        ds = xr.concat(monthly_means, dim="f_day").set_coords("f_day").rename_vars(
            {self.variable: "mean_abs_err"}
        )
        # Re-chunk data
        return ds.chunk({"f_day": -1, "latitude": 360, "longitude": 480})

```

---

<div class="post-metadata">

**Author:** ![norlandrhagen](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/norlandrhagen/32/1966_2.png) [@norlandrhagen](https://discourse.pangeo.io/u/norlandrhagen)\
**Post date:** [August 20, 2024, 10:36pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/2 "2024-08-20T22:36:51Z")

</div>

Hey @vbalza 👋

Are you writing your zarr store to local storage or to cloud storage? There is usually quite a difference in speed between the two.

Here is an [older pangeo post](https://discourse.pangeo.io/t/extremly-slow-write-to-s3-bucket-with-xarray-dataset-to-zarr/2262/20) about xarray & zarr write speeds. Tons of info here on timing strategies, latency etc.

FWIW, I tried recreating a similar dataset to yours to get some timings:

```python
import xarray as xr
import numpy as np

time = np.arange(16)
lat = np.linspace(-90, 90, 721)
lon = np.linspace(-180, 180, 1440)

mean_abs_error = np.random.randint(0, 100, size=(time_dim, lat_dim, lon_dim), dtype=np.int32)

ds = xr.Dataset(
    data_vars={
        "mean_abs_error": (("time", "lat", "lon"), mean_abs_error)
    },
    coords={
        "time": time,
        "lat": lat,
        "lon": lon
    }
)
ds

```

This mock dataset is about 66MB.

Saving this dataset to local disk takes about ~1.2 seconds:

```auto
%%time
ds.to_zarr('tmp.zarr',mode='w',consolidated=True)

```

Since your dataset is chunked, you could try using dask to speed-up your write:

```python
from distributed import Client

num_workers 8 # adjust as needed
client = Client(n_workers=num_workers)
client

...
ds.to_zarr(<path.zarr>, consolidated=True)

```

Hope it helps!

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [August 20, 2024, 11:09pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/3 "2024-08-20T23:09:57Z")

</div>

10 minutes for 64MB is absurdly slow - are you sure it’s the write step that is actually taking up the time, and not the opening/loading/compute step?

See also this comment

> [@Code hangs while saving dataset to disk using .to\_netcdf()](https://discourse.pangeo.io/t/code-hangs-while-saving-dataset-to-disk-using-to-netcdf/4413/4):
>
> Are you sure this is a problem with saving to disk? Your code basically does three things: Reads data from google cloud Does the interpolation (gc.interpolation.interp\_hybrid\_to\_pressure) Writes it to local disk Because you’re using Dask and operating lazily, all three of these things happen at once when you call to\_zarr or to\_netcdf. For netcdf, it may allocate disk space for the file at the beginning, but it may still be computing the actual data for a long time after that initial step. To…

---

<div class="post-metadata">

**Author:** ![vbalza](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/vbalza/32/2884_2.png) [@vbalza](https://discourse.pangeo.io/u/vbalza)\
**Post date:** [August 20, 2024, 11:21pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/4 "2024-08-20T23:21:31Z")

</div>

@TomNicholas You’re right—it might actually be the compute speed because once I call `.compute()` upon taking the mean, the writing step speeds up. Though I’m still trying to figure out how to speed up the mean computation.

---

<div class="post-metadata">

**Author:** ![maawoo](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/maawoo/32/1972_2.png) [@maawoo](https://discourse.pangeo.io/u/maawoo)\
**Post date:** [August 22, 2024, 7:32am UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/5 "2024-08-22T07:32:38Z")

</div>

How is the ERA5 data loaded? Maybe that is somehow limiting your performance. Also, I would try again without the last re-chunk step to see if that makes a difference.

---

<div class="post-metadata">

**Author:** ![dcherian](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/dcherian/32/2235_2.png) [@dcherian](https://discourse.pangeo.io/u/dcherian)\
**Post date:** [August 22, 2024, 3:23pm UTC](https://discourse.pangeo.io/t/extremely-slow-xarray-zarr-writes/4434/6 "2024-08-22T15:23:52Z")

</div>

Interesting, can you show us all the code please (perhaps upload a reproducible notebook somewhere)?
