# High resolution time series; open\_zarr question

**URL:** <https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603>\
**Category:** Science\
**Created:** [May 11, 2020, 5:23pm UTC](https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603 "2020-05-11T17:23:21Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![lsetiawan](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/lsetiawan/32/515_2.png) [@lsetiawan](https://discourse.pangeo.io/u/lsetiawan)\
**Post date:** [May 11, 2020, 5:23pm UTC](https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603/1 "2020-05-11T17:23:21Z")

</div>

Hey all,

I have a question about xarray.open\_zarr. My understanding is that that for timeseries data saved in zarr format, when opened xarray reads the .zmetadata and the contents of the time variable and load those into memory.

My question is, I have a very high density data (seconds) resolution for 5 - 6 years. For some, this can result to time variable of size 3GB+. Is there a way where I can distribute this initial read to dask workers? Or is there a known way to only grab a smaller subset of the data at open\_zarr?

Thanks in advance. Any help/comments are much appreciated.

Best,  
Don

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [May 13, 2020, 2:20pm UTC](https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603/2 "2020-05-13T14:20:59Z")

</div>

You can disable dask chunking as follows

```python
ds = xr.open_zarr(path_to_store, chunks=False)

```

Then you can subset and apply chunking manually. However, it sounds like your problem is that your _coordinate variable_`time` is too large. Xarray will always read the coordinate variable eagerly (i.e. directly into memory) in order to create an index. This is a known issue with xarray

> <https://github.com/pydata/xarray/issues/1094>
>
> (Follow-up of discussion here #1024 (comment)).
> xarray + dask.array successfully enable out-of-core computation for very large variables that doesn't fit in memory....

Until that issue is fixed, you have two options for workaround.

- Don’t use xarray. Open the zarr arrays directly using dask. You will lose indexing capabilities and other fancy xarray stuff, but you may be able to do what you need.
- Drop the time coordinate from the dataset before opening in xarray, so that you don’t need to generate an index for time. (You will lose time-indexing functions.)

---

<div class="post-metadata">

**Author:** ![lsetiawan](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/lsetiawan/32/515_2.png) [@lsetiawan](https://discourse.pangeo.io/u/lsetiawan)\
**Post date:** [May 13, 2020, 4:09pm UTC](https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603/3 "2020-05-13T16:09:48Z")

</div>

Thank you @rabernat. That really helps. I can’t quite getaway from having a smaller time variable. The sensor that I’m dealing with outputs at a high sampling rate. But, I will try your workaround suggestions.

---

<div class="post-metadata">

**Author:** ![NickMortimer](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/nickmortimer/32/438_2.png) [@NickMortimer](https://discourse.pangeo.io/u/NickMortimer)\
**Post date:** [July 2, 2020, 9:02am UTC](https://discourse.pangeo.io/t/high-resolution-time-series-open-zarr-question/603/4 "2020-07-02T09:02:40Z")

</div>

I don’t know if this would work but could you create a multi-level index eg. by day have a fixed index tile for the second of the day, grid the data onto a second of the day grid? and if you had multiple sensors you could stack them in this tile.

I do something similar extracting data from drone images into zarr all the tiles have the same dimensions in meters and then they have a time for each tile.

You could do hourly tiles of the data? that would reduce your master make the master index 3600 times smaller and depending on how you want to use the data e.g. get me an hourly mean! might make things work well
