# CMIP6 Zarr datasets on AWS — useful for interactive exploration?

**URL:** <https://discourse.pangeo.io/t/cmip6-zarr-datasets-on-aws-useful-for-interactive-exploration/1542>\
**Category:** Data\
**Created:** [June 10, 2021, 10:31am UTC](https://discourse.pangeo.io/t/cmip6-zarr-datasets-on-aws-useful-for-interactive-exploration/1542 "2021-06-10T10:31:45Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![guigrpa](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/guigrpa/32/1028_2.png) [@guigrpa](https://discourse.pangeo.io/u/guigrpa)\
**Post date:** [June 10, 2021, 10:31am UTC](https://discourse.pangeo.io/t/cmip6-zarr-datasets-on-aws-useful-for-interactive-exploration/1542/1 "2021-06-10T10:31:45Z")

</div>

Hi there,

I’m working on the development of new, interactive ways to exploit cloud-based Zarr datasets, and obviously got very excited when a lot of CMIP6 datasets were released on Amazon (AWS) in that format (cmip6-pds bucket).

However, looking closely into the ASDI CMIP6 buckets, I’ve seen that the datasets are only chunked across the time dimension and are quite large (10-100 MB). This makes fast, interactive analysis very hard, since (as far as I know) obtaining a time series for a single location would require downloading all chunks, even in Python with xarray; from JavaScript, it would be even more of a show-stopper.

Maybe we’re missing something? We thought about range header requests (à la COG), but (1) I’m not sure they’re supported in Zarr or Zarr libraries; and (2) for geographic subsetting (say we want a small AOI) we would still be sending a lot of requests (one for each lat coordinate, since they won’t be contiguously stored in the file).

Any ideas? Thanks in advance!

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [June 10, 2021, 12:11pm UTC](https://discourse.pangeo.io/t/cmip6-zarr-datasets-on-aws-useful-for-interactive-exploration/1542/2 "2021-06-10T12:11:50Z")

</div>

Thanks for this interesting and useful question @guigrpa – and welcome to the forum!

I’ll try to write a detailed response within a few days. In the meantime, this other post contains many points that are relevant to your question:

> [@Best practices to go from 1000s of netcdf files to analyses on a HPC cluster?](https://discourse.pangeo.io/t/best-practices-to-go-from-1000s-of-netcdf-files-to-analyses-on-a-hpc-cluster/588):
>
> What? 8759 netcdf files totalling 17TB of HYCOM ocean model (u,v) velocity data at two depth levels (and bottom velocity), at hourly time steps. So the total data arrays of interest are [9000 (X) by 7055 (Y) by 8759 (time) by 2 (Depth)] for both u and v; Using nco tools, I can reduce these files as an example to 365 netcdf files, totaling 4.3TB, one file per day (24 hourly steps), for one depth level only, so data arrays are 9000 (X) by 7055 (Y) by 8759 (time) for the u,v components. How? Us…
