# High unmanaged memory warning with Dask when measuring NetCDF read throughput from Google Cloud Storage

**URL:** <https://discourse.pangeo.io/t/high-unmanaged-memory-warning-with-dask-when-measuring-netcdf-read-throughput-from-google-cloud-storage/2495>\
**Category:** Data\
**Created:** [May 31, 2022, 2:44pm UTC](https://discourse.pangeo.io/t/high-unmanaged-memory-warning-with-dask-when-measuring-netcdf-read-throughput-from-google-cloud-storage/2495 "2022-05-31T14:44:07Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [June 3, 2022, 1:13pm UTC](https://discourse.pangeo.io/t/high-unmanaged-memory-warning-with-dask-when-measuring-netcdf-read-throughput-from-google-cloud-storage/2495/5 "2022-06-03T13:13:11Z")

</div>

Ok, I think the answer is straightforward. This file `ETOPO1_Ice_g_gmt4.nc` is a netCDF3 file. NetCDF3 files do not support internal chunking of the data. All data variables are stored as simple flat binary data in C-order. The only way to read it over the internet is with `engine='scipy'`. This engine does not support lazy loading of data from fsspec filesystems, as documented in this issue:

> <https://github.com/pangeo-forge/pangeo-forge-recipes/issues/361>
>
> I am trying to create a recipe with some netCDF3 files based on this example - h…ttps://discourse.pangeo.io/t/dask-xarray-and-swap-memory-polution-on-local-linux-cluster/2453
> 
> I have discovered an issue with the way netCDF3 / scipy / fsspec interact. The main issue is the \[scipy netcdf\](https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.io.netcdf.netcdf\_file.html) function has an \`mmap\` option which provides lazy access which only works with local files. It seems that otherwise we are eagerly loading the data.  
> 
> I profiled the following cases using fil.
> 
> \### Local file
> 
> \`\`\`python
> %%filprofile
> url = "gcs://leap-scratch/rabernat/ERA5\_HiRes\_Hourly/cache/de66f7c4c230a196f5fad34f35355df2-https\_cluster.klima.uni-
> with fsspec.open("simplecache::" + url, "rb") as fp:
> ds = xr.open\_dataset(fp.name, engine='scipy', backend\_kwargs={'mmap': True})
> \`\`\`
> 
> \<img width="1051" alt="image" src="https://user-images.githubusercontent.com/1197350/167677145-ea2fa76b-2828-41e2-b00a-61a497757e52.png"\>
> 
> 
> \## fsspec local filesystem
> 
> \`\`\`python
> %%filprofile
> with fsspec.open("simplecache::" + url, "rb") as fp:
> ds = xr.open\_dataset(fp, engine='scipy', backend\_kwargs={'mmap': False})
> \`\`\`
> 
> \<img width="1054" alt="image" src="https://user-images.githubusercontent.com/1197350/167677348-fd8a7df1-db68-4549-b451-200291da78a6.png"\>
> 
> 
> \### Using gcsfs
> 
> \`\`\`python
> %%filprofile
> bremen.de\_fmaussion\_teaching\_climate\_dask\_exps\_hires\_hourly\_surf\_era5\_hires\_hourly\_tp\_2000\_01.nc"
> with fsspec.open(url, "rb") as fp:
> ds = xr.open\_dataset(fp, engine='scipy', backend\_kwargs={'mmap': False})
> \`\`\`
> 
> \<img width="1066" alt="image" src="https://user-images.githubusercontent.com/1197350/167676904-fc71fb8f-5313-4d96-bd1b-c90be3480704.png"\>

The chunks that get created when you call `dsa.from_array(da)` are spurious and not helpful. Every time you try to compute a single chunk, the entire array has to be read. So since you have 70 chunks, you will use 70x the memory of the original array.

I would retry this exercise with a NetCDF4 file with appropriately configured internal chunks and see if you do any better.

---

_[View the full topic](https://discourse.pangeo.io/t/high-unmanaged-memory-warning-with-dask-when-measuring-netcdf-read-throughput-from-google-cloud-storage/2495)._
