# Question on DASK efficiency

**URL:** <https://discourse.pangeo.io/t/question-on-dask-efficiency/2366>\
**Category:** Data\
**Created:** [March 30, 2022, 3:52pm UTC](https://discourse.pangeo.io/t/question-on-dask-efficiency/2366 "2022-03-30T15:52:57Z")\
**Posts on this page:** 2\
**Page:** 2

<div class="post-metadata">

**Author:** ![rwegener2](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rwegener2/32/887_2.png) [@rwegener2](https://discourse.pangeo.io/u/rwegener2)\
**Post date:** [April 28, 2022, 2:50pm UTC](https://discourse.pangeo.io/t/question-on-dask-efficiency/2366/21 "2022-04-28T14:50:19Z")

</div>

Hey @oceanusofi! I came across this thread and just wanted to share that I have also had a lot of trouble with that package. I thought it would speed up my processing because I could use dask, but for reasons I don’t understand I’ve found it to be considerably slower than [the standard package](https://github.com/ecjoliver/marineHeatWaves) for calculating marine heatwaves (which xmhw is built off of).

I find the standard package to be a bit more awkward to use (it’s numpy based, not xarray based), but it runs much faster and you can still parallelize the code using dask arrays. Here is a snippet of doing that with some fake data:

```auto
# Create fake input data
lat_size, long_size = 100, 100
data = da.random.random_integers(0, 30, size=(1_000, long_size, lat_size), chunks=(-1, 10, 10)) # size = (time, longitude, latitude)
time = np.arange(730_000, 731_000) # time in ordinal days

# define a wrapper for the climatology
def func1d_climatology(arr, time):
   _, point_clim = mhw.detect(time, arr)
   # return climatology
   return point_clim['seas']

# define a wrapper for the threshold
def func1d_threshold(arr, time):
   _, point_clim = mhw.detect(time, arr)
   # return threshold
   return point_clim['thresh']

# output arrays
full_climatology = da.zeros_like(data)
full_threshold = da.zeros_like(data)

climatology = da.apply_along_axis(func1d_climatology, 0, data, time=time, dtype=data.dtype, shape=(1000,))
threshold = da.apply_along_axis(func1d_threshold, 0, data, time=time, dtype=data.dtype, shape=(1000,))
climatology = climatology.compute()
threshold = threshold.compute()

```

That code runs the function twice, once for each of the outputs that I want (the climatology and the threshold values). I am almost positive that dask has a way where you don’t have to do that, but I haven’t figured it out.

I hope this is helpful!

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [April 29, 2022, 2:48pm UTC](https://discourse.pangeo.io/t/question-on-dask-efficiency/2366/22 "2022-04-29T14:48:27Z")

</div>

Thanks for sharing @rwegener2!

Has someone opened an issue on xmhw to report that the package is not working as advertised?

When she was working in our group, @hillaryscannell also developed a package for MHW tracking:

> **[GitHub - ocetrac/ocetrac: Python package to label and track unique geospatial...](https://github.com/ocetrac/ocetrac)**
>
> Python package to label and track unique geospatial features from gridded datasets - GitHub - ocetrac/ocetrac: Python package to label and track unique geospatial features from gridded datasets

[Previous page](https://discourse.pangeo.io/t/question-on-dask-efficiency/2366.md?page=1)
