# Go multi regional with Dask (AWS)

**URL:** <https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037>\
**Category:** Cloud\
**Created:** [January 4, 2023, 12:09pm UTC](https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037 "2023-01-04T12:09:29Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![oconpa](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/oconpa/32/1997_2.png) [@oconpa](https://discourse.pangeo.io/u/oconpa)\
**Post date:** [January 4, 2023, 12:09pm UTC](https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037/1 "2023-01-04T12:09:29Z")

</div>

OpenSource project I’ve been heavily involved in.  
Some cool things it does. Reduce data replication/minimise data movement, Integrate with Dask, Cloud solution, Dask across AWS Regions & Jupyter Notebook interface

[Code](https://github.com/aws-samples/distributed-compute-on-aws-with-cross-regional-dask)

---

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [January 4, 2023, 2:41pm UTC](https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037/2 "2023-01-04T14:41:44Z")

</div>

Very cool. In [distributed-compute-on-aws-with-cross-regional-dask/ux\_notebook.ipynb at 836dfb4678b83167f41ccee311e3bcc161efac46 · aws-samples/distributed-compute-on-aws-with-cross-regional-dask · GitHub](https://github.com/aws-samples/distributed-compute-on-aws-with-cross-regional-dask/blob/836dfb4678b83167f41ccee311e3bcc161efac46/lib/SagemakerCode/ux_notebook.ipynb), how where does `dask_worker_pools` come from? I couldn’t find it.

When we did [GitHub - pangeo-data/multicloud-demo: Notebooks and infrastructure for Earthcube2020: Multi-Cloud workflows with Pangeo and Dask Gateway](https://github.com/pangeo-data/multicloud-demo) a few years back, a difficulty was that the user had to know where the data lived and was responsible for ensuring that the “right” Dask workers were used for each computation. So we had things like

```auto
era5_tp_hist_ = gcp_client.compute(tp_hist, retries=5)
lens_hist_ = aws_client.compute(lens_hist)

```

i.e. a region-specific client, and users are explicitly computing results with that client. It’d be fun (and extremely challenging, I think) to use dataset metadata to inform where compute resources should be created, and which ones should be used for a particular stage of the computation.

---

<div class="post-metadata">

**Author:** ![oconpa](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/oconpa/32/1997_2.png) [@oconpa](https://discourse.pangeo.io/u/oconpa)\
**Post date:** [January 4, 2023, 3:19pm UTC](https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037/3 "2023-01-04T15:19:58Z")

</div>

Good call Tom. We’re using this library [GitHub - gjoseph92/dask-worker-pools: Assign tasks to pools of workers in dask](https://github.com/gjoseph92/dask-worker-pools) by gjoseph92 to selectively tell Dask which pool of workers to run on. It sounds like your previous project did a version of this by using the resource tag mechanism in Dask.

For the region location piece yes that was an annoyance having to select the region you wanted to run so how we bypassed this was by indexing metadata for each of the datasets we wanted to connect into OpenSearch as an index.  
You configure these datasets before launching the stack [here](https://github.com/aws-samples/distributed-compute-on-aws-with-cross-regional-dask/blob/836dfb4678b83167f41ccee311e3bcc161efac46/bin/variables.ts#L16). You can customise this variable workers to include more datasets, in more regions etc.  
The CDK package will set up some sync jobs to keep the index up to date (currently a daily refresh, but I think this could be optimised). You just need to query the index of the data you want, and the backend should figure out the region to query.

---

<div class="post-metadata">

**Author:** ![oconpa](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/oconpa/32/1997_2.png) [@oconpa](https://discourse.pangeo.io/u/oconpa)\
**Post date:** [January 4, 2023, 8:52pm UTC](https://discourse.pangeo.io/t/go-multi-regional-with-dask-aws/3037/5 "2023-01-04T20:52:29Z")

</div>

This notebook has an example in the repo. Search for function query\_nc and you’ll see how you can abstract the region. Turn this into a library for Jupyter and you don’t need to worry

File: lib/SagemakerCode/get\_historical\_data.ipynb
