# Pangeo + Globus Labs Meeting and Discussion

**URL:** <https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308>\
**Category:** Meta\
**Created:** [March 11, 2022, 7:02pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308 "2022-03-11T19:02:29Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 11, 2022, 7:02pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/1 "2022-03-11T19:02:29Z")

</div>

This thread is to organize an informal meeting between folks from Pangeo and folks from Globus Labs to discuss potential areas of collaboration and common interests around open science

### Globus Labs

[https://labs.globus.org/](https://labs.globus.org/)

> Globus Labs is a research group led by Prof. Ian Foster and Dr. Kyle Chard that spans the [Department of Computer Science](https://www.cs.uchicago.edu/) at the University of Chicago and the [Data Science and Learning Division](https://www.anl.gov/dsl) at Argonne National Laboratory. Our modest goal is to realize a world in which **all research data are reliably, rapidly, and securely accessible, discoverable, and usable**. To this end, we work on a broad range of research problems in data-intensive computing and research data management.

I have been in contact with both Ian and Ben Blaiszik, who shared that they are beginning to work on some projects in the weather / climate space. This work involves MODIS, CMIP6, and processing large volumes of data on the Argonne supercomputers. Overall I get the impression that our communities share similar aims and values around open science, so I am eager to stimulate some dialog!

A particular project of relevance is [Foundry](https://ai-materials-and-chemistry.gitbook.io/foundry/v/docs/)

> Foundry is a Python package that simplifies the discovery and usage of machine-learning ready datasets and published models in materials science and chemistry. We provide software tools that make it easy to load datasets and work with them in local or cloud environments and to perform inference using published ML models.

Ben shared these slides about some work that they have done using Foundry in materials science research workflows that are pretty inspiring.

> **[20220301-FAIRUS-extended.pdf](https://drive.google.com/file/d/1NXeHnVk5-RwiWHld990wf4eVKWv1z9oh/view)**
>
> Google Drive file.

This work has some parallels with Pangeo and Pangeo Forge in particular. From the Pangeo side, we have discussed leveraging Globus’ file transfer technology several times:

- [Transfer inputs using Globus · Issue #222 · pangeo-forge/pangeo-forge-recipes · GitHub](https://github.com/pangeo-forge/pangeo-forge-recipes/issues/222)
- [Configure Globus Connect Personal on ocean.pangeo.io · Issue #489 · pangeo-data/pangeo-cloud-federation · GitHub](https://github.com/pangeo-data/pangeo-cloud-federation/issues/489)

but have not yet managed to integrate well.

### Meeting Goals and Agenda

The goal of the meeting is _to raise mutual awareness of what the each project is doing and identify possible areas of collaboration_. With that in mind, I would suggest an agenda that looks something like this:

- Brief presentations (\< 10 min) from each group to introduce the broader aims.
- Deeper dive into specific projects (5 min presentation each)
  - Pangeo Forge
  - Foundry

- Open discussion (30 min)

If you are interested in participating in such a meeting, please fill out this poll for the week of March 21. (If this week is not good, let me know and we can try something else.)

[https://www.when2meet.com/?14931462-lMArL](https://www.when2meet.com/?14931462-lMArL)

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 14, 2022, 1:05am UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/2 "2022-03-14T01:05:39Z")

</div>

Note: _I will leave this poll up through Wed, Mar. 16 (weekly Pangeo telecon)._ After that we will pick a time and announce it here…

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 17, 2022, 5:32pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/3 "2022-03-17T17:32:48Z")

</div>

Ok the meeting is confirmed for _Wednesday, March 23, 2pm ET_.

I did not have email for most folks, here is a calendar event you can add to your schedule: [https://calendar.google.com/event?action=TEMPLATE&tmeid=NjdnbXJvdGpwYjhzMjkwbGJvNmtjNmhtYWcgcnBhQGxkZW8uY29sdW1iaWEuZWR1&tmsrc=rpa%40ldeo.columbia.edu](https://calendar.google.com/event?action=TEMPLATE&tmeid=NjdnbXJvdGpwYjhzMjkwbGJvNmtjNmhtYWcgcnBhQGxkZW8uY29sdW1iaWEuZWR1&tmsrc=rpa%40ldeo.columbia.edu)

Zoom details below. All are welcome to join

* * *

Topic: Pangeo + Globus Labs Meeting  
Time: Mar 23, 2022 02:00 PM Eastern Time (US and Canada)

Join Zoom Meeting

> **[Join our Cloud HD Video Meeting](https://columbiauniversity.zoom.us/j/96776260336?pwd=dG4xd1liQlZJb25xamgySTJRektaZz09)**
>
> Zoom is the leader in modern enterprise video communications, with an easy, reliable cloud platform for video and audio conferencing, chat, and webinars across mobile, desktop, and room systems. Zoom Rooms is the original software-based conference...

Meeting ID: 967 7626 0336  
Passcode: 016725  
One tap mobile  
+13126266799,96776260336#,\*016725# US (Chicago)  
+13462487799,96776260336#,\*016725# US (Houston)

Dial by your location  
+1 312 626 6799 US (Chicago)  
+1 346 248 7799 US (Houston)  
+1 646 876 9923 US (New York)  
+1 669 900 6833 US (San Jose)  
+1 253 215 8782 US (Tacoma)  
+1 301 715 8592 US (Washington DC)  
Meeting ID: 967 7626 0336  
Passcode: 016725  
Find your local number: [Zoom International Dial-in Numbers - Zoom](https://columbiauniversity.zoom.us/u/acGRufIfuz)

Join by SIP  
[96776260336@zoomcrc.com](mailto:96776260336@zoomcrc.com)

Join by H.323  
162.255.37.11 (US West)  
162.255.36.11 (US East)  
221.122.88.195 (China)  
115.114.131.7 (India Mumbai)  
115.114.115.7 (India Hyderabad)  
213.19.144.110 (Amsterdam Netherlands)  
213.244.140.110 (Germany)  
103.122.166.55 (Australia Sydney)  
103.122.167.55 (Australia Melbourne)  
209.9.211.110 (Hong Kong SAR)  
64.211.144.160 (Brazil)  
69.174.57.160 (Canada Toronto)  
65.39.152.160 (Canada Vancouver)  
207.226.132.110 (Japan Tokyo)  
149.137.24.110 (Japan Osaka)  
Meeting ID: 967 7626 0336  
Passcode: 016725

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 17, 2022, 9:53pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/4 "2022-03-17T21:53:24Z")

</div>

Note I mistakenly said Thursday, but in fact the consensus time was **Wednesday**. The invitation has been corrected.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 23, 2022, 2:49pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/5 "2022-03-23T14:49:41Z")

</div>

I’m looking forward to this meeting today.

I’ve put an agenda up here:

> **[Notes - Pangeo + Globus Labs Meeting](https://docs.google.com/document/d/11Ec7Xj3B8IS4jf9eq0ZftAsAwhSKDFNiBhbmgI_A858/edit?usp=sharing)**
>
> | Attendees: Agenda Quick introduction to all participants (&lt; 5 min) Presentation from Globus Labs (10 min) Including Foundry Presentation from Pangeo (Ryan) (10 min) Including Pangeo Forge Open discussion of collaboration points How can...

Anyone should feel free to add any topics they want to address.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 23, 2022, 7:39pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/6 "2022-03-23T19:39:42Z")

</div>

Thanks everyone for the great meeting. Here were a couple of different collaboration ideas that we discussed.

- Continue working to enable Pangeo Forge to pull data from globus. Already in progress at [Transfer inputs using Globus · Issue #222 · pangeo-forge/pangeo-forge-recipes · GitHub](https://github.com/pangeo-forge/pangeo-forge-recipes/issues/222)
- Explore using Foundry as a way for HPC users to publish simulation data directly from their HPC site. Would be useful for cases like [Proposed Recipes for CESM2 Superparameterization Emulator · Issue #100 · pangeo-forge/staged-recipes · GitHub](https://github.com/pangeo-forge/staged-recipes/issues/100)
- Explore whether Pangeo Forge can execute recipes via [FuncX](https://funcx.org/) and / or [parsl](https://parsl-project.org/).
- Work with Foundry team to enable Xarray datasets to be loaded in Foundry.
- Help @mgrover1 and DOE folks with their CMIP6 @ Argonne project, leveraging our experience with CMIP6 in the cloud.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [March 25, 2022, 3:08pm UTC](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308/7 "2022-03-25T15:08:47Z")

</div>

I opened this issue to follow up on the Foundry-in-geoscience discussion:

> <https://github.com/MLMI2-CSSI/foundry/issues/172>
>
> This issue is a follow up to the great discussion we had at the \[Pangeo + Globus… Labs meeting\](https://discourse.pangeo.io/t/pangeo-globus-labs-meeting-and-discussion/2308) earlier this week.
> 
> Foundry seems like a great project that would fill and important niche in the geosciences: allowing scientists to publish simulation data from their globus-connected HPC systems to share with the broader world. As I understand it, Foundry is currently focused exclusively on supervised machine-learning datasets that fit into the \`intput\` / \`target\` paradigm. This certainly captures some of the needs in geosciences, but not all of them. Sometimes people just want to publish a dataset for general consumption, not specifically for ML. So my first question is \_whether there is scope in the project for more generic dataset publishing via globus?\_ If not, can you point me towards any other projects in that space? 
> 
> Even if the answer is no, I think we will still want to use Foundry to publish ML-focused data in the geosciences.
> 
> Leaving that question aside for now, here are some random thoughts on what might make Foundry useful / appealing for geoscience / ocean / weather / climate / etc. users.
> 
> \---
> 
> \### NetCDF is our data model
> 
> Foundry currently supports two data types: tablular data and hierarchical data. In geosciences, we tend to use a similar schema to distinguish between data types. However, we tend to say tabular data vs. array data. And in the geosciences, 99% of array data is encoded in the \[NetCDF data model\](https://docs.unidata.ucar.edu/netcdf-c/current/netcdf\_data\_model.htm). And probably 80% uses \[CF Conventions\](https://cfconventions.org/).
> 
> When data follow these conventions, the dataset metadata already contain fields like \`standard\_name\`, \`description\`, \`units\`, etc. etc. within the data files themselves (i.e. self-describing). Consequentially, some of the \[metadata that Foundry requires for describing datasets\](https://ai-materials-and-chemistry.gitbook.io/foundry/publishing/publishing-datasets#describing-datasets) may be redundant. So a major design question for incorporating NetCDF-type data into Foundry would be \_how to harmonize Foundry's metadata requirements with the metadata standards commonly used in geosciences via NetCDF / CF conventions.\_ Perhaps some of the Foundry metadata could be automatically discovered by examining the data.
> 
> \### Xarray is our python API
> 
> Just like Pandas provides data structures and a computational API that matches the tabular data model, \[xarray\](https://docs.xarray.dev/en/stable/) provides data structures and a computational API that matches the NetCDF data model. So just like Foundry maps tabular data to be opened in Pandas, we would want to \_map NetCDF-style data to be opened in Xarray\_.
> 
> Xarray also serves as swiss-army knife of files formats. Similarly to Pandas, It can read \[dozens of different data formats\](https://docs.xarray.dev/en/stable/user-guide/io.html) and load them all into the same data model. This is a huge cognitive boost for scientists, who can then easily write analysis code that interoperates with any of these formats.
> 
> \### Our files are NetCDF, \[Cloud Optimized\] GeoTIFF or Zarr
> 
> For the most part, I imagine scientists wanting to share data via Foundry will be using one of three formats.
> \- \*\*NetCDF\*\*, often with hundreds / thousands of files in folder, arranged sequentially along some dimension such as time. (Note that Xarray \[can open these collections\](https://docs.xarray.dev/en/stable/user-guide/io.html#reading-multi-file-datasets) as single Dataset object, leveraging Dask. I see this feature has come up elsewhere in Foundry: #52. Under the hood, NetCDF4 files are HDF5, but as a rule we almost never open them directly with h5py, preferring to use the NetCDF model instead.
> \- \*\*Cloud Optimized GeoTIFF\*\* (or \[COG\](https://www.cogeo.org/)) is the predominant format for \_geospatial imagery data\_. Each file holds a single multiscale image. The files are often catalogued using \[Spatio-temporal asset catalog\](https://stacspec.org/). The \[Radiant Earth MLHub\](https://mlhub.earth/) holds lots of remote sensing ML training datasets and shares similar aims to Foundry, so that might be a worthwhile project to investigate.
> \- \*\*\[Zarr\](https://zarr.readthedocs.io/)\*\* - Zarr is a hierarchical format similar to HDF5. However, rather than storing everything in a single file, Zarr explodes the data into many individual files / objects, including separate json metadata. This has \[advantages and disadvantages\](https://par.nsf.gov/servlets/purl/10177969) depending on your context. (HPC filesystems tend not to love lots of small files; cloud object stores work great with it). One clear advantage is that it is trivial to append to these datasets along any dimensions. So where we might have 1000 individual netCDF files, we would only have one single Zarr store. When sharing Zarr data via Foundry, we would probably just share a single Zarr group (potentially comprised of thousands of individual files, which have no meaning outside the context of the Zarr group / array). Like with NetCDF, there is a question of redundancy between the metadata stored in Zarr itself vs. input to Foundry.
> 
> \---
> 
> If you're interested in moving this forward, we would be happy to serve as guinea pigs 🐹 in exploring Foundry for geosciences. We have many datasets sitting on globus-connected systems that we would like to publish.
> 
> cc @mgrover1, @scollis, @cisaacstern, @jbusecke,
