# Get Involved in Pangeo: Entry Points for New Contributors

**URL:** <https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643>\
**Category:** Meta\
**Created:** [June 10, 2020, 1:43pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643 "2020-06-10T13:43:57Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [June 10, 2020, 1:43pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/1 "2020-06-10T13:43:57Z")

</div>

The success of Pangeo derives from the diversity of our contributors. We have successfully assembled a community that crosses traditional disciplinary boundaries, and this has enabled us to do some innovative stuff. However, as the project has grown, our activities have sprawled across dozens of GitHub repos, making it hard to identify what needs to be done and where new contributors can have an impact.

In addition to disciplinary diversity, we also must continue to tackle other dimensions of diversity, particularly gender and race. I’m proud of the steps our community has taken in this direction. The first paragraph of our [code of conduct](https://github.com/pangeo-data/governance/blob/master/conduct/code_of_conduct.md) reads

> We strive to be a community that welcomes and supports people of all backgrounds and identities. This includes, but is not limited to, members of any race, ethnicity, culture, national origin, color, immigration status, social and economic class, educational level, sex, sexual orientation, gender identity and expression, age, physical appearance, family status, technological or professional choices, academic discipline, religion, mental ability, and physical ability.

Being welcoming is a first step. But we must do more to actively recruit and support diverse contributors to Pangeo. This will benefit our project of course, but it is also a concrete action we can take to combat systematic racism. (See the [ShutdownSTEM post](https://discourse.pangeo.io/t/shutdownstem-and-strike4blacklives-on-wednesday-june-10/638) for more context.) Let’s use Pangeo as a vehicle to help members of underrepresented groups build their skills and gain recognition in the field of geoscience / big data / software engineering.

**Let’s use this thread as a place to collect potential projects for new contributors.**  
Let’s also collect information about internships, fellowships, etc. that can provide paid support to such contributors. While some people may be able to volunteer, we should not assume that everyone has this privilige.

## Template

Please try to use this template for all posts.

* * *

```markdown
# Project Title

link to GitHub issue (recommended)

## Description

One or two paragraph description of the project.

## Required Skills

What technical skills are needed in order to contribute? For example
- Basic python programming
- Some familiarity with kubernetes

## Mentors

(All projects need at least one mentor who is willing to help out the new contributors.)

- Name | Email Address

```

## For Potential Contributors

Please email the mentor to express interest in a project and learn more about how to get started.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [June 10, 2020, 2:13pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/2 "2020-06-10T14:13:21Z")

</div>

# Matrix of Kubespawner `profile_list` Options

> <https://github.com/jupyterhub/kubespawner/issues/307>
>
> I would like to give uses an option to select independently between machine options (e.g. cpu\_limit) and notebook image.
> I imagine that...

## Description

Our cloud-based Jupyter hubs allow users to choose among different options for the environment in which the notebook server will run. For example, on [ocean.pangeo.io](http://ocean.pangeo.io), we see

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/1X/d885d949aa5a6c02c54436ef5fa9f98a6e9fdce5.png)

These options, called `profile_list` are passed to kubespawner (see [docs](https://jupyterhub-kubespawner.readthedocs.io/en/latest/spawner.html?highlight=profile_list)). They include both hardware (CPU, memory, etc.) and software (specifically the docker image to use). _We would like to be able to separate the hardware part from the software part._ This would require making some changes to the kubespawner package to enable more flexible configuration of profiles.

## Required Skills

- Intermediate python programming
- Basic HTML

## Mentors

- Ryan Abernathey | rpa@ldeo.columbia.edu

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [June 10, 2020, 2:25pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/3 "2020-06-10T14:25:31Z")

</div>

# Contribute Example Notebooks to Pangeo Gallery

[http://gallery.pangeo.io/contributing.html](http://gallery.pangeo.io/contributing.html)

Pangeo gallery is our new approach to sharing reproducible scientific content in the cloud. We are always looking for more examples of how to apply Pangeo tools (e.g. Xarray, Dask, etc.) to real-world scientific problems. If you already use these tools, creating an example gallery is a great way to get started as a new contributor.

## Required Skills

- Some domain scientific knowledge (e.g. oceanography, atmospheric science)
- Basic scientific programming
- Familiarity with Jupyter notebooks
- Comfortable working with git / github

## Mentors

- Ryan Abernathey | rpa@ldeo.columbia.edu

---

<div class="post-metadata">

**Author:** ![rsignell](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rsignell/32/447_2.png) [@rsignell](https://discourse.pangeo.io/u/rsignell)\
**Post date:** [June 15, 2020, 1:10pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/4 "2020-06-15T13:10:44Z")

</div>

# Create an app for viewing terrain data using Xarray-spatial

[Xarray-spatial](https://github.com/makepath/xarray-spatial) is a new high performance package for raster-based spatial analysis for Python. The USGS has a large collection of terrain data in raster format, and is pushing this data to AWS in COG format ([example here](https://www.cogeo.org/map/#/url/https%3A%2F%2Fprd-tnm.s3.amazonaws.com%2FStagedProducts%2FElevation%2F1%2FTIFF%2Fn42w071%2FUSGS_1_n42w071.tif/center/-70.7041,41.6089/zoom/11)). This project would build a dashboard for exploring terrain data in Python using xarray-spatial and [Panel](https://panel.holoviz.org), a high-level app and dashboarding solution for Python.

![composite_map](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/1X/196836a93c69455aea03daf1dbffc36c948cf798.gif)

## Required Skills

- Basic knowledge of Python
- Willingness to work on a cool project! 🕶

## Mentors

- Rich Signell | [rsignell@usgs.gov](mailto:rsignell@usgs.gov)

---

<div class="post-metadata">

**Author:** ![bradyrx](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/bradyrx/32/94_2.png) [@bradyrx](https://discourse.pangeo.io/u/bradyrx)\
**Post date:** [June 23, 2020, 4:31pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/5 "2020-06-23T16:31:34Z")

</div>

# Contribute to the `climpred` package for analyzing climate predictions

`climpred` is a package that uses pangeo-supported software like `xarray` and `dask` to make evaluating climate predictions easier. Many institutions are running climate models similar to a weather model to predict the Earth system anywhere from 2 weeks to decades in advance. These projects produce massive datasets and require users to tediously write code to assess how well the forecasts did. `climpred` automates a lot of the analysis (like aligning forecast times with real-world times and computing statistical metrics) so that users can get right to answering the scientific questions they care about.

## What to contribute

> **[bradyrx/climpred](https://github.com/bradyrx/climpred/issues)**
>
> :earth\_americas: an xarray wrapper for analysis of ensemble forecast models for climate prediction :earth\_africa: - bradyrx/climpred

Look for tags “Help Wanted”, “ASP Projects”, or “Good First Issue”

## Required Skills

What technical skills are needed in order to contribute?

- Intermediate python programming
- Basic knowledge of git (Although see [https://climpred.readthedocs.io/en/stable/contributing.html](https://climpred.readthedocs.io/en/stable/contributing.html) for instructions on how to contribute)
- Some domain-specific knowledge of forecasting/forecast evaluation

Note that Aaron and I are eager to mentor anyone who is a first time contributor. The code review is friendly and you’ll learn a lot from the process. Feel free to email us if you have ideas or just want to help out and want some guidance on where to get started.

## Mentors

- Riley Brady (PhD candidate at CU Boulder)  
riley.brady@colorado.edu  
[https://www.github.com/bradyrx](https://www.github.com/bradyrx)

- Aaron Spring (PhD candidate at MPI in Hamburg, Germany)  
@aaronspring  
[aaron.spring@mpimet.mpg.de](mailto:aaron.spring@mpimet.mpg.de)  
[https://www.github.com/aaronspring](https://www.github.com/aaronspring)

---

<div class="post-metadata">

**Author:** ![nicholaskgeorge](https://avatars.discourse-cdn.com/v4/letter/n/dec6dc/32.png) [@nicholaskgeorge](https://discourse.pangeo.io/u/nicholaskgeorge)\
**Post date:** [July 15, 2020, 3:32pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/6 "2020-07-15T15:32:04Z")

</div>

## Making Terrain Data Viewer with Xarray

Hello everyone! My name is Nicholas George and I am currently working on a short internship with USGS to create an interactive viewer for terrain data using Xarray in tandem with Panel and hvPlot. Our goal is to use these tools to make a viewing window for COG data which is both interactive and informative. If you are interested in seeing what we have done, the GitHub repo can be found [here](https://github.com/USGS-CMG/dem-dashboard)

---

<div class="post-metadata">

**Author:** ![cgentemann](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/cgentemann/32/530_2.png) [@cgentemann](https://discourse.pangeo.io/u/cgentemann)\
**Post date:** [August 27, 2020, 9:11pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/7 "2020-08-27T21:11:21Z")

</div>

# Test cloud optimized data formats for swath (L2) satellite data

## Description

As NASA and NOAA data are moving to the cloud, what is the best cloud-optimized format for swath data? NASA has explored different formats and written a [report](https://ntrs.nasa.gov/citations/20200001178) that presents some viable options for swath data. For this project, we will stage a sample of the MODIS L2P data (from PO.DAAC) on Pangeo, transform it to a couple different formats (eg. Zarr, cloud-optimized HDF) and test access and analysis times for a few different likely patterns of analysis, such as collocation with random points that are globally distributed, finding all data within a bounding box, etc.

## Required Skills

What technical skills are needed in order to contribute? For example

- Basic-Intermediate python programming (Xarray, matplotlib)

## Mentors

- Chelle Gentemann, [cgentemann@faralloninstitute.org](mailto:cgentemann@faralloninstitute.org)

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [September 2, 2020, 2:33pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/8 "2020-09-02T14:33:51Z")

</div>

# NGINX Proxy for Cloud Storage

## Description

In Pangeo, we love using Zarr as a cloud optimized data format, and are trying to push data providers to start serving data in Zarr directly from cloud object storage. So far, our Zarr cloud datasets have been either

- Totally public (no authentication required at all)
- Requester pays (requiring credentials from the cloud provider, but any credentials will do)

However, many data providers (e.g. NASA) want to have more fine-grained control over who can access which datasets. They also want logging of who is downloading their data. While this is theoretically possible using IAM roles, that is probably not scalable. These services can have thousands of users, and creating a unique identity for each one using the cloud-provider’s identity system is probably not feasible. It also might present security issues.

_So we need some way to manage and restrict access to the underlying object store using an external (e.g. oauth2) identity provider._

My idea is to use NGINX for this. NGINX can pretty easily be configured to proxy cloud storage. Here’s an example:

> **[presslabs/gs-proxy](https://github.com/presslabs/gs-proxy)**
>
> Simple nginx proxy to google cloud storage. Contribute to presslabs/gs-proxy development by creating an account on GitHub.

If we put the NGINX proxies inside an auto-scaling kubernetes cluster, we should be able to scale up and down in response to load to avoid excessive compute charges.

What we would need to add to this would be JWT authentication support. Based on the NGINX docs, that seems relatively straightforward

> **[NGINX Docs | Setting up JWT Authentication](https://docs.nginx.com/nginx/admin-guide/security-controls/configuring-jwt-authentication/)**
>
> Control access using JWT authentication.

Logging cloud also be configured to track usage and downloads. Here’s a diagram of how it might work.

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/1X/3cd15ded2b80cff1432507f2883cdd5cdc0a67b3.png)

## Required Skills

Perhaps I’m underestimating the difficulty, but I feel like this would be a \< 1 day project for the right person. We’re looking for someone who understands

- NGINX configuration
- Oauth2 and JWT
- Kubernetes

## Mentors

- Ryan Abernathey | rpa@ldeo.columbia.edu

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [September 2, 2020, 9:02pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/9 "2020-09-02T21:02:11Z")

</div>

On this last topic, at our latest meeting, people suggested signed URLs, which could solve the problem more elegantly:

- [https://cloud.google.com/storage/docs/access-control/signed-urls](https://cloud.google.com/storage/docs/access-control/signed-urls)
- [https://docs.aws.amazon.com/AmazonS3/latest/dev/ShareObjectPreSignedURL.html](https://docs.aws.amazon.com/AmazonS3/latest/dev/ShareObjectPreSignedURL.html)

---

<div class="post-metadata">

**Author:** ![jukent](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/jukent/32/297_2.png) [@jukent](https://discourse.pangeo.io/u/jukent)\
**Post date:** [September 24, 2020, 4:22pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/10 "2020-09-24T16:22:24Z")

</div>

@bradyrx Would you be interested in turning this into a SCIParCS project ([Internship Projects](https://discourse.pangeo.io/t/internship-projects/855/2))? We can discuss more if you are!

---

<div class="post-metadata">

**Author:** ![bradyrx](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/bradyrx/32/94_2.png) [@bradyrx](https://discourse.pangeo.io/u/bradyrx)\
**Post date:** [September 24, 2020, 4:36pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/11 "2020-09-24T16:36:39Z")

</div>

@jukent, yes that would be great. We will be releasing our next version in the next couple weeks as well as a paper to JOSS so it will be in a great place for contributing next summer for SCIParCS. We also just migrated over to `pangeo-data`.

@aaronspring would have to serve as a mentor most likely. I’m still trying to navigate what IP law and my free time will look like at my new job. But definitely can get this spun up now and then will know in the spring about IP/time.

---

<div class="post-metadata">

**Author:** ![jukent](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/jukent/32/297_2.png) [@jukent](https://discourse.pangeo.io/u/jukent)\
**Post date:** [September 24, 2020, 4:50pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/12 "2020-09-24T16:50:07Z")

</div>

Someone at NCAR needs to be the official mentor (likely me), but we would reach out for questions and help much like any other Pangeo project. I would ask you to read the project proposal I write to make sure it is in line with your goals, then I would try to make at least some contribution to get myself spun up on the package (and I will probably need some help at this stage), and then the intern and I could be as independent or collaborate as much as need be with the uncertainty of next summer for you and @aaronspring.

The internship proposals are due October 7th, before you’ll know about your time availability, so we will write it without any promises of engagement from you. Congratulations!

---

<div class="post-metadata">

**Author:** ![bradyrx](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/bradyrx/32/94_2.png) [@bradyrx](https://discourse.pangeo.io/u/bradyrx)\
**Post date:** [September 24, 2020, 5:01pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/13 "2020-09-24T17:01:51Z")

</div>

That sounds perfect to me Julia! Let’s go for it.

---

<div class="post-metadata">

**Author:** ![jukent](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/jukent/32/297_2.png) [@jukent](https://discourse.pangeo.io/u/jukent)\
**Post date:** [September 24, 2020, 5:10pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/14 "2020-09-24T17:10:33Z")

</div>

Is [https://climpred.readthedocs.io/en/stable/contributing.html](https://climpred.readthedocs.io/en/stable/contributing.html) up to date or is there a new document over at pangeo-data?

---

<div class="post-metadata">

**Author:** ![bradyrx](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/bradyrx/32/94_2.png) [@bradyrx](https://discourse.pangeo.io/u/bradyrx)\
**Post date:** [September 24, 2020, 5:26pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/15 "2020-09-24T17:26:03Z")

</div>

Yes the docs are up-to-date with the new repo location. We have a lot of updates rolling out for the next release that are on [https://climpred.readthedocs.io/en/latest](https://climpred.readthedocs.io/en/latest) (latest, not stable). Plenty of un-addressed issues in the tracker as well: [https://github.com/pangeo-data/climpred/issues](https://github.com/pangeo-data/climpred/issues).

The CI is really good and the code base is generally good. I’m trying to work on converting the whole code base over to an inheritance-based system. I am sketching it out at [https://github.com/bradyrx/climpred\_skeleton](https://github.com/bradyrx/climpred_skeleton).

Anyways, if you end up finding the code base to be hard to understand, please let us know. Always looking for ways to make contributing easier. I feel like switching to inheritance will be hard in some ways for contributors, but should clean up a lot of redundant code that we currently have.

---

<div class="post-metadata">

**Author:** ![jukent](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/jukent/32/297_2.png) [@jukent](https://discourse.pangeo.io/u/jukent)\
**Post date:** [September 30, 2020, 10:49pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/16 "2020-09-30T22:49:19Z")

</div>

@cgentemann Would you be interested in turning this into a SCIParCS project ([Internship Projects](https://discourse.pangeo.io/t/internship-projects/855/2))? I am very familiar with MODIS satellite but don’t have the most experience with the cloud so I could use the expertise of a co-mentor. We can discuss more if you are interested.

---

<div class="post-metadata">

**Author:** ![cgentemann](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/cgentemann/32/530_2.png) [@cgentemann](https://discourse.pangeo.io/u/cgentemann)\
**Post date:** [September 30, 2020, 11:56pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/17 "2020-09-30T23:56:51Z")

</div>

@jukent, sure! Let me know what you need from me to do this.

---

<div class="post-metadata">

**Author:** ![Michael\_Sumner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/michael_sumner/32/1608_2.png) [@Michael\_Sumner](https://discourse.pangeo.io/u/Michael_Sumner)\
**Post date:** [September 8, 2024, 9:00pm UTC](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/18 "2024-09-08T21:00:57Z")

</div>

Did this work or anything related go ahead? I’m trying to find similar followups 🙏
