# Cloud-optimized Eulerian+Lagrangian dataset freely accessible

**URL:** <https://discourse.pangeo.io/t/cloud-optimized-eulerian-lagrangian-dataset-freely-accessible/4496>\
**Category:** News & Announcements\
**Created:** [September 11, 2024, 2:20pm UTC](https://discourse.pangeo.io/t/cloud-optimized-eulerian-lagrangian-dataset-freely-accessible/4496 "2024-09-11T14:20:02Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![selipot](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/selipot/32/359_2.png) [@selipot](https://discourse.pangeo.io/u/selipot)\
**Post date:** [September 11, 2024, 2:20pm UTC](https://discourse.pangeo.io/t/cloud-optimized-eulerian-lagrangian-dataset-freely-accessible/4496/1 "2024-09-11T14:20:02Z")

</div>

Hello Pangeo community!

I wanted to share with you what I think is an exciting freely-accessible, cloud-optimized, dataset: [HYCOM-OceanTrack](https://registry.opendata.aws/hycom-global-drifters/)! This dataset comprises sea surface height and near-surface velocity **at hourly time steps on an Eulerian grid from a global [HYCOM](https://www.hycom.org), 1-year, 1/25 degree** simulation resolving the ocean general circulation and the tides. In addition, it comes with nearly **13M Lagrangian trajectories of particles** released in the Eulerian velocity fields thanks to [Ocean Parcels](https://oceanparcels.org/#developmentstatus).

This dataset is made available as zarr archived in an AWS S3 bucket thanks to the [AWS Open Data program](https://registry.opendata.aws) so you can open it in a few lines of code. The dataset is described in details in a recent publication ([Elipot et al. 2024b](https://doi.org/10.1038/s41597-024-03813-z) in Scientific Data) and we provide [python tutorial notebooks](https://github.com/selipot/hycom-oceantrack) to get you started. Some of the tools for analyzing the Lagrangian data come from the [clouddrift package](https://github.com/Cloud-Drift/clouddrift) which we describe in a recent publication in the Journal of Open Source Software ([Elipot et al. 2024a](https://joss.theoj.org/papers/10.21105/joss.06742)).

We have a couple of papers in the pipeline analyzing these data but we believe in open science and open data so we are sharing these data now. Please use them and let us know. We hope this effort contributes to democratize numerical ocean data. All this funded and supported by AWS, NSF, ONR, and the University of Miami.

It was a big effort to produce the Lagrangian trajectories (thank you to the Parcels community and @erikvansebille in particular) and to transform the Eulerian data into a cloud-optimized format (thanks @rabernat et al. for the [rechunker](https://rechunker.readthedocs.io/en/latest/)). Some of you may remember the most epic discussion from my post about 4 years ago … 👇

> [@Best practices to go from 1000s of netcdf files to analyses on a HPC cluster?](https://discourse.pangeo.io/t/best-practices-to-go-from-1000s-of-netcdf-files-to-analyses-on-a-hpc-cluster/588):
>
> What? 8759 netcdf files totalling 17TB of HYCOM ocean model (u,v) velocity data at two depth levels (and bottom velocity), at hourly time steps. So the total data arrays of interest are [9000 (X) by 7055 (Y) by 8759 (time) by 2 (Depth)] for both u and v; Using nco tools, I can reduce these files as an example to 365 netcdf files, totaling 4.3TB, one file per day (24 hourly steps), for one depth level only, so data arrays are 9000 (X) by 7055 (Y) by 8759 (time) for the u,v components. How? Us…

Ok I am done with these advertisements, let me know if you have any question!

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [September 11, 2024, 2:47pm UTC](https://discourse.pangeo.io/t/cloud-optimized-eulerian-lagrangian-dataset-freely-accessible/4496/2 "2024-09-11T14:47:55Z")

</div>

This is fantastic Shane! Thanks so much for sharing this amazing dataset. It’s so cool to see how your lab has tackled this extremely challenging data problem over the past years. Congrats on the release. 🚀
