# Pangeo Showcase: "insitubatch: Streaming ML Batches from Cloud Zarr Without Resharding" (September 16, 12 PM EDT / 16:00 UTC)

**URL:** <https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812>\
**Category:** Pangeo Showcase\
**Created:** [September 8, 2026, 2:09pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812 "2026-09-08T14:09:23Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![maxrjones](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/maxrjones/32/3190_2.png) [@maxrjones](https://discourse.pangeo.io/u/maxrjones)\
**Post date:** [September 8, 2026, 2:09pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/1 "2026-09-08T14:09:23Z")

</div>

[![](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/8/8bd0f05e944bcfd52188baf91b416ce38f7be716.jpeg "insitubatch: Streaming ML Batches from Cloud Zarr Without Resharding") ](https://www.youtube.com/watch?v=LSXxZxPZtPs)

**Title** : “insitubatch: Streaming ML Batches from Cloud Zarr Without Resharding”  
**Speaker** : David Stuebe  
**When:** Wednesday, September 16, 2026 at 12 PM EDT (2026-09-16T16:00:00Z)  
**Where:** [Launch Meeting - Zoom](https://numfocus-org.zoom.us/j/83234473026?pwd=mgM3uAuru1AE1aFPlRjW2dAAAyo8IY.1)  
**Details:**

Join us next Wednesday, September 16 at 12 PM EDT for the kickoff of the fall Pangeo showcase series!

**Abstract** :

Between a dataset small enough to hold in memory and one worth building a purpose-built ETL pipeline for, there is a wide middle: archives too large to load, not yours to rewrite, or read in ways that keep changing. The usual move is to reshard into a sample-oriented format, a full ETL copy that throws away the archive’s chunk locality, must be rebuilt as the archive grows, and only pays for itself if you run the same job over the same data many times. Keeping the data in place runs into the loader: PyTorch’s `DataLoader` puts parallelism in worker processes, so each carries its own cache and IO budget, and a chunk feeding four workers is fetched and decoded four times.

The IO is no longer the hard part: obstore and Icechunk over Zarr v3’s async store already saturate the NIC. `insitubatch` is the loader-orchestration layer on top of it, reading any zarr `Store`, whether obstore, fsspec or Icechunk backs it. It plans reads chunk-first, so Python work scales with the **chunks a batch touches, not the samples it contains** , and one async event loop streams them under a single concurrency budget into a bounded pool that is residency tier and decode-once cache at once. A stored chunk is fetched and decoded exactly once, however many samples, batches or epochs reference it. Splits and shuffle live in coordinate space over the existing store, so there is no second copy of the archive, and memory is a budget you set rather than the working set.

The same ETL amortization argument applies to process parallelism: a worker pool earns back its startup over a long training run but never over a single inference pass. With no processes to launch, insitubatch reaches its first batch while a worker stack is still starting, which makes it as much an inference loader as a training loader. Handoff to PyTorch, JAX or TensorFlow is a thin DLPack adapter, and the sample axis is a role rather than a fixed dimension, so one engine spans domains: ERA5 forecasting, OME-NGFF microscopy segmentation, and Hubble and SDSS telescope archives that were never Zarr, streamed in place via VirtualiZarr. I’ll be equally clear about the price: a block-local shuffle rather than a global one, and the regimes where a tuned worker pool still wins.

**Agenda** :

- ~15 minutes - Showcase presentation
- 10 - 30 minutes - Discussion
- 15 - 30 minutes - Community check-in

---

<div class="post-metadata">

**Author:** ![emfdavid](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/emfdavid/32/2598_2.png) [@emfdavid](https://discourse.pangeo.io/u/emfdavid)\
**Post date:** [September 11, 2026, 4:37am UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/2 "2026-09-11T04:37:31Z")

</div>

[insitubatch · PyPI](https://pypi.org/project/insitubatch/0.2.0/) version 0.2.0 is live on pypi!  
Architecture and design docs are still being updated before the presentation.

---

<div class="post-metadata">

**Author:** ![emfdavid](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/emfdavid/32/2598_2.png) [@emfdavid](https://discourse.pangeo.io/u/emfdavid)\
**Post date:** [September 16, 2026, 2:56pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/3 "2026-09-16T14:56:10Z")

</div>

The data loader is a logistics problem.  
Notebook forecast demo [gist](https://gist.github.com/emfdavid/66db4785127f84335c4f8f5753cda314)  
[Slides](https://claude.ai/artifact/52mfs9TA7VqzDbh3np2BQv) - updated with one new slide for Tile Shape.

---

<div class="post-metadata">

**Author:** ![martindurant](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/martindurant/32/388_2.png) [@martindurant](https://discourse.pangeo.io/u/martindurant)\
**Post date:** [September 16, 2026, 6:39pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/4 "2026-09-16T18:39:10Z")

</div>

Please update with recording, when it is ready

---

<div class="post-metadata">

**Author:** ![maxrjones](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/maxrjones/32/3190_2.png) [@maxrjones](https://discourse.pangeo.io/u/maxrjones)\
**Post date:** [September 28, 2026, 7:08pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/5 "2026-09-28T19:08:39Z")

</div>

The recording is at the top of the post. thanks again of your fantastic presentation @emfdavid!

---

<div class="post-metadata">

**Author:** ![emfdavid](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/emfdavid/32/2598_2.png) [@emfdavid](https://discourse.pangeo.io/u/emfdavid)\
**Post date:** [September 28, 2026, 8:27pm UTC](https://discourse.pangeo.io/t/pangeo-showcase-insitubatch-streaming-ml-batches-from-cloud-zarr-without-resharding-september-16-12-pm-edt-16-00-utc/5812/6 "2026-09-28T20:27:05Z")

</div>

Thank you Max!

David
