# Decode GeoTIFF to GPU memory

**URL:** <https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214>\
**Category:** Data\
**Tags:** machine-learning\
**Created:** [June 24, 2025, 1:05pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214 "2025-06-24T13:05:21Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [June 24, 2025, 1:05pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/1 "2025-06-24T13:05:21Z")

</div>

Sharing this [blog post](https://weiji14.xyz/blog/geotiffs-to-gpus-part-1:-an-unexpected-conflict-multi-threaded-reads-too-many-jobs) on speeding up Cloud-optimized GeoTIFF (COG) reads using a CUDA GPU library called [nvTIFF](https://docs.nvidia.com/cuda/nvtiff) that I’ve been playing around with for the past month. Hopefully it’ll be useful for folks struggling to use GDAL effectively, or those who don’t want to convert from COG → Format X because … data duplication.

Excerpt from the blog post:

> Preliminary benchmark results from reading a 318MB Sentinel-2 True-Colour Image (TCI) Cloud-optimized GeoTIFF file ([S2A\_37MBV\_20241029\_0\_L2A](https://sentinel-cogs.s3.us-west-2.amazonaws.com/sentinel-s2-l2a-cogs/37/M/BV/2024/10/S2A_37MBV_20241029_0_L2A/TCI.tif)) with DEFLATE compression:
> 
> ![Violin plot of decoding speed benchmarks, GPU vs CPU](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/9/9ef298d99c30ac600c5f3ca63b47d5791df717e5.png)
> 
> Top row shows the GPU-based `nvTIFF`+`nvCOMP` taking about 0.35 seconds (~900MB/s throughput), compared to the CPU-based `GDAL` taking 1.05 seconds (~300MB/s throughput), and `image-tiff` taking 1.75 seconds (~180MB/s throughput). These Cloud-optimized GeoTIFF reads were done reading from a local disk rather than the network.
> 
> Do take this 3x speed up of nvTIFF over GDAL with a grain of salt. On one hand, I probably haven’t optimized the GDAL or nvTIFF code that much yet. On the other hand, I haven’t even included the CPU → GPU transfer cost if I were using this for a machine learning workflow!

For those interested, the code is currently in an experimental PR [here](https://github.com/weiji14/cog3pio/pull/27). Hoping to polish things up in the next few months, but appreciate any feedback!

---

<div class="post-metadata">

**Author:** ![Michael\_Sumner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/michael_sumner/32/1608_2.png) [@Michael\_Sumner](https://discourse.pangeo.io/u/Michael_Sumner)\
**Post date:** [June 24, 2025, 4:15pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/2 "2025-06-24T16:15:07Z")

</div>

awesome really appreciate all this detail! could you please run this in python and report timing so I can gauge what’s comparable on your system?

```python
from osgeo import gdal
gdal.UseExceptions()

import time
t0 = time.time()
ds = gdal.Open("TCI.tif") ## in an empty directory, or set GDAL_DISABLE_READDIR_ON_OPEN=EMPTY_DIR
d = ds.ReadRaster()
t1 = time.time()
print(t1 - t0) ## about 2 seconds for me

type(d)
#<class 'bytearray'>
len(d)
#361681200
ds.RasterXSize * ds.RasterXSize * ds.RasterCount
#361681200

```

I can get the same timing in R by parallelizing read of blocks (16cpus), but on my system 2s seems to be the best I can see, and increasing block size (multiples of native 1024) had no impact. (I have better disk perf on another system but don’t have access to that rn). I’m excited to explore this Rust code.

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [June 24, 2025, 4:53pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/3 "2025-06-24T16:53:36Z")

</div>

Awesome! I’m curious - IIUC correctly we already have zarr-to-GPU, and we have VirtualiZarr, so could we just use the zarr-to-GPU + [VirtualiZarr as a runtime translation layer](https://github.com/zarr-developers/VirtualiZarr/issues/603) to achieve the same thing that this GeoTIFF-to-GPU code does?

---

<div class="post-metadata">

**Author:** ![Michael\_Sumner](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/michael_sumner/32/1608_2.png) [@Michael\_Sumner](https://discourse.pangeo.io/u/Michael_Sumner)\
**Post date:** [June 24, 2025, 6:43pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/4 "2025-06-24T18:43:01Z")

</div>

also fwiw, using LIBERTIFF provides quite a bit more perf:

```python
from osgeo import gdal
gdal.UseExceptions()

ds = gdal.OpenEx("TCI.tif", allowed_drivers = ["LIBERTIFF"], open_options = ["NUM_THREADS=16"])
d = ds.ReadRaster()

```

> **[LIBERTIFF -- GeoTIFF File Format — GDAL documentation](https://gdal.org/en/stable/drivers/raster/libertiff.html)**

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [June 24, 2025, 9:20pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/5 "2025-06-24T21:20:38Z")

</div>

Here’s the timings @Michael_Sumner 😆

## Standard GTiff driver (GDAL 3.10.3)

```python
from osgeo import gdal
import time
import os

gdal.UseExceptions()
os.environ["GDAL_DISABLE_READDIR_ON_OPEN"] = "EMPTY_DIR"

# %%
%%timeit
t0 = time.perf_counter()
ds = gdal.Open("benches/TCI.tif")
d = ds.ReadRaster()
t1 = time.perf_counter()
print(t1 - t0)
# 1.29 s ± 37.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

```

1.29s (Python) - 1.05s (Rust) means about 0.25s of extra overhead from Python.

## LiberTIFF driver (GDAL 3.11.0)

```python
# %%
%%timeit
t0 = time.perf_counter()
ds = gdal.OpenEx("benches/TCI.tif", allowed_drivers = ["LIBERTIFF"], open_options = ["NUM_THREADS=16"])
d = ds.ReadRaster()
t1 = time.perf_counter()
print(t1 - t0)
# 192 ms ± 2.68 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

```

0.2s (LiberTIFF) - 0.35s (nvTIFF) is about 0.15s faster! Guess I’ve got some work to do (still need to benchmark true CPU → GPU timings). My guess is that for small COGs, GDAL+GTiff/LiberTIFF might be performant enough, but larger COGs could benefit from nvTIFF’s GPU-based decoding. But I’ll run the numbers to verify.

Edit: I will note though, as mentioned in the blog post, that multi-threaded GDAL+LiberTIFF will clash with Pytorch multiprocessing, so there’s still value in off-loading decoding to GPU instead of staying on the CPU. Single-threaded LiberTIFF takes `879 ms ± 3.18 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)` on my laptop, 0.88s (LiberTIFF 1 thread) - 0.35s (nvTIFF) = 0.53s gap.

> [@TomNicholas](#):
>
> IIUC correctly we already have zarr-to-GPU, and we have VirtualiZarr, so could we just use the zarr-to-GPU + [VirtualiZarr as a runtime translation layer](https://github.com/zarr-developers/VirtualiZarr/issues/603) to achieve the same thing that this GeoTIFF-to-GPU code does?

That’s what I’ve been wondering for years since this [post](https://discourse.pangeo.io/t/favorite-way-to-go-from-netcdf-xarray-to-torch-tf-jax-et-al/2663/6) (whether we can use kerchunk-at-that-time, VirtualiZarr now, to do direct-to-GPU reads).

My understanding is that we would need `zarr-python`/`VirtualiZarr` to support these GPU-native libs:

| | CPU | GPU |
| --- | --- | --- |
| TIFF metadata/IFD decoding | [`async-tiff`](https://github.com/developmentseed/async-tiff) (Rust) + [`virtual-tiff`](https://github.com/virtual-zarr/virtual-tiff) (Python) | ? |
| Decompression | [`numcodecs`](https://github.com/zarr-developers/numcodecs) (Python/Cython) | [`nvCOMP`](https://developer.nvidia.com/nvcomp) (C++) |

The ? is the key part. I’m proposing that [`nvTIFF`](https://docs.nvidia.com/cuda/nvtiff/index.html) is the more direct way of read COGs to the GPU. The Virtualizarr-way would go through [`kvikio.zarr.GDSStore`](https://docs.rapids.ai/api/kvikio/stable/api/#kvikio.zarr.GDSStore) and if it works, that could be faster in theory since it’s using [`cuFile`](https://docs.nvidia.com/gpudirect-storage/api-reference-guide/index.html). Sadly, nvTIFF doesn’t actually use cuFile yet, but I think it’s only a matter of time.

My hot take is that reading L2 GeoTIFF data to the GPU shouldn’t need to rely on Zarr or wait for the GeoZarr spec. Also, virtualizarr is Python-only for now, and I do think we should be building something that is cross-language compatible, which GDAL+LiberTIFF is doing for CPU workflows, and I’m hoping that Rust-bindings to `nvTIFF` will play that role for GPU workflows.

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [June 24, 2025, 9:49pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/6 "2025-06-24T21:49:54Z")

</div>

> [@weiji14](#):
>
> My understanding is that we would need `zarr-python`/`VirtualiZarr` to support these GPU-native libs

That seems right.

> [@weiji14](#):
>
> shouldn’t need to rely on Zarr

reasonable.

> [@weiji14](#):
>
> wait for the GeoZarr spec

I think this is totally orthogonal.

> [@weiji14](#):
>
> cross-language compatible, which GDAL+LiberTIFF is doing for CPU workflows

In what sense is that cross-language compatible? That you can bind to it from other low-level languages?

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [June 24, 2025, 10:59pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/7 "2025-06-24T22:59:32Z")

</div>

> [@TomNicholas](#):
>
> > [@weiji14](#):
> >
> > cross-language compatible, which GDAL+LiberTIFF is doing for CPU workflows
> 
> In what sense is that cross-language compatible? That you can bind to it from other low-level languages?

Yes, writing these I/O libraries in C/Rust allows us to create bindings in Python/R/Javascript(WebAssembly)/etc. See e.g. what Arrow/[GeoArrow](https://github.com/geoarrow/geoarrow-rs) has done for tabular data.

Besides cross-language, I’m also keen on getting **cross-device** compatibility working, and as mentioned [here](https://github.com/zarr-developers/zarr-python/issues/2658), I’m pushing on [DLPack](https://github.com/dmlc/dlpack) to be the standard in-memory tensor format that will allow for data exchange between Intel CPUs/CUDA GPUs/AMD ROCm/Apple Sillicon/etc. This will enable better ‘separation of storage and compute’ (Zarr is almost exclusively tied to zarr-python/xarray; same with GeoTIFF and GDAL), because then you can store data in any format that can go into DLPack, and then use whatever compute engine that reads from DLPack (Torch/JaX/MLX/etc) to run your algorithms.

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [June 25, 2025, 3:24am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/8 "2025-06-25T03:24:57Z")

</div>

> virtualizarr is Python-only for now, and I do think we should be building something that is cross-language compatible

I mean maybe we should have rust-powered virtualizarr parsers… 😁

From your blogpost:

> I’m proposing we build composable pieces to handle every layer of decoding a Cloud-optimized GeoTIFF:
> 
> - Network/Disk transfer - via [`object_store`](https://crates.io/crates/object_store) or [`kvikio Remote IO`](https://developer.nvidia.com/blog/high-performance-remote-io-with-nvidia-kvikio/)
> - Decompression of raw bytes - through [`numcodecs`](https://numcodecs.readthedocs.io/) or [`nvcomp`](https://github.com/NVIDIA/nvcomp)
> - Parsing of TIFF tag metadata - handled by [`geotiff`](https://crates.io/crates/geotiff) or [`nvtiff-sys`](https://crates.io/crates/nvtiff-sys)
> 
> The first two steps are general enough to be used by other data formats, Zarr, HDF5, etc. It is only the third step - TIFF tag metadata parsing, which requires custom logic.

But the last step is exactly what a VirtualiZarr [`Parser`](https://virtualizarr--619.org.readthedocs.build/en/619/custom_parsers/) for TIFF is meant to do! And that approach is not restricted to COGs (which zero people outside of the geospatial community use or ever will use). It effectively isolates the absolute minimum amount of code that needs to be format-specific (the parser). That’s what I imagine full composability would look like.

I’m probably missing something but couldn’t you do something like:

- Parse the COG’s TIFF metadata in python using virtualizarr
- Now use zarr-python / zarrs + numcodecs / nvCOMP to read and decompress actual bytes, either on CPU or GPU (using dlpack)

And if you wanted all the logic to be cross-language compatible the only part left to do is port the virtualizarr parser to be in rust/C, which would presumably be fairly straightforward if you had started by wrapping a rust-powered tiff parser like `async-tiff`.

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [June 25, 2025, 10:31pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/9 "2025-06-25T22:31:55Z")

</div>

> [@TomNicholas](#):
>
> But the last step is exactly what a VirtualiZarr [`Parser`](https://virtualizarr--619.org.readthedocs.build/en/619/custom_parsers/) for TIFF is meant to do! And that approach is not restricted to COGs (which zero people outside of the geospatial community use or ever will use). It effectively isolates the absolute minimum amount of code that needs to be format-specific (the parser). That’s what I imagine full composability would look like.
> 
> I’m probably missing something but couldn’t you do something like:
> 
> - Parse the COG’s TIFF metadata in python using virtualizarr
> - Now use zarr-python / zarrs + numcodecs / nvCOMP to read and decompress actual bytes, either on CPU or GPU (using dlpack)
> 
> And if you wanted all the logic to be cross-language compatible the only part left to do is port the virtualizarr parser to be in rust/C, which would presumably be fairly straightforward if you had started by wrapping a rust-powered tiff parser like `async-tiff`.

I think we’re both tackling the problem (GeoTIFF to GPU parsing) from two ends, and eventually things will converge 🔀. Virtualizarr is approaching it from a **protocol-based** Python abstraction, what you have in `Parser` is essentially what is called a [Trait](https://doc.rust-lang.org/book/ch10-02-traits.html) in Rust, difference being Python does runtime-checks (to see if you conform to the protocol), whereas Rust does compile-time enforcement. I’m tackling things from a more **low-level implementation** side, which is either `nvTIFF` (CUDA GPU-based) or `async-tiff` (CPU-based, which I’ve also got a foot in) or whatever custom parser logic that still needs to exist to read the GeoTIFF format.

Protocols are tricky to define, and I do have a lot of respect for how things have evolved from kerchunk’s JSON-based format to parquet to Virtualizarr’s Parser protocol! I think we’ve more or less settled on a network protocol (fsspec/object\_store), and buffer/bytes decompression protocol (numcodecs/[compress trait](https://timclicks.dev/article/how-to-use-compression-algorithms-in-rust)), it’s that last mile of parsing custom n-dimensional file formats (HDF5/TIFF/Zarr/etc) that I see Virtualizarr trying to solve, and I think you’re doing a good job for CPU-based parsing in Python at the moment, but we might need more work (in the future) to support cross-device (CUDA/ROCm/Metal) and cross-language (Python/R/Javascript/etc) which is where I’m getting at with [DLPack](https://dmlc.github.io/dlpack/latest/).

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [June 26, 2025, 3:57pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/10 "2025-06-26T15:57:29Z")

</div>

> Virtualizarr is approaching it from a **protocol-based** Python abstraction […] I’m tackling things from a more **low-level implementation** side

I don’t disagree, but I see the difference as more that you’re concentrating on efficiently decompressing and moving array chunk bytes into CPU/GPU memory (which on it’s own is a format-agnostic wish), and I’m focusing on a general framework for parsing the file format metadata needed to know the locations/compression options of those array chunks. I agree they are complementary.

> [@weiji14](#):
>
> I think we’ve more or less settled on a network protocol (fsspec/object\_store), and buffer/bytes decompression protocol (numcodecs/[compress trait](https://timclicks.dev/article/how-to-use-compression-algorithms-in-rust)), it’s that last mile of parsing custom n-dimensional file formats (HDF5/TIFF/Zarr/etc) that I see Virtualizarr trying to solve

I think that’s a good way to frame it.

I guess what I don’t quite understand is why `DLPack` standardization has anything to do with the specific file format at all? IIUC DLPack is a solution for making moving array bytes cross-device and cross-language. That seems totally unrelated to the metadata parsing?

FYI I think there is at least one rust-powered tabular-to-arrow equivalent of what we are trying to build: [GitHub - abdenlab/oxbow: Oxbow makes genomic data ready for high-performance analytics.](https://github.com/abdenlab/oxbow) .

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [June 26, 2025, 4:10pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/11 "2025-06-26T16:10:15Z")

</div>

Clarification that I should have asked earlier: In your blog post above - is it assumed that we are reading data from a local SSD, rather than from cloud object storage? If so then do we have a way to fetch data from object storage direct to GPU using rust?

---

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [June 26, 2025, 7:17pm UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/12 "2025-06-26T19:17:51Z")

</div>

> If so then do we have a way to fetch data from object storage direct to GPU using rust?

This would require technologies like [GPUDirect RDMA](https://docs.nvidia.com/cuda/gpudirect-rdma/index.html), which I think require some specific hardware and drivers that [some cloud VMs apparently support](https://aws.amazon.com/ec2/instance-types/p4/). I’m not 100% sure whether that will help with reads from S3 or not, or whether just RDMA between two EC2 instances is supported.

FWIW, I think our software stack is a ways away from being bound by OS overhead (which RDMA bypasses) and I’m hopeful that the overhead of host to device memory transfers can be masked by using things like pinned memory (nice intro [here](https://docs.pytorch.org/tutorials/intermediate/pinmem_nonblock.html)  
) and overlapping I/O and computation.

---

<div class="post-metadata">

**Author:** ![RichardScottOZ](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/richardscottoz/32/752_2.png) [@RichardScottOZ](https://discourse.pangeo.io/u/RichardScottOZ)\
**Post date:** [July 3, 2025, 5:54am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/13 "2025-07-03T05:54:17Z")

</div>

Saw this recently too @weiji14[GitHub - microsoft/pytorch-cloud-geotiff-optimization: A toolkit for optimizing cloud GeoTIFF streaming in PyTorch. Achieves 20x throughput and 90% GPU utilization through optimized data loading and compression. Paper: https://arxiv.org/pdf/2506.06235](https://github.com/microsoft/pytorch-cloud-geotiff-optimization)

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [August 2, 2025, 3:02am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/14 "2025-08-02T03:02:48Z")

</div>

[Part 2](https://weiji14.xyz/blog/geotiffs-to-gpus-part-2:-barrels-out-of-bytes-streaming-rust-bits-to-the-gpu/) is now out, covering the technical details of how the Rust bindings to `nvTIFF` (C++) works.

> [@TomAugspurger](#):
>
> > If so then do we have a way to fetch data from object storage direct to GPU using rust?
> 
> This would require technologies like [GPUDirect RDMA](https://docs.nvidia.com/cuda/gpudirect-rdma/index.html), which I think require some specific hardware and drivers that [some cloud VMs apparently support](https://aws.amazon.com/ec2/instance-types/p4/). I’m not 100% sure whether that will help with reads from S3 or not, or whether just RDMA between two EC2 instances is supported.
> 
> FWIW, I think our software stack is a ways away from being bound by OS overhead (which RDMA bypasses) and I’m hopeful that the overhead of host to device memory transfers can be masked by using things like pinned memory (nice intro [here](https://docs.pytorch.org/tutorials/intermediate/pinmem_nonblock.html)  
> ) and overlapping I/O and computation.

I think Tom’s covered this well. Will just add that the OpenDAL folks are looking into remote S3 reads too at [new feature: Explore if OpenDAL can support KvikIO (aka Nvidia GPUDirect Storage) · Issue #5090 · apache/opendal · GitHub](https://github.com/apache/opendal/issues/5090#issuecomment-3050998325) . Also, the[rust bindings to cuFile in cudarc](https://github.com/coreylowman/cudarc/issues/403) was just merged yesterday, so maybe just a matter of time!

> [@RichardScottOZ](#):
>
> Saw this recently too @weiji14 [GitHub - microsoft/pytorch-cloud-geotiff-optimization: A toolkit for optimizing cloud GeoTIFF streaming in PyTorch. Achieves 20x throughput and 90% GPU utilization through optimized data loading and compression. Paper: https://arxiv.org/pdf/2506.06235](https://github.com/microsoft/pytorch-cloud-geotiff-optimization)

Yes, that was linked in the first blog post 😃 I actually picked 1 (out of the 10) of the GeoTIFF images they used in their benchmark, specifically [here](https://github.com/microsoft/pytorch-cloud-geotiff-optimization/blob/5fb6d1294163beff822441829dcd63a3791b7808/README.md?plain=1#L32)

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [May 27, 2026, 10:54am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/15 "2026-05-27T10:54:46Z")

</div>

Finally released [cog3pio v0.1.0](https://cog3pio.readthedocs.io/en/v0.1.0/changelog/#010-2026-05-25) this week, and wrote up the third and final blog post in the trilogy 🎉

> **[GeoTIFFs to GPUs part 3: The Last Stage - via DLPack into the world](https://weiji14.xyz/blog/geotiffs-to-gpus-part-3:-the-last-stage-via-dlpack-into-the-world/)**

The xarray integration part is still work in progress at [TIFF to GPU memory via cog3pio backend entrypoint by weiji14 · Pull Request #81 · xarray-contrib/cupy-xarray · GitHub](https://github.com/xarray-contrib/cupy-xarray/pull/81), so open to any feedback before I merge that. Specifically, please test it out if you have 1) access to a CUDA GPU and 2) can run things on linux-x86\_64 or linux-aarch64 (TIFF to GPU decoding not supported on other platforms yet 😛)

---

<div class="post-metadata">

**Author:** ![csaybar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/csaybar/32/3107_2.png) [@csaybar](https://discourse.pangeo.io/u/csaybar)\
**Post date:** [September 21, 2026, 1:10am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/16 "2026-09-21T01:10:46Z")

</div>

Hi @weiji14! I’ve been working on [rumi](https://github.com/asterisk-labs/rumi). It’s a file format that puts GeoZL frames in a GeoTIFF-like container.

I ran it on the same image you benchmarked (S2A\_37MBV\_20241029 TCI). Decompress only, files already in RAM, Apple M5.

| | size | 1 thread | 10 threads |
| --- | --- | --- | --- |
| GDAL GTiff | 237 MB | 2708 ms | 345 ms |
| GDAL LIBERTIFF | 237 MB | 1458 ms | 211 ms |
| rumi, PFOR | 250 MB | 271 ms | 64 ms |
| rumi, MED + entropy | 210 MB | 964 ms | 213 ms |

This is all on CPU. The only thing that really changes is the compression and interleave. GeoTIFF has a fixed list of codecs, rumi doesn’t, every frame carries its own OpenZL graph, so we can keep adding new codecs or combinations!

**PFOR** = plane predictor (W+N−NW), zigzag, then bit-packing.  
**MED** = the JPEG-LS median predictor, zigzag, then an entropy coder.

> **[Google Colab](https://colab.research.google.com/github/asterisk-labs/rumi/blob/main/examples/rumi-vs-geotiff.ipynb)**

---

<div class="post-metadata">

**Author:** ![weiji14](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/weiji14/32/389_2.png) [@weiji14](https://discourse.pangeo.io/u/weiji14)\
**Post date:** [September 25, 2026, 2:38am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/17 "2026-09-25T02:38:40Z")

</div>

Yes, I’ve been following your work on RUMI (which I see as an evolution of TACOTIFF?), and have mentioned it in a few talks already 😄

Have you got any SIMD code in RUMI yet? I see you’re benchmarking on Apple Sillicon, so the GDAL GTiff multi-threaded benchmark might be slower than expected compared to if it was running on an Intel chip (since GDAL has SIMD code for some predictor/compression combinations). That said, the shared/unified memory for Apple Sillicon means you effectively have data available for the Apple Metal GPU already, so you’re also faster in some sense already without needing CPU-GPU transfer. I’m also wondering if there are OpenZL decoders that work on the Apple Sillicon GPU yet, which would be even cooler.

---

<div class="post-metadata">

**Author:** ![csaybar](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/csaybar/32/3107_2.png) [@csaybar](https://discourse.pangeo.io/u/csaybar)\
**Post date:** [September 25, 2026, 4:18am UTC](https://discourse.pangeo.io/t/decode-geotiff-to-gpu-memory/5214/18 "2026-09-25T04:18:07Z")

</div>

Hi @weiji14! Nice to hear rumi came up in your talks, thanks for that 🙇 . Yes it is a evolution of some initial ideas we put in TACOTIFF 🙂

Yes, there is SIMD in there, and most of the speed comes from how the codecs are designed around it. Planar is the easiest one to show, so let me go step by step because it’s a fun trick.

The predictor is `x = W + N - NW`. When you decode, it looks serial, because to get pixel _i_ you need pixel _i-1_ (the W). So at first sight you would decode one pixel at a time.

But if you move W to the other side you get

```
x[i] - x[i-1] = r[i] + N[i] - N[i-1]

```

The right side only uses the residual and the row above, and the row above is already decoded! So you can compute it for the whole row at once. Call it c.

Small example. The row above is `10 12 15 15` and the residuals are `1 0 -1 2`.

Step 1, compute c for every pixel at the same time (at the left edge W and NW are zero)

```
c = 11 2 2 2

```

Step 2, run a prefix sum over c

```
x = 11 13 15 17

```

and that’s the decoded row. You can check any pixel with the normal formula and it matches.

Step 1 is plain vector math, 8 pixels per NEON instruction for uint16. Step 2 is a prefix sum inside the register. You shift the vector by 1 lane and add, then by 2 and add, then by 4 and add, and all 8 lanes now hold the running sum. The last lane gets carried into the next vector. So the only serial part left is one add every 8 pixels instead of one every pixel. There are SSE2 and AVX2 versions too, the kernel is [here](https://github.com/asterisk-labs/geozl/blob/main/core/src/planar/decode_planar_kernel.c) if you want to take a look.

And if you think about it in 2D, the inverse of planar is just a cumsum down the columns and then a cumsum along the rows. That’s two very standard GPU operations, so planar should be easy to move to CUDA.

PFOR is basically bitpacking with exceptions. Each block of 256 values uses one bit width, and the few values that don’t fit keep their high bits in a small side list. The design is heavily inspired by [TurboPFor](https://github.com/powturbo/TurboPFor-Integer-Compression), especially the vertical SIMD layout. The bits are interleaved in 16 byte lanes, so one vector load gives you a slice of every lane, and unpacking is only shift, and, or. No shuffles at all. For 32 bit values the plain C code is enough and the compiler turns it into NEON by itself. Code is [here](https://github.com/asterisk-labs/geozl/tree/main/core/src/pfor).

About OpenZL on the GPU, I’ve been following their changes for a while. It’s CUDA only for now, nothing for Metal, and nothing in a release yet ([v0.2.0](https://github.com/facebook/openzl/releases/tag/v0.2.0) is still the latest). But, on the dev branch there is a `ZL_GPU_decompress()` entry point in [contrib/gpu](https://github.com/facebook/openzl/tree/dev/contrib/gpu). It takes a frame that’s already in device memory and builds a decode plan, but it doesn’t run it yet. Only a few standard codecs are wired in, and the only real kernel so far is bf16 float\_deconstruct.

The part I like most/I’m quite excited is [PivCo-Huffman](https://github.com/facebook/openzl/tree/dev/src/openzl/codecs/pivco_huffman), which they just added ([paper](https://arxiv.org/abs/2606.05765), [original repo](https://github.com/MarcinZukowski/pivco-huffman)). It’s a Huffman variant that decodes a whole block using bitmaps and popcounts instead of one code at a time. They have a [CUDA proof of concept](https://github.com/facebook/openzl/tree/dev/contrib/pivco-huffman/gpu) doing 100 to 190 GiB/s on an A100, and it uses the same bytes as the CPU version. On my M5 with NEON it decoded 2 to 3x faster than the current Huffman in a quick test, same ratio. So one file can be fast on both CPU and GPU. It needs a newer frame format, so we have to wait for the next OpenZL release.

For the GPU side of [GeoZL](https://github.com/asterisk-labs/geozl) I’m thinking of a plain C API that takes device pointers and a CUDA stream, and then going through [DLPack](https://github.com/dmlc/dlpack), not tied to torch. Pretty much the path from your [part 3 post](https://weiji14.xyz/blog/geotiffs-to-gpus-part-3-the-last-stage-via-dlpack-into-the-world/).

Metal with unified memory would be really cool, I agree. I haven’t seen anyone working on it yet though.

Sorry this got so long, but nobody around me wants to talk about prefix sums! 🐛
