# Tables, (x)arrays, and rasters¶

**URL:** <https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945>\
**Category:** Uncategorized\
**Created:** [November 22, 2021, 2:37pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945 "2021-11-22T14:37:49Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [November 22, 2021, 2:37pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/1 "2021-11-22T14:37:49Z")

</div>

Hi all,

I’ve been playing around with some ideas for working with geospatial raster data. I’d be curious for any feedback you have.

The core question: _what’s the best data model for raster data in Python?_ Unsurprisingly, I think the answer is “it depends”. Let’s use work through a concrete task and evaluate the various options. Suppose we wanted to compute NDVI for all the scenes captured by Landsat 8 over a couple of hours. (Full notebook at [Jupyter Notebook Viewer](https://nbviewer.org/github/TomAugspurger/rasterpandas/blob/main/example.ipynb))

We’ll use the Planetary Computer’s STAC API to find the scenes, and geopandas to plot the bounding boxes of each scene on a map.

```python
catalog = pystac_client.Client.open(
    "https://planetarycomputer.microsoft.com/api/stac/v1"
)

items = catalog.search(
    collections=["landsat-8-c2-l2"],
    datetime="2021-07-01T08:00:00Z/2021-07-01T10:00:00Z"
).get_all_items()

items = [planetary_computer.sign(item) for item in items]
items = pystac.ItemCollection(items, clone_items=False)
df = geopandas.GeoDataFrame.from_features(items.to_dict(), crs="epsg:4326")

# https://github.com/geopandas/geopandas/issues/1208
df["id"] = [x.id for x in items]
m = df[["geometry", "id", "datetime"]].explore()
m

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/c/cd90dc3a05e5d1e7035f59f3e7bdc9ee24bcce20.jpeg)

This type of data _can_ be represented as an xarray DataArray. But it’s not the most efficient way to store the data:

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/1/13cef84a05792cb27063e7fe565c0dde4469123d.png)

To build this `(time, band, y, x)` DataArray, we end up with many missing values. If you think about the data_cube_ literally, with some “volume” of observed pixels, we have a lot of empty space. In this case, the DataArray takes 426 TiB to store.

Even if we collapse the time dimension, which probably makes sense for this dataset, we still have empty space in the “corners”

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/9/937ad82d914b16a697534d9d64aeb26cd1831b0c.png)

This helps a lot, getting us down to 3.6 TiB (the curse of dimensionality works in reverse too!) But it’s still not as efficient as possible because of that empty space in the corners for this dataset. To actually load all these rasters into, say, a list would take much less memory: 218.57 GiB.

So _for this dataset_ (I cannot emphasize that enough; this example was deliberatly designed to look bad for a data cube) it doesn’t make sense to model the data as a DataArray.

| data model | memory (TiB) |
| --- | --- |
| xarray `(time, band, y, x)` | 426 |
| xarray `(band, y, x)` | 3.6 |
| list | 0.2 |

In the Python data science space, we’re fortunate to have both xarray and pandas (and geopandas and dask.dataframe). So we have choices! pandas provides an [extension array interface](https://pandas.pydata.org/docs/development/extending.html#extension-types) to store non-NumPy arrays inside a pandas DataFrame. What would it look like to store STAC items (and more interestingly, rasters stored as DataArrays) inside a pandas DataFrame? Here’s a prototype:

Let’s load those STAC items into a an “ItemArray”, which can be put in a `pandas.Series`

```python
>>> import rasterpandas

>>> sa = rasterpandas.ItemArray(items)
>>> series = pd.Series(sa, name="stac_items")
>>> series
0 <Item id=LC08_L2SP_191047_20210701_02_T1>
1 <Item id=LC08_L2SP_191046_20210701_02_T1>
2 <Item id=LC08_L2SP_191045_20210701_02_T1>
3 <Item id=LC08_L2SP_191044_20210701_02_T1>
4 <Item id=LC08_L2SP_191043_20210701_02_T1>
                         ...                    
112 <Item id=LC08_L2SP_175010_20210701_02_T1>
113 <Item id=LC08_L2SP_175006_20210701_02_T1>
114 <Item id=LC08_L2SP_175005_20210701_02_T2>
115 <Item id=LC08_L2SP_175001_20210701_02_T1>
116 <Item id=LC08_L2SR_159248_20210701_02_T2>
Name: stac_items, Length: 117, dtype: stac

```

Notice the `stac` dtype. Pandas lets you register accessors. For example, we could have a `stac` accessor that knows how to do stuff with STAC metadata, for example adding a column for each asset in the collection.

```python
rdf = series[:10].stac.with_rasters(assets=["SR_B2", "SR_B3", "SR_B4", "SR_B5"])
rdf

```

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/a/a9fada7fa14e2b93362a91dc749df15ed9053f37.png)

Now things are getting more interesting! The repr is a bit messy, but this new DataFrame has a column for each of the blue, green, red, and nir bands. Each of those is a column of rasters. And each raster is just an xarray.DataArray!

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/8/89e28123618ec846acce82595ba4f1f42fd216c9.png)

```python
>>> ndvi = rdf.raster.ndvi("SR_B4", "SR_B5")
>>> type(ndvi)
pandas.core.series.Series

```

Each element of the output is again a raster (stored in a DataArray).

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/9/9e3024b322c438e5ec32b37fc552ed0e8197c5b5.png)

So that’s the prototype (source code is at [GitHub - TomAugspurger/rasterpandas](https://github.com/TomAugspurger/rasterpandas)). It’s a fun demonstration of pandas’ extension arrays. Is it useful? Maybe. The Spark / Scala world have found that model useful, as implemented by [rasterframes](https://rasterframes.io/). We have xarray, which lowers the _need_ for something like this. We could also use lists of DataArrays, but the DataFrame concept is pretty handy, so I think this might still be useful.

Anyway, if you all have thoughts on whether something like this seems useful, or if you have workflows for similar kinds of data then I’d love to hear your feedback.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [November 23, 2021, 1:58am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/2 "2021-11-23T01:58:27Z")

</div>

Tom just a quick note to say that I really think you’re onto something here. You have articulated a central problem which is hard to tackle with our current stack. (Also closely related to the L2 satellite data processing challenge identified by @cgentemann in [Get Involved in Pangeo: Entry Points for New Contributors - #7 by cgentemann](https://discourse.pangeo.io/t/get-involved-in-pangeo-entry-points-for-new-contributors/643/7)). Your solution elegantly leverages the existing tools (Pandas, Xarray, and STAC) in a really neat way. I would definitely keep pursuing this.

I think the next step would be to define a workflow–say to create a cloud-free mosaic of a region or something like that–and see how this approach maps. I would talk to Chelle, @RichardScottOZ, and others working with L2 data to identify some candidate workflows.

---

<div class="post-metadata">

**Author:** ![RichardScottOZ](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/richardscottoz/32/752_2.png) [@RichardScottOZ](https://discourse.pangeo.io/u/RichardScottOZ)\
**Post date:** [November 23, 2021, 2:20am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/3 "2021-11-23T02:20:07Z")

</div>

Yes, this is really interesting Tom.

[and I have some Databricks compute to burn, so might be interesting to run a rasterframes version too]

Something I have been considering recently is seismic data stored around the country - and 2D lines are basically what is of mining interest - what is the best structure/format for analysis there, similarly.

---

<div class="post-metadata">

**Author:** ![benbovy](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/benbovy/32/592_2.png) [@benbovy](https://discourse.pangeo.io/u/benbovy)\
**Post date:** [November 24, 2021, 4:23pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/4 "2021-11-24T16:23:30Z")

</div>

This is indeed an interesting approach! It seems quite complementary to @TomNicholas’s [DataTree](https://github.com/TomNicholas/datatree) project, with all the advantages here that a pandas DataFrame may have over a more complex, hierarchical structure.

Regarding your use case (i.e., compute NDVI for all the scenes captured by Landsat 8), an alternative approach would be to directly produce a “dense” DataArray object that doesn’t really represents a DataCube but that would rather stack all raster pixels along one spatial dimension. We would loose all raster implicit topology, but I don’t think it’s a big deal for pixelwise operations, which IMO would be easier to perform with this approach than using two nested data structures (DataFrame + DataArray). Moreover, I think we’ll be able to do powerful things by combining both Xarray accessors and Xarray custom indexes (ready soon, hopefully!), e.g., provide some convenient API for selecting and re-transforming the data back into a more conventional raster format.

Leveraging Xarray with non-DataCube-friendly data is (sort of) what we’ve experimented with [Xoak](https://github.com/xarray-contrib/xoak) using ocean model data on curvilinear grids or unstructured meshes as use cases. (note: Xoak will eventually provide custom, Xarray-compatible indexes instead of implementing everything in an accessor).

---

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [November 24, 2021, 9:29pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/5 "2021-11-24T21:29:58Z")

</div>

Thanks for the link to DataTree. I’ll see if I can make a comparison.

> [@benbovy](#):
>
> an alternative approach would be to directly produce a “dense” DataArray object that doesn’t really represents a DataCube but that would rather stack all raster pixels along one spatial dimension.

Yeah, that’s potentially worth exploring too. In [https://github.com/TomAugspurger/planetary-computer-deep-dives/blob/main/Geospatial%20Machine%20Learning.ipynb](https://github.com/TomAugspurger/planetary-computer-deep-dives/blob/main/Geospatial%20Machine%20Learning.ipynb) we do a bit of that to concatenate many scenes together (that also stacks them, which doesn’t have to happen). I think that works very well when you have identically-sized “items”, since you can keep the `y` / `x` labels around in a non-dimension coordinate. I’m less sure how it would work if you have variable-sized items (somewhat common if you’re just working with “raw” satellite scenes).

---

<div class="post-metadata">

**Author:** ![benbovy](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/benbovy/32/592_2.png) [@benbovy](https://discourse.pangeo.io/u/benbovy)\
**Post date:** [November 25, 2021, 10:17am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/6 "2021-11-25T10:17:28Z")

</div>

> [@TomAugspurger](#):
>
> I’m less sure how it would work if you have variable-sized items (somewhat common if you’re just working with “raw” satellite scenes).

In theory you could have an additional coordinate along the stacked spatial dimension that represents raster (or stac item) id values, but yeah getting back the individual rasters as 2D DataArrays would be at the cost of many `unstack` operations. E.g., with this toy example of two concatenated (and stacked) rasters of variable size:

```auto
stacked_raster_id_coords = [
    [0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1],
    [0., 0., 1., 1., 3., 3., 3., 4., 4., 4., 5., 5., 5.],
    [0., 1., 0., 1., 3., 4., 5., 3., 4., 5., 3., 4., 5.]
]
midx = pd.MultiIndex.from_arrays(
    stacked_raster_id_coords, names=("raster_id", "y", "x")
)
data = xr.DataArray(
    np.random.uniform(size=(4+9, 3)),
    coords={"space": midx, "band": [1, 2, 3]}, dims=("space", "band")
)

print(data)

```

```auto
<xarray.DataArray (space: 13, band: 3)>
array([[0.14189417, 0.58131929, 0.35225833],
       [0.97512305, 0.36157414, 0.64397666],
       [0.32119297, 0.93799738, 0.37879202],
       [0.44954318, 0.69639306, 0.65414954],
       [0.21128336, 0.18527528, 0.478475],
       [0.96264641, 0.89752556, 0.17590214],
       [0.534152 , 0.64762276, 0.91105744],
       [0.28872876, 0.64881346, 0.57804076],
       [0.855744 , 0.83493434, 0.01243857],
       [0.99203143, 0.38273195, 0.72474495],
       [0.78296827, 0.30311979, 0.41697918],
       [0.92853232, 0.46370617, 0.6316989],
       [0.86078516, 0.22592683, 0.35725647]])
Coordinates:
  * space (space) MultiIndex
  - raster_id (space) int64 0 0 0 0 1 1 1 1 1 1 1 1 1
  - y (space) float64 0.0 0.0 1.0 1.0 3.0 3.0 ... 4.0 4.0 5.0 5.0 5.0
  - x (space) float64 0.0 1.0 0.0 1.0 3.0 4.0 ... 4.0 5.0 3.0 4.0 5.0
  * band (band) int64 1 2 3

```

Get one raster as a 2D DataArray:

```auto
raster0 = data.sel(raster_id=0).unstack()
print(raster0)

```

```auto
<xarray.DataArray (band: 3, y: 2, x: 2)>
array([[[0.39386449, 0.45238512],
        [0.15979935, 0.40926863]],

       [[0.07529218, 0.96071543],
        [0.15065281, 0.50611628]],

       [[0.11695715, 0.34876804],
        [0.39314546, 0.89090393]]])
Coordinates:
  * band (band) int64 1 2 3
  * y (y) float64 0.0 1.0
  * x (x) float64 0.0 1.0

```

Get all rasters as a list of 2D DataArrays:

```auto
raster_list = [
    da.reset_index("raster_id", drop=True).unstack()
    for _, da in data.groupby("raster_id")
]

raster_list

```

```auto
[<xarray.DataArray (band: 3, y: 2, x: 2)>
 array([[[0.39386449, 0.45238512],
         [0.15979935, 0.40926863]],
 
        [[0.07529218, 0.96071543],
         [0.15065281, 0.50611628]],
 
        [[0.11695715, 0.34876804],
         [0.39314546, 0.89090393]]])
 Coordinates:
   * band (band) int64 1 2 3
   * y (y) float64 0.0 1.0
   * x (x) float64 0.0 1.0,
 <xarray.DataArray (band: 3, y: 3, x: 3)>
 array([[[0.42742665, 0.30037112, 0.41910118],
         [0.27372398, 0.21368011, 0.32303203],
         [0.61795123, 0.15125551, 0.14177116]],
 
        [[0.58677836, 0.17936755, 0.78520273],
         [0.20886481, 0.81725117, 0.5277651],
         [0.53852069, 0.56333296, 0.30122524]],
 
        [[0.03892721, 0.16013487, 0.75304758],
         [0.52124076, 0.14832651, 0.95243819],
         [0.53577372, 0.35633979, 0.59934574]]])
 Coordinates:
   * band (band) int64 1 2 3
   * y (y) float64 3.0 4.0 5.0
   * x (x) float64 3.0 4.0 5.0]

```

Unstack may be expensive and not really necessary, but still required if we want to reuse existing fonctions like `xrspatial.multispectral.ndvi` that only accepts 2D DataArrays. That said, with stacked rasters we can compute the same NDVI for the whole collection of rasters by just writing

```auto
red = data.sel(band=1)
nir = data.sel(band=2)
ndvi = (nir - red) / (nir + red)

```

which is much more expressive (and perhaps more efficient too, depending on how data chunks are affected by stacking the stac items?).

So in summary, to answer your core question in your top comment, I think that both data models (one DataArray with the stacked space dimension vs. one DataFrame of DataArrays) have their pros and cons. The question is whether we really care about the individual 2D rasters or we’re only interested in the individual samples / pixels (i.e., the collection of rasters is only viewed as a data acquisition detail). Sometimes both matter and/or we’d need to switch between one or the other model if we want to reuse some existing libraries (e.g., `xrspatial` vs. `scikit-learn`). So it would be great if we could have some convenient API for that! I guess the conversion would be also pretty cheap as it only requires metadata.

Note that using a `pandas.MultiIndex` for the concat/stacked rasters may not be optimal here, but with the Xarray indexes refactor I guess that it will be possible to create a custom Xarray index tailored to this case (which could also potentially hold other metadata for each individual raster that would be cumbersome to concatenate along extra DataArray coordinates). Or maybe it will be enough to just set multiple `pd.Index` instances for each coordinate along the `space` dimension (not possible today but it will be allowed after the indexes refactor)…

---

<div class="post-metadata">

**Author:** ![ghiggi](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ghiggi/32/1288_2.png) [@ghiggi](https://discourse.pangeo.io/u/ghiggi)\
**Post date:** [November 26, 2021, 1:00pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/7 "2021-11-26T13:00:39Z")

</div>

I think another sensible problem (which partially intersects with the above one) is the lack of a standard workflow for the processing of geolocated irregular time series (i.e. weather sensors, irregular satellite overpass).  
A naive approach would be to loop over each single time series, read it into a pd.Series or xr.DataArray, place it into a list/dict and apply function by looping over it.  
We could instead develop an approach similar to rasterpandas (i.e. tspandas) enabling pandas/geopandas Dataframe to have column storing time series (as pd.Series/DaFrame / xr.DataArray/Dataset) and methods to apply distributed custom computations to all timeseries (i.e. homogenization, statistical analysis, regularization / termporal interpolation, )  
Each tspandas row would correspond to a geolocated observation, the standard dataframe columns would encode “static” sensor/timeseries attributes/features, while the multiple time series columns would represent multiple measured variables, which could eventually be condensed into a single time series column if they share same timesteps.

---

<div class="post-metadata">

**Author:** ![martindurant](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/martindurant/32/388_2.png) [@martindurant](https://discourse.pangeo.io/u/martindurant)\
**Post date:** [November 26, 2021, 6:27pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/8 "2021-11-26T18:27:46Z")

</div>

Quick note: almost all astronomy image processing looks like this, involving jitter (small offsets of a few/non-integer pixels) and mosaicing (large offsets or about the image size) and often multiple detectors in an image and dead-space. Often over time and multiple passbands.

---

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [November 27, 2021, 5:36pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/9 "2021-11-27T17:36:28Z")

</div>

> [@ghiggi](#):
>
> I think another sensible problem (which partially intersects with the above one) is the lack of a standard workflow for the processing of geolocated irregular time series (i.e. weather sensors, irregular satellite overpass).

Just to clarify: would this usecase be satisfied by a long-form table with columns like this?

| timestamp | geometry | value |
| --- | --- | --- |

(or `value` might be multiple columns like `precipitation`, `temperature`, etc.)

Or the nested / grouped nature of each timeseries important?

This proposed workflow sounds a bit like the R `sits` library’s data model: [Setup | sits: Satellite Image Time Series Analysis on Earth Observation Data Cubes](https://e-sensing.github.io/sitsbook/setup.html#the-time-series-table).

---

<div class="post-metadata">

**Author:** ![ghiggi](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ghiggi/32/1288_2.png) [@ghiggi](https://discourse.pangeo.io/u/ghiggi)\
**Post date:** [November 27, 2021, 10:58pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/14 "2021-11-27T22:58:27Z")

</div>

Hi Tom!

Thanks for pointing me to the R package `sits` … I was not aware of it. It is exactly the idea that I was referring to. I actually also already implemented in R a personal library (not packaged yet) to perform time series processing in such a way, with the difference that the list columns contain `zoo` or `xts` time series objects, and it wraps not only `data.frame` but also `sf` spatial objects (the python equivalent of `geopandas`).

Based on my experience with that, when dealing with hundred of thousands of timeseries with high-temporal resolution, the long table format

|ID| \<ID\_attributes\> | geometry | timestep | \<ts\_values\>|

is not efficient for multiple reasons:

- The table starts to have millions/billions of rows

- The need to always group over ID before applying computation is highly inefficient. One could implement some smart caching/indexing but then would need to update it during filtering/subset calls.

- Each time series is often accompanied by static attributes (# timestep, start/end time, flags, min/mean/max/std values, …) that do not change over time. The long table format is highly memory inefficient in the first place because requires duplicating such information over and over. Having a small dataset of 1000 10-years long time series, with an hourly resolution, with just 4 static attributes per timeseries would result in repeating/wasting 1 x 24 x 365 x 10 x 4 x 1000 = 350’000’000 cell values.

- In a second place, if the user desire to select/filter the table based on some of the static attributes, it should go over millions of rows, which is clearly also highly inefficient. Obviously, an alternative would be to store the static attributes separately, but the nested structure directly solves the problem.

- If the timeseries of a given ID do not share the same timesteps, the table starts to be filled by NaNs, which is also memory inefficient.

- Also, if the user desires to have columns with timeseries with different temporal resolutions (i.e. daily/monthly/annual), then a long-format table would require also some sort of temporal\_resolution\_ID column to avoid messing up when doing temporal groupby operation.

- The use of a long-format table constrains each timeseries to have the same data type (a single data type per column), while a nested solution would free from such constrain.

- If each timeseries is saved on disk in separate folders/files, the creation of the table involves a lot of concatenation and a lot of checks (common data type, filling NaN when missing column, …)

- Some considerations should also involve the overhead of the split-apply-combine approach when the number of rows increases. R `data.frame/tibble/data.table` have a column-oriented data storage format (similar to `vaex`), while `pandas` is row-based.

An important point that is also important to bring up is that while long-form tables are easy to be saved on disk efficiently (i.e. parquet), the envisioned development of a `tspandas` class with nested objects would require some thoughts regarding the design of the hierarchical disk storage format. An idea would be a nested structure with \<ts\_database\_name\> / / \<Timeseries\_Column\_Names\>.

The | ID | \<ID\_attributes\> | geometry | table could be saved in [Parquet](https://github.com/geopandas/geo-arrow-spec), while each timeseries in `Parquet`/`Zarr` /`NetCDF` depending on the nested object type (`pandas` or `xarray`).

---

<div class="post-metadata">

**Author:** ![geynard](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/geynard/32/2924_2.png) [@geynard](https://discourse.pangeo.io/u/geynard)\
**Post date:** [December 13, 2021, 4:00pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/15 "2021-12-13T16:00:08Z")

</div>

Hi everyone, and thanks @TomAugspurger for this very interesting topic!

Just wanted to say that I would be really interested by a trade-off in this case between the Pandas approach described by Tom, and a Xarray but non (real) Datacube approach mentioned above by @benbovy. What would be the Pros and Cons of each approach?

Just to be sure I understood correctly, advancing with the Pandas approach would mean building accessors, or links between the Pandas API and the DataArray or stac items series?

---

<div class="post-metadata">

**Author:** ![Material-Scientist](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/material-scientist/32/716_2.png) [@Material-Scientist](https://discourse.pangeo.io/u/Material-Scientist)\
**Post date:** [January 16, 2022, 2:15pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/16 "2022-01-16T14:15:25Z")

</div>

I would also prefer to retain the dense representation, but with tricks to keep the data of sparse type in memory.

Look at the following example with pandas multiindex & sparse dtype:

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/5/5b1600ad14db7af0fababe48013b8725e0d39c0e.jpeg)

The dense data uses ~40 MB of memory, while the dense representation with sparse dtypes uses only ~0.5 kB of memory!

And while you can import dataframes with the `sparse=True` keyword, the size seems to be displayed inaccurately (both are the same size?), and we cannot examine the data like we can with pandas multiindex + sparse dtype:

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/b/be826b5f3d81ff8711824610e443a61b1aea689c.png)

Besides, a lot of operations are not available on sparse xarray data variables (i.e. if I wanted to group by price level for ffill & downsampling):

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/8/80a6047c166c79ac4e9ca44a19ee1869aee1d91a.png)

So, it would be nice if xarray adopted pandas’ approach of unstacking sparse data.

In the end, you could extract all the non-NaN values and write them to a sparse storage format, such as TileDB sparse arrays.

---

<div class="post-metadata">

**Author:** ![TomNicholas](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomnicholas/32/3189_2.png) [@TomNicholas](https://discourse.pangeo.io/u/TomNicholas)\
**Post date:** [April 21, 2022, 9:59pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/17 "2022-04-21T21:59:14Z")

</div>

> Thanks for the [link to DataTree](https://github.com/xarray-contrib/datatree). I’ll see if I can make a comparison.

Once I have the `arrays` variable from your notebook @TomAugspurger , this is what it looks like if I load your data into DataTree:

```bash
!pip install git+https://github.com/xarray-contrib/datatree.git

```

```python
dt = DataTree.from_dict(arrays)
dt

```

```auto
DataTree('root', parent=None)
├── DataTree('LC08_L2SP_191047_20210701_02_T1')
│ Dimensions: (time: 1, band: 4, x: 7532,
│ y: 7692)
│ Coordinates: (12/30)
│ * time (time) datetime64[ns] 2021-07...
│ id (time) <U31 'LC08_L2SP_191047...
│ * band (band) <U5 'SR_B2' ... 'SR_B5'
│ * x (x) float64 6.981e+05 ... 9.2...
│ * y (y) float64 2.195e+06 ... 1.9...
│ landsat:scene_id <U21 'LC81910472021182LGN00'
│ ... ...
│ title (band) <U27 'Blue Band (B2)' ...
│ proj:transform object {0.0, -30.0, 698085.0,...
│ common_name (band) <U5 'blue' ... 'nir08'
│ center_wavelength (band) float64 0.48 ... 0.86
│ full_width_half_max (band) float64 0.06 ... 0.03
│ epsg int64 32631
│ Data variables:
│ stackstac-00624d8333968b598c716b2f5bafc447 (time, band, y, x) float64 dask.array<chunksize=(1, 1, 7692, 7532), meta=np.ndarray>
├── DataTree('LC08_L2SP_191046_20210701_02_T1')
│ Dimensions: (time: 1, band: 4, x: 7702,
│ y: 7852)
│ Coordinates: (12/30)
│ * time (time) datetime64[ns] 2021-07...
│ id (time) <U31 'LC08_L2SP_191046...
│ * band (band) <U5 'SR_B2' ... 'SR_B5'
│ * x (x) float64 1.008e+05 ... 3.3...
│ * y (y) float64 2.357e+06 ... 2.1...
│ landsat:scene_id <U21 'LC81910462021182LGN00'
│ ... ...
│ title (band) <U27 'Blue Band (B2)' ...
│ proj:transform object {0.0, -30.0, 2356815.0...
│ common_name (band) <U5 'blue' ... 'nir08'
│ center_wavelength (band) float64 0.48 ... 0.86
│ full_width_half_max (band) float64 0.06 ... 0.03
│ epsg int64 32632
│ Data variables:
│ stackstac-4b8a8ffb8c81eabc1282a4a440cad38d (time, band, y, x) float64 dask.array<chunksize=(1, 1, 7852, 7702), meta=np.ndarray>
│   
AND SO ON....

```

I can then use datatree to map xarray operations over the different nodes:

```python
dt.sel(band='SR_B2')

```

```auto
DataTree('root', parent=None)
├── DataTree('LC08_L2SP_191047_20210701_02_T1')
│ Dimensions: (time: 1, x: 7532, y: 7692)
│ Coordinates: (12/30)
│ * time (time) datetime64[ns] 2021-07...
│ id (time) <U31 'LC08_L2SP_191047...
│ band <U5 'SR_B2'
│ * x (x) float64 6.981e+05 ... 9.2...
│ * y (y) float64 2.195e+06 ... 1.9...
│ landsat:scene_id <U21 'LC81910472021182LGN00'
│ ... ...
│ title <U27 'Blue Band (B2)'
│ proj:transform object {0.0, -30.0, 698085.0,...
│ common_name <U5 'blue'
│ center_wavelength float64 0.48
│ full_width_half_max float64 0.06
│ epsg int64 32631
│ Data variables:
│ stackstac-00624d8333968b598c716b2f5bafc447 (time, y, x) float64 dask.array<chunksize=(1, 7692, 7532), meta=np.ndarray>
├── DataTree('LC08_L2SP_191046_20210701_02_T1')
│ Dimensions: (time: 1, x: 7702, y: 7852)
│ Coordinates: (12/30)
│ * time (time) datetime64[ns] 2021-07...
│ id (time) <U31 'LC08_L2SP_191046...
│ band <U5 'SR_B2'
│ * x (x) float64 1.008e+05 ... 3.3...
│ * y (y) float64 2.357e+06 ... 2.1...
│ landsat:scene_id <U21 'LC81910462021182LGN00'
│ ... ...
│ title <U27 'Blue Band (B2)'
│ proj:transform object {0.0, -30.0, 2356815.0...
│ common_name <U5 'blue'
│ center_wavelength float64 0.48
│ full_width_half_max float64 0.06
│ epsg int64 32632
│ Data variables:
│ stackstac-4b8a8ffb8c81eabc1282a4a440cad38d (time, y, x) float64 dask.array<chunksize=(1, 7852, 7702), meta=np.ndarray>
│
AND SO ON...

```

I should be able to do the same for `.where`, `.mean` etc.

The datatree can be thought of as a dict-like container of datasets, which can be arbitrarily nested and will map standard xarray operations recursively over all nodes. Each node contains the contents of one `xarray.Dataset`, so can hold individual `DataArrays` if you wish, as well as node-specific metadata.

Is this of any use to you?

(@rabernat turns out it was easy to load it in)

---

<div class="post-metadata">

**Author:** ![kirill.kzb](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/kirill.kzb/32/1276_2.png) [@kirill.kzb](https://discourse.pangeo.io/u/kirill.kzb)\
**Post date:** [May 10, 2022, 5:03am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/18 "2022-05-10T05:03:21Z")

</div>

I wonder if Dask `DataArray` can be coerced into supporting “block sparse” operations. Internally it is basically a dictionary of `index tuples -> delayed np.ndarray`. If you know which blocks have no data at graph construction time (what `odc-stac` is doing for example) then you could skip populating those. Problem is that this is not a supported mode of DataArray, it seems to assume that every block has some data as far as I can tell. So you have to populate those slots with a cheaper to compute function that just returns `np.full` of the right shape and value. But that means your dask graph is still huge and you are forced to process those empty slots as you perform further computations.

It would be cool if dask datarray supported block level “nan”, unpopulated blocks would then be skipped when doing map\_blocks and all operations that use that, you should still be able to turn empties into real data with compute. Slicing and rechunking logic might get even more complicated but should still be possible I think. Only extra bit of information needed from the user is “nan” value to use when pixel data types is not float.

Anyone knows what zarr or tiledb do, do they support missing blocks?

---

<div class="post-metadata">

**Author:** ![TomAugspurger](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/tomaugspurger/32/21_2.png) [@TomAugspurger](https://discourse.pangeo.io/u/TomAugspurger)\
**Post date:** [May 10, 2022, 11:48am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/19 "2022-05-10T11:48:38Z")

</div>

Dask (distributed) does that optimization in a minor way, and it’s used in stackstac at [stackstac/to\_dask.py at 0bc305f4cb64e2b35cb6b997eff9342a5cfa6a1d · gjoseph92/stackstac · GitHub](https://github.com/gjoseph92/stackstac/blob/0bc305f4cb64e2b35cb6b997eff9342a5cfa6a1d/stackstac/to_dask.py#L163-L171). NumPy arrays created using `broadcast_to` will be efficiently serialized and deserialized (I don’t think the same applies to `np.full`).

You do still pay the task overhead, and operations on that tends to fully materialize the array, so something like you describe might still be helpful.

---

<div class="post-metadata">

**Author:** ![kirill.kzb](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/kirill.kzb/32/1276_2.png) [@kirill.kzb](https://discourse.pangeo.io/u/kirill.kzb)\
**Post date:** [May 10, 2022, 8:33pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/20 "2022-05-10T20:33:07Z")

</div>

thanks for that, I’ll switch to np.broadcast, but that’s a relatively minor optimisation in our case as we also re-use the same empty block so it doesn’t get copied a lot anyway or duplicated in RAM.

---

<div class="post-metadata">

**Author:** ![kirill.kzb](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/kirill.kzb/32/1276_2.png) [@kirill.kzb](https://discourse.pangeo.io/u/kirill.kzb)\
**Post date:** [May 26, 2022, 3:29am UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/21 "2022-05-26T03:29:19Z")

</div>

Somewhat related, so I thought I’ll share it here.

By working in a rotated space that matches native orientation of the data we can significantly reduce memory requirements. It is not a generic solution, but it can still be useful in a lot of cases. That’s why next version of `odc-stac==0.3.0rc1` supports arbitrary destination image planes. Produced arrays do not have to be axis aligned to CRS units.

This narrow and tall image displayed below in QGIS was produced on Planetary Computer and is only about 4.7Gb counting all the overviews, 30m pixels in EPSG:3857. Uncompressed size is about 9Gb so it fits into memory of a single instance without any issues leaving enough space to perform COG construction in RAM.

 ![image](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/8/8a2adbaf3fd09af531586bc89ccf91869da15208.jpeg)

To read more see:

> **[Generating Rotated Images to Save Space · opendatacube/odc-stac Wiki](https://github.com/opendatacube/odc-stac/wiki/Generating-Rotated-Images-to-Save-Space)**
>
> Load STAC items into xarray Datasets. Contribute to opendatacube/odc-stac development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![sharkinsspatial](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sharkinsspatial/32/992_2.png) [@sharkinsspatial](https://discourse.pangeo.io/u/sharkinsspatial)\
**Post date:** [October 19, 2022, 5:32pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/22 "2022-10-19T17:32:11Z")

</div>

Some further discussion on the stackstac repository about using [sparse](https://sparse.pydata.org/en/stable/) backing to improve memory efficiency for sparse geospatial data [Consider optional support for sparse array backing. · Discussion #178 · gjoseph92/stackstac · GitHub](https://github.com/gjoseph92/stackstac/discussions/178).

---

<div class="post-metadata">

**Author:** ![benbovy](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/benbovy/32/592_2.png) [@benbovy](https://discourse.pangeo.io/u/benbovy)\
**Post date:** [November 15, 2022, 12:29pm UTC](https://discourse.pangeo.io/t/tables-x-arrays-and-rasters/1945/23 "2022-11-15T12:29:00Z")

</div>

Linking some interesting discussion on the Arrow side: [GeoArrow + Raster? · Issue #24 · geoarrow/geoarrow · GitHub](https://github.com/geoarrow/geoarrow/issues/24). I haven’t read it in detail yet but it looks like the proposed storage model (i.e., “Array of rasters”) is very close to the “DataFrame of DataArrays” concept you have been experimenting with, @TomAugspurger.
