# What's the best file format to chose for raster imagery and masks products

**URL:** <https://discourse.pangeo.io/t/whats-the-best-file-format-to-chose-for-raster-imagery-and-masks-products/4555>\
**Category:** Data\
**Created:** [October 1, 2024, 7:00am UTC](https://discourse.pangeo.io/t/whats-the-best-file-format-to-chose-for-raster-imagery-and-masks-products/4555 "2024-10-01T07:00:48Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![Basile\_Goussard](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/basile_goussard/32/2839_2.png) [@Basile\_Goussard](https://discourse.pangeo.io/u/Basile_Goussard)\
**Post date:** [October 2, 2024, 3:13pm UTC](https://discourse.pangeo.io/t/whats-the-best-file-format-to-chose-for-raster-imagery-and-masks-products/4555/5 "2024-10-02T15:13:12Z")

</div>

The easiest solution for us (a startup providing a product based on satellite data) is to rely on STAC + COG. It’s easy to handle and, from my perspective, covers almost all the use cases.  
Retrieving data (by selecting the relevant pixels) can be done using `odc.stac` and `rioxarray`.  
Visualization can be achieved by relying on `titiler`, as well as by directly adding the URL into QGIS (which works great for quick visualization), and large-scale processing can easily be handled using `coiled` (or HPC).

However, we haven’t succeeded in making it work for one use case:  
 → Retrieving the time series of all Sentinel-2 data over more than 60,000 points.  
We used `xvec` (a pretty awesome library), but it was still too slow…

> [@Compute time series for 70,000 locations (Speed up the processing)](https://discourse.pangeo.io/t/compute-time-series-for-70-000-locations-speed-up-the-processing/4436/4):
>
> Hi @Basile_Goussard, I have some additional remarks: Do you consider all the time steps or do you filter out cloudy scenes beforehand using the eo:cloud\_cover property? If not applied already, this could save you some time, reducing the number of Items you need to load. Are you running the computation in a serial manner? If you have more than 1 CPU you could easily parallelize the whole computation using libraries like joblib, multiprocessing or dask You could also think about grouping the geo…

I’m not sure if `zarr` would be a better candidate for this use case.

Regarding ML, I don’t know if batching COG (with `xbatch`) would have the same capabilities as `zarr`. I think not, but I’m not sure how many users will try it.

---

_[View the full topic](https://discourse.pangeo.io/t/whats-the-best-file-format-to-chose-for-raster-imagery-and-masks-products/4555)._
