Cloud Storage#
HydroMT can read data directly from cloud object stores — Amazon S3, Google Cloud Storage, and Microsoft Azure Blob Storage / ADLS Gen2 — without downloading files manually. All cloud access is built on fsspec, so any protocol that fsspec supports can be used.
Install the optional io dependencies to enable cloud storage access:
pip install "hydromt[io]"
This installs s3fs (AWS), gcsfs (GCS), adlfs (Azure), and
azure-identity / azure-ai-ml (Azure authentication and AzureML
datastore support).
Quick comparison#
Provider |
fsspec protocol |
Required package |
Example URI |
|---|---|---|---|
Amazon S3 |
|
|
|
Google Cloud Storage |
|
|
|
Azure Blob / ADLS Gen2 |
|
|
|
For private S3 buckets, see Reading from private S3 buckets. For private Azure storage, see Azure Blob Storage.
Simple cloud access (any provider)#
The simplest way to read from any cloud store is to set the filesystem
on the driver — exactly as you would for a local file, but with a cloud
protocol. This works identically for S3, GCS, and Azure and uses the
default ConventionResolver.
AWS S3 (anonymous)
esa_worldcover:
data_type: RasterDataset
uri: s3://esa-worldcover/v100/2020/ESA_WorldCover_10m_2020_v100_Map_AWS.vrt
driver:
name: rasterio
filesystem:
protocol: s3
anon: true
Google Cloud Storage
cmip6_historical:
data_type: RasterDataset
uri: gs://cmip6/CMIP6/CMIP/MPI-ESM1-2-HR/historical/r1i1p1f1/day/tas/*/*
driver:
name: raster_xarray
filesystem:
protocol: gcs
Azure Blob Storage (anonymous)
noaa_isd:
data_type: DataFrame
uri: abfs://isdweatherdatacontainer/ISDWeather/year=2020/month=1/*.parquet
driver:
name: pandas
filesystem:
protocol: abfs
account_name: azureopendatastorage
anon: true
In all three cases, HydroMT:
Creates an fsspec filesystem from the
filesystem:block (e.g.adlfs.AzureBlobFileSystem(account_name=..., anon=True)).Passes the URI to
ConventionResolver, which callsfs.glob()to resolve wildcards.Hands the resolved URIs to the driver for reading.
The Convention Resolver is cloud-agnostic — it doesn’t know or care which
provider is behind the filesystem. All provider-specific logic lives in the
fsspec implementation (s3fs, gcsfs, adlfs).
When to use this approach:
Public / anonymous containers
Containers where you manage credentials via environment variables that the fsspec implementation picks up automatically
Private S3 buckets, using credentials from your AWS configuration files (see Reading from private S3 buckets)
Simple
abfs://URIs without SAS tokens, HTTPS blob URLs, or AzureML datastore URIs
Reading from private S3 buckets#
Private S3 buckets (including S3-compatible object stores with a custom
endpoint) are accessed through the same filesystem: block as public
buckets, but with credentials. HydroMT hands the filesystem: options to
s3fs, which reads credentials from your AWS configuration files, so no
secrets need to end up in the data catalog.
Configure credentials#
Create the AWS credentials file. On Windows this is
%USERPROFILE%\.aws\credentials, on Linux and macOS ~/.aws/credentials.
Add one section per profile, for example one per bucket or account:
[<profile-name>]
aws_access_key_id = <your-access-key>
aws_secret_access_key = <your-secret-key>
[default]
aws_access_key_id = <your-access-key>
aws_secret_access_key = <your-secret-key>
Then create the AWS config file (%USERPROFILE%\.aws\config on Windows,
~/.aws/config on Linux and macOS). Note that in this file, named
profiles use the profile prefix, while [default] does not:
[profile <profile-name>]
region = <region>
endpoint_url = <endpoint-url>
[default]
region = <region>
endpoint_url is only needed for S3-compatible object stores that are not
hosted on AWS; omit it for buckets on AWS S3.
Warning
The credentials file contains secrets. Never commit it, or the keys in it, to version control.
Verify the configuration#
Check that AWS picks up your configuration using python.
The io optional dependency group of hydromt is required for this.
import s3fs
fs = s3fs.S3FileSystem(profile="<profile-name>")
print(fs.ls("<bucket-name>")[:5])
To check that the AWS CLI picks up your configuration, you first need to install aws. Which can be done by adding awscli to your python environment or by doing a system installation.
aws configure list
Then verify that you can list the bucket with the matching profile:
aws s3 ls s3://<bucket-name> --profile <profile-name>
Use the profile in a data catalog#
Select the profile with the profile option of the filesystem: block
and set anon: false so that credentials are used:
my_dataset:
data_type: RasterDataset
uri: s3://<bucket-name>/<path-to-data>/<file>.nc
driver:
name: raster_xarray
filesystem:
protocol: s3
anon: false
profile: <profile-name>
Alternatively, leave out profile and select the profile with the
AWS_PROFILE environment variable. If neither is set, the [default]
profile is used.
$env:AWS_PROFILE = "<profile-name>"
Tip
If you get a NoCredentialsError or a 403 error, run
pixi run aws configure list to check that the profile is found, and
check that the bucket name and profile in your catalog match the commands
you used to verify access above.
Azure Blob Storage#
Public Azure containers and containers that authenticate through environment variables can be read with the generic approach above. For everything Azure-specific — SAS tokens, HTTPS blob URLs, AzureML datastore URIs, the Azure credential chain, and signed HTTPS URLs for rasterio / GDAL — HydroMT provides a dedicated resolver. See Azure Blob Storage for details, including Choosing between resolvers and a step-by-step guide to Step-by-step: accessing private Azure Blob Storage with a SAS token.