Skip to content

EM27 Metadata

To use the pipeline, you have to specify the metadata of your EM27/SUN network: “How was each instrument set up over time?”. The EM27 metadata is stored in a config/em27_metadata.toml file that looks like this:

# LOCATIONS
[[locations]]
...
[[locations]]
...
# SENSORS
[[sensors]]
...
[[sensors]]
...
# CAMPAIGNS
[[campaigns]]
...
[[campaigns]]
...
# EVENTS
[[events]]
...
[[events]]
...

The metadata consists of two required sections: locations and sensors and two optional sections campaigns and events:

  • Locations: The locations where any of your instruments have been deployed. Each location has a unique location_id and coordinates so you can reuse it when defining where the sensors have been deployed.
  • Sensors: When has each sensors been deployed where over time. This list uses the locations defined in the locations section and defines a list of “deployments” for each sensor. In each deployment you define where the sensor has been from which time to which time. Each sensor has a unique sensor_id and when retrieving the EM27/SUN data the pipeline will look for pressure data from the same sensor id. Alternatively you can specify to use a different pressure source during that time which is helpful for side-by-side deployments of multiple EM27/SUNs at the same location.
  • Campaigns: Which sensors/times/locations belong together - e.g., “Hamburg 2023 Campaign”. This is only relevant for the exported bundles which contain a column “campaign_ids” so that the bundles can easily be filtered.
  • Events: Events where something special happened - e.g., “instrument was malfunctioning” or “measuring with a special new setup”, etc.. For each event you can mark whether the data from that period should be used or not. Same as for campaigns, the events are only relevant for the exported bundles which contain two columns “event_description” and “event_data_quality_flag”.

The API Reference section contains a full example file for the metadata as well as a complete specification of the schema.

To configure the pipeline with the metadata, you have two options: Save them locally or store them in a GitHub repository. The latter option is a bit more work to set up, but then the metadata can be edited anywhere and is version-controlled.

Section titled “Option 1: Local Files (Recommended for New Users)”

For this option, you can save the three files em27_metadata.toml file to the config/ folder of the pipeline directory.

1. Create a repository

Use the repository tum-esm/em27-metadata-storage-template as a template to create your own metadata-storage repository. On the top right is a button “Use this template”.

The repository has already been configured with a GitHub Actions workflow to test whether the metadata matches the required schema.

2. Connect the repository to the pipeline

In the pipeline configuration file at config/config.toml, you have to specify the GitHub repository. If you want to keep the repository private, you can use an access token with read access to this repository.

3. Test the connection

You can test the connection to the repository and the integrity of the data in it by running the integration tests with pytest -m metadata.