Skip to content

config.toml Reference

The config.toml file contains all information on how the EM27 Retrieval Pipeline should run. The key version specifies the version of the EM27 Retrieval Pipeline that this config file is compatible with. This is an example config. Below it you can find all possible configuration options supported by RETRO’s latest version.

Example:

version = "1.11"
# METADATA
[metadata]
source = "local"
# DATA
[data.atmospheric_profiles]
path = "path-to-atmospheric-profiles"
[data.ground_pressure]
path = "path-to-ground-pressure-data"
file_regex = "^ground-pressure-$(SENSOR_ID)-$(YYYY)-$(MM)-$(DD).csv$"
separator = ","
pressure_column = "pressure"
pressure_column_format = "hPa"
date_column = "UTCdate_____"
date_column_format = "%Y-%m-%d"
time_column = "UTCtime_____"
time_column_format = "%H:%M:%S"
[data.interferograms]
path = "path-to-interferogram-directory"
ifg_file_regex = "^$(SENSOR_ID)$(DATE).*\\.\\d+$"
[data.results]
path = "path-to-results-directory"
# GGG PROFILES DOWNLOADER
[ggg_profiles_downloader.server]
email = "...@..."
max_parallel_requests = 25
[ggg_profiles_downloader.scope]
from_date = "2022-01-01"
to_date = "2022-01-05"
models = [ "GGG2014", "GGG2020" ]
force_download_locations = [ "TUM_I" ]
[[ggg_profiles_downloader.ggg2020_standard_sites]]
identifier = "mu"
lat = 48.151
lon = 11.569
from_date = "2019-01-01"
to_date = "2099-12-31"
# RETRIEVAL
[retrieval.general]
max_process_count = 9
queue_verbosity = "compact"
[retrieval.jobs.0]
retrieval_algorithm = "proffast-1.0"
atmospheric_profile_model = "GGG2014"
sensor_ids = [ "ma", "mb", "mc", "md", "me" ]
from_date = "2019-01-01"
to_date = "2022-12-31"
store_binary_spectra = true
dc_min_threshold = 0.05
dc_var_threshold = 0.1
use_local_pressure_in_pcxs = true
use_ifg_corruption_filter = false
[retrieval.jobs.0.custom_ils.ma]
channel1_me = 0.9892
channel1_pe = -0.001082
channel2_me = 0.9892
channel2_pe = -0.001082
[retrieval.jobs.0.custom_ils.mb]
channel1_me = 0.9893
channel1_pe = -0.001083
channel2_me = 0.9893
channel2_pe = -0.001083
[retrieval.jobs.1]
retrieval_algorithm = "proffast-2.4"
atmospheric_profile_model = "GGG2020"
sensor_ids = [ "ma", "mb", "mc", "md", "me" ]
from_date = "2019-01-01"
to_date = "2099-12-31"
# BUNDLE GENERATOR
[[bundle_exports]]
dst_dir = "directory-to-write-the-bundles-to"
output_formats = [ "csv", "parquet" ]
from_datetime = "2022-01-01T00:00:00Z"
to_datetime = "2022-12-31T23:59:59Z"
retrieval_algorithms = [ "proffast-1.0", "proffast-2.4" ]
atmospheric_profile_models = [ "GGG2014", "GGG2020" ]
sensor_ids = [ "ma", "mb", "mc", "md", "me" ]
parse_dc_timeseries = true
parse_retrieval_diagnostics = true
# GEOMS GENERATOR
[[geoms_exports]]
sensor_ids = [ "ma", "mb", "mc", "md", "me" ]
retrieval_algorithms = [ "proffast-1.0", "proffast-2.4" ]
atmospheric_profile_models = [ "GGG2014", "GGG2020" ]
from_datetime = "2022-01-01T00:00:00Z"
to_datetime = "2022-12-31T23:59:59Z"
parse_dc_timeseries = false
max_sza = 80
min_xair = 0.98
max_xair = 1.02
conflict_mode = "replace"

Definition on which metadata to use, where to find input data and where to store output data.


Contains: How and where to get the metadata from.
source
[string]
* required

Where to source the metadata from. If `local`, it will use `config/em27_metadata.toml`. If `github`, it will download the metadata from the GitHub repository specified in the `github_repository` field.

Allowed values:
- "local"
- "github"
github_repository
[string]

GitHub repository name, e.g. `my-org/my-repo`.

Default: null

Regex Pattern: "^[a-z0-9-_]+/[a-z0-9-_]+$"

github_access_token
[string]

GitHub access token with read access to the repository, only required if the repository is private.

Default: null

Min. Length: 1


Contains: Where to find the atmospheric profile files.
path
[string]
* required

Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.


Contains: Format of the ground pressure files. We support any text file that stores one data point per row and separates the columns with a comma, space, or tab, i.e. CSV, TSV, or space-separated files. Using the `file_regex` field, you specify which files to consider for a given sensor id and date. You have to specify the columns that contain the date and time of the data. There is three options to specify this - the CLI will complain if you configure none or more than one of these options: * One column with a datetime string -> configure `datetime_column` and `datetime_column_format` * Two columns, one with the date and one with the time -> configure `date_column`, `date_column_format`, `time_column`, and `time_column_format` * One column with a unix timestamp -> configure `unix_timestamp_column` and `unix_timestamp_column_format`
path
[string]
* required

Directory path to ground pressure files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.

file_regex
[string]
* required

A regex string to match the ground pressure file names. In this string, you can use the placeholders `$(SENSOR_ID)`, `$(YYYY)`, `$(YY)`, `$(MM)`, and `$(DD)` to make this regex target a certain station and date. The placeholder `$(DATE)` is a shortcut for `$(YYYY)$(MM)$(DD)`.

Examples:
- "^$(DATE).tsv$"
- "^$(SENSOR_ID)_$(DATE).dat$"
- "^ground-pressure-$(SENSOR_ID)-$(YYYY)-$(MM)-$(DD).csv$"

Min. Length: 1

separator
[string]
* required

Separator used in the ground pressure files. Only needed and used if the file format is `text`.

Examples:
- ","
- "\t"
- " "
- ";"

Min. Length: 1

Max. Length: 1

datetime_column
[string | null]

Column name in the ground pressure files that contains the datetime.

Default: null

Examples:
- "datetime"
- "dt"
- "utc-datetime"
datetime_column_format
[string | null]

Format of the datetime column in the ground pressure files.

Default: null

Examples:
- "%Y-%m-%dT%H:%M:%S"
date_column
[string | null]

Column name in the ground pressure files that contains the date.

Default: null

Examples:
- "date"
- "d"
- "utc-date"
date_column_format
[string | null]

Format of the date column in the ground pressure files.

Default: null

Examples:
- "%Y-%m-%d"
- "%Y%m%d"
- "%d.%m.%Y"
time_column
[string | null]

Column name in the ground pressure files that contains the time.

Default: null

Examples:
- "time"
- "t"
- "utc-time"
time_column_format
[string | null]

Format of the time column in the ground pressure files.

Default: null

Examples:
- "%H:%M:%S"
- "%H:%M"
- "%H%M%S"
unix_timestamp_column
[string | null]

Column name in the ground pressure files that contains the unix timestamp.

Default: null

Examples:
- "unix-timestamp"
- "timestamp"
- "ts"
unix_timestamp_column_format
[string]

Format of the unix timestamp column in the ground pressure files. I.e. is the Unix timestamp in seconds, milliseconds, etc.?

Default: null

Allowed values:
- "s"
- "ms"
- "us"
- "ns"
pressure_column
[string]
* required

Column name in the ground pressure files that contains the pressure.

Examples:
- "pressure"
- "p"
- "ground_pressure"
pressure_column_format
[string]
* required

Unit of the pressure column in the ground pressure files.

Allowed values:
- "hPa"
- "Pa"
- "bar"
- "mbar"
- "atm"
- "psi"
- "inHg"
- "mmHg"

Contains: Where to find the interferograms.
path
[string]
* required

Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.

ifg_file_regex
[string]
* required

A regex string to match the ifg file names. In this string, `$(SENSOR_ID)`, `$(YYYY)`, `$(YY)`, `$(MM)`, and `$(DD)` are placeholders to target a certain station and date. The placeholder `$(DATE)` is a shortcut for `$(YYYY)$(MM)$(DD)`. They don't have to be used - you can also run the retrieval on any file it finds in the directory using `.*`

Examples:
- "^*\\.\\d+$^$(SENSOR_ID)$(DATE).*\\.\\d+$"
- "^$(SENSOR_ID)-$(YYYY)-$(MM)-$(DD).*\\.nc$"

Min. Length: 1


Contains: Where to find the results.
path
[string]
* required

Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.


Settings for downloading vertical profiles from the ccycle FTP server.

Contains: Settings for accessing the ccycle ftp server. Besides the `email` field, these can be left as default in most cases.
email
[string]
* required

Email address to use to log in to the ccycle ftp server.

Min. Length: 3

max_parallel_requests
[integer]

Maximum number of requests to put in the queue on the ccycle server at the same time. Only when a request is finished, a new one can enter the queue.

Default: 25

Minimum: 1

Maximum: 200


Contains: From when to when to request the vertical profile data and which models to request. It use the em27 metadata to determine which sensors are located at which locations and request all profiles for these locations in the period specified in this scope.
from_date
[string]

Date in format `YYYY-MM-DD` from which to request vertical profile data.

Default: "1900-01-01"

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"

to_date
[string]

Date in format `YYYY-MM-DD` until which to request vertical profile data.

Default: "2100-01-01"

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"

models
[array]
* required

list of data types to request from the ccycle ftp server.

force_download_locations
[array]

List of locations to force-download data for. These will be downloaded even at times where no instrument in the metadata is located there.

Default: []


ggg_profiles_downloader.ggg2020_standard_sites[i]

Section titled “ggg_profiles_downloader.ggg2020_standard_sites[i]”
Contains: A list item of the `ggg2020_standard_sites` list describing for which standard site to download data for.
identifier
[string]
* required

Identifier of the standard site used on the ginput server.

Min. Length: 1

lat
[number]
* required

Minimum: -90

Maximum: 90

lon
[number]
* required

Minimum: -180

Maximum: 180

from_date
[string]
* required

Date in format `YYYY-MM-DD` from which this standard site is active.

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"

to_date
[string]

Date in format `YYYY-MM-DD` until which this standard site is active. Default is yesterday.

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"


Settings for automated proffast processing. It contains general settings applied to every retrieval job and a list of retrieval jobs to run.

Contains: Settings applied to all retrieval jobs.
max_process_count
[integer]

How many parallel processes to dispatch. There will be one process per sensor-day. With hyper-threaded CPUs, this can be higher than the number of physical cores.

Default: 1

Minimum: 1

Maximum: 1024

queue_verbosity
[string]

How much information the retrieval queue should print out. In `verbose` mode it will print out the full list of sensor-days for each step of the filtering process. This can help when figuring out why a certain sensor-day is not processed.

Default: "compact"

Allowed values:
- "compact"
- "verbose"
container_dir
[string | null]

Directory to store the containers in. If not set, it will use `./data/containers` inside the pipeline directory. If your system has enough memory, you could also use `/dev/shm` which is a memory-based file system where files are stored in memory and never written to disk.

Default: null


Contains: Settings for filtering the storage data. Only used if `config.data_sources.storage` is `true`.
retrieval_algorithm
[string]
* required

Which retrieval algorithms to use. Proffast 2.X uses the Proffast Pylot under the hood to dispatch it. Proffast 1.0 uses a custom implementation by us similar to the Proffast Pylot.

Allowed values:
- "proffast-1.0"
- "proffast-2.2"
- "proffast-2.3"
- "proffast-2.4"
- "proffast-2.4.1"
atmospheric_profile_model
[string]
* required

Which vertical profiles to use for the retrieval.

Allowed values:
- "GGG2014"
- "GGG2020"
sensor_ids
[array]
* required

Sensor ids to consider in the retrieval.

Min. Items: 1

from_date
[string]
* required

Date string in format `YYYY-MM-DD` from which to consider data in the storage directory.

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"

to_date
[string]

Date string in format `YYYY-MM-DD` until which to consider data in the storage directory. Default is yesterday.

Regex Pattern: "^\d{4}-\d{2}-\d{2}$"

store_binary_spectra
[boolean]

Whether to store the binary spectra files. These are the files that are used by the retrieval algorithm. They are not needed for the output files, but can be useful for debugging.

Default: false

dc_min_threshold
[number]

Value used for the `DC_min` threshold in Proffast. If not set, defaults to the Proffast default.

Default: 0.05

Minimum: 0.001

Maximum: 0.999

dc_var_threshold
[number]

Value used for the `DC_var` threshold in Proffast. If not set, defaults to the Proffast default.

Default: 0.1

Minimum: 0.001

Maximum: 0.999

use_local_pressure_in_pcxs
[boolean]

Whether to use the local pressure in the pcxs files. If not used, it will tell PCXS to use the pressure from the atmospheric profiles (set the input value in the `.inp` file to `9999.9`). If used, the pipeline computes the solar noon time using `skyfield` and averages the local pressure over the time period noon-2h to noon+2h.

Default: false

use_ifg_corruption_filter
[boolean]

Whether to use the ifg corruption filter. This filter is a program based on `preprocess4` and is part of the `tum-esm-utils` library: https://tum-esm-utils.netlify.app/api-reference#tum_esm_utilsinterferograms. If activated, we will only pass the interferograms to the retrieval algorithm that pass the filter - i.e. that won't cause it to crash.

Default: true

custom_ils
[object]

Maps sensor IDs to ILS correction values. If not set, the pipeline will use the values published inside the Proffast Pylot codebase (https://github.com/coccon/proffastpylot).

Default: {}

Examples:
- { "ma": { "channel1_me": 0.9892, "channel1_pe": -0.001082, "channel2_me": 0.9892, "channel2_pe": -0.001082 }, "mb": { "channel1_me": 0.9893, "channel1_pe": -0.001083, "channel2_me": 0.9893, "channel2_pe": -0.001083 } }
output_suffix
[string | null]

Suffix to append to the output folders. If not set, the pipeline output folders are named `sensorid/YYYYMMDD/`. If set, the folders are named `sensorid/YYYYMMDD_suffix/`. This is useful when having multiple retrieval jobs processing the same sensor dates with different settings.

Default: null

pressure_calibration_factors
[object]

Maps sensor IDS to pressure calibration factors. If not set, it is set to 1 for each sensor. `corrected_pressure = input_pressure * calibration_factor + calibration_offset`

Default: {}

Examples:
- "{\"ma\": 0.99981}"
- "{\"ma\": 1.00019, \"mb\": 0.99981}"
pressure_calibration_offsets
[object]

Maps sensor IDS to pressure calibration offsets. If not set, it is set to 0 for each sensor. `corrected_pressure = input_pressure * calibration_factor + calibration_offset`

Default: {}

Examples:
- "{\"ma\": -0.00007}"
- "{\"ma\": -0.00007, \"mb\": 0.00019}"

The pipeline can bundle a set of retrieval results into GEOMS files (one file per sensor and day) or into bundle files (one file per sensor). You can define a list of bundle/geoms targets to be generated. It uses the metadata and data paths defined in the input/output section above.


Contains: There will be one file per sensor id and atmospheric profile and retrieval algorithm combination. The final name looks like `em27-retrieval-bundle-$SENSOR_ID-$RETRIEVAL_ALGORITHM-$ATMOSPHERIC_PROFILE-$FROM_DATE-$TO_DATE$BUNDLE_SUFFIX.$OUTPUT_FORMAT`, e.g.`em27-retrieval-bundle-ma-GGG2020-proffast-2.4-20150801-20240523-v2.1.csv`. The bundle suffix is optional and can be used to distinguish between different internal datasets.
dst_dir
[string]
* required

Directory to write the bundeled outputs to. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.

output_formats
[array]
* required

List of output formats to write the merged output files in. Allowed values are `csv` and `parquet`.

Examples:
- [ "csv" ]
- [ "parquet" ]
- [ "csv", "parquet" ]
from_datetime
[string]
* required

UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` from which to bundle data

Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"

to_datetime
[string]
* required

UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` to which to bundle data

Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"

retrieval_algorithms
[array]
* required

The retrieval algorithms for which to bundle the outputs

atmospheric_profile_models
[array]
* required

The atmospheric profile models for which to bundle the outputs

sensor_ids
[array]
* required

The sensor ids for which to bundle the outputs

bundle_suffix
[string]

Suffix to append to the output bundles.

Default: null

Examples:
- "v2.1"
- "v2.2"
- "oco2-gradient-paper-2021"

Min. Length: 1

retrieval_job_output_suffix
[string | null]

When you ran the retrieval with a custom suffix, you can specify it here to only bundle the outputs of this suffix. Use the same value here as in the field `config.retrieval.jobs[i].settings.output_suffix`.

Default: null

parse_dc_timeseries
[boolean]

Whether to parse the DC timeseries from the results directories. This is an output only available in this Pipeline for Proffast2.4. We adapted the preprocessor to output the DC min/mean/max/variation values for each record of data. If you having issues with a low signal intensity on one or both channels, you can run the retrieval with a very low DC_min threshold and filter the data afterwards instead of having to rerun the retrieval.

Default: false

parse_retrieval_diagnostics
[boolean]

Whether to parse the retrieval diagnostics from the results directories - `niter`, `rms`, and `scl` for each retrieval job.

Default: false


Contains: There will be one file per retrieval output directory and the h5 files will be stored in the individual output directories of the results folders. Most code of this exporter originates from the GEOMS export code of the PROFFAST Pylot, but we adapted it to fit the output of this pipeline.
sensor_ids
[array]
* required

The sensor ids for which to generate the GEOMS outputs

retrieval_algorithms
[array]
* required

The retrieval algorithms for which to generate the GEOMS outputs

atmospheric_profile_models
[array]
* required

The atmospheric profile models for which to generate the GEOMS outputs

from_datetime
[string]
* required

UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` from which to generate GEOMS data

Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"

to_datetime
[string]
* required

UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` to which to generate GEOMS data

Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"

parse_dc_timeseries
[boolean]

Whether to parse the DC timeseries from the results directories. This is an output only available in this Pipeline for Proffast2.4. We adapted the preprocessor to output the DC min/mean/max/variation values for each record of data. If you having issues with a low signal intensity on one or both channels, you can run the retrieval with a very low DC_min threshold and filter the data afterwards instead of having to rerun the retrieval.

Default: false

dc_min_xco2
[number]

Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCO2 records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).

Default: 0.05

dc_min_xch4
[number]

Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCH4 records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).

Default: 0.05

dc_min_xh2o
[number]

Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XH2O records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).

Default: 0.05

dc_min_xco
[number]

Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCO records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).

Default: 0.05

max_sza
[number | null]

Maximum solar zenith angle to consider in the GEOMS outputs. If not set, it will consider all solar zenith angles.

Default: null

min_xair
[number | null]

Minimum XAIR required to consider in the GEOMS outputs. If not set, it will consider all XAIR values.

Default: null

max_xair
[number | null]

Maximum XAIR required to consider in the GEOMS outputs. If not set, it will consider all XAIR values.

Default: null

conflict_mode
[string]

What to do if an output file already exist.

Default: "replace"

Allowed values:
- "error"
- "skip"
- "replace"
min_datapoints_per_day
[integer]

Minimum number of data points per day required to generate a GEOMS file for that day. If not enough data points are available, no GEOMS file will be generated for that day.

Default: 11

Minimum: 1