config.toml Reference
The config.toml file contains all information on how the EM27 Retrieval Pipeline should run. The key version specifies the version of the EM27 Retrieval Pipeline that this config file is compatible with. This is an example config. Below it you can find all possible configuration options supported by RETRO’s latest version.
Example:
version = "1.11"
# METADATA
[metadata]source = "local"
# DATA
[data.atmospheric_profiles]path = "path-to-atmospheric-profiles"
[data.ground_pressure]path = "path-to-ground-pressure-data"file_regex = "^ground-pressure-$(SENSOR_ID)-$(YYYY)-$(MM)-$(DD).csv$"separator = ","pressure_column = "pressure"pressure_column_format = "hPa"date_column = "UTCdate_____"date_column_format = "%Y-%m-%d"time_column = "UTCtime_____"time_column_format = "%H:%M:%S"
[data.interferograms]path = "path-to-interferogram-directory"ifg_file_regex = "^$(SENSOR_ID)$(DATE).*\\.\\d+$"
[data.results]path = "path-to-results-directory"
# GGG PROFILES DOWNLOADER
[ggg_profiles_downloader.server]email = "...@..."max_parallel_requests = 25
[ggg_profiles_downloader.scope]from_date = "2022-01-01"to_date = "2022-01-05"models = [ "GGG2014", "GGG2020" ]force_download_locations = [ "TUM_I" ]
[[ggg_profiles_downloader.ggg2020_standard_sites]]identifier = "mu"lat = 48.151lon = 11.569from_date = "2019-01-01"to_date = "2099-12-31"
# RETRIEVAL
[retrieval.general]max_process_count = 9queue_verbosity = "compact"
[retrieval.jobs.0]retrieval_algorithm = "proffast-1.0"atmospheric_profile_model = "GGG2014"sensor_ids = [ "ma", "mb", "mc", "md", "me" ]from_date = "2019-01-01"to_date = "2022-12-31"store_binary_spectra = truedc_min_threshold = 0.05dc_var_threshold = 0.1use_local_pressure_in_pcxs = trueuse_ifg_corruption_filter = false
[retrieval.jobs.0.custom_ils.ma]channel1_me = 0.9892channel1_pe = -0.001082channel2_me = 0.9892channel2_pe = -0.001082
[retrieval.jobs.0.custom_ils.mb]channel1_me = 0.9893channel1_pe = -0.001083channel2_me = 0.9893channel2_pe = -0.001083
[retrieval.jobs.1]retrieval_algorithm = "proffast-2.4"atmospheric_profile_model = "GGG2020"sensor_ids = [ "ma", "mb", "mc", "md", "me" ]from_date = "2019-01-01"to_date = "2099-12-31"
# BUNDLE GENERATOR
[[bundle_exports]]dst_dir = "directory-to-write-the-bundles-to"output_formats = [ "csv", "parquet" ]from_datetime = "2022-01-01T00:00:00Z"to_datetime = "2022-12-31T23:59:59Z"retrieval_algorithms = [ "proffast-1.0", "proffast-2.4" ]atmospheric_profile_models = [ "GGG2014", "GGG2020" ]sensor_ids = [ "ma", "mb", "mc", "md", "me" ]parse_dc_timeseries = trueparse_retrieval_diagnostics = true
# GEOMS GENERATOR
[[geoms_exports]]sensor_ids = [ "ma", "mb", "mc", "md", "me" ]retrieval_algorithms = [ "proffast-1.0", "proffast-2.4" ]atmospheric_profile_models = [ "GGG2014", "GGG2020" ]from_datetime = "2022-01-01T00:00:00Z"to_datetime = "2022-12-31T23:59:59Z"parse_dc_timeseries = falsemax_sza = 80min_xair = 0.98max_xair = 1.02conflict_mode = "replace"Input / Output
Section titled “Input / Output”Definition on which metadata to use, where to find input data and where to store output data.
metadata
Section titled “metadata”Where to source the metadata from. If `local`, it will use `config/em27_metadata.toml`. If `github`, it will download the metadata from the GitHub repository specified in the `github_repository` field.
GitHub repository name, e.g. `my-org/my-repo`.
Default: null
Regex Pattern: "^[a-z0-9-_]+/[a-z0-9-_]+$"
GitHub access token with read access to the repository, only required if the repository is private.
Default: null
Min. Length: 1
data.atmospheric_profiles
Section titled “data.atmospheric_profiles”Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.
data.ground_pressure
Section titled “data.ground_pressure”Directory path to ground pressure files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.
A regex string to match the ground pressure file names. In this string, you can use the placeholders `$(SENSOR_ID)`, `$(YYYY)`, `$(YY)`, `$(MM)`, and `$(DD)` to make this regex target a certain station and date. The placeholder `$(DATE)` is a shortcut for `$(YYYY)$(MM)$(DD)`.
Min. Length: 1
Separator used in the ground pressure files. Only needed and used if the file format is `text`.
Min. Length: 1
Max. Length: 1
Column name in the ground pressure files that contains the datetime.
Default: null
Format of the datetime column in the ground pressure files.
Default: null
Column name in the ground pressure files that contains the date.
Default: null
Format of the date column in the ground pressure files.
Default: null
Column name in the ground pressure files that contains the time.
Default: null
Format of the time column in the ground pressure files.
Default: null
Column name in the ground pressure files that contains the unix timestamp.
Default: null
Format of the unix timestamp column in the ground pressure files. I.e. is the Unix timestamp in seconds, milliseconds, etc.?
Default: null
Column name in the ground pressure files that contains the pressure.
Unit of the pressure column in the ground pressure files.
data.interferograms
Section titled “data.interferograms”Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.
A regex string to match the ifg file names. In this string, `$(SENSOR_ID)`, `$(YYYY)`, `$(YY)`, `$(MM)`, and `$(DD)` are placeholders to target a certain station and date. The placeholder `$(DATE)` is a shortcut for `$(YYYY)$(MM)$(DD)`. They don't have to be used - you can also run the retrieval on any file it finds in the directory using `.*`
Min. Length: 1
data.results
Section titled “data.results”Directory path to atmospheric profile files. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.
GGG Profiles Downloader
Section titled “GGG Profiles Downloader”ggg_profiles_downloader.server
Section titled “ggg_profiles_downloader.server”Email address to use to log in to the ccycle ftp server.
Min. Length: 3
Maximum number of requests to put in the queue on the ccycle server at the same time. Only when a request is finished, a new one can enter the queue.
Default: 25
Minimum: 1
Maximum: 200
ggg_profiles_downloader.scope
Section titled “ggg_profiles_downloader.scope”Date in format `YYYY-MM-DD` from which to request vertical profile data.
Default: "1900-01-01"
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
Date in format `YYYY-MM-DD` until which to request vertical profile data.
Default: "2100-01-01"
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
list of data types to request from the ccycle ftp server.
List of locations to force-download data for. These will be downloaded even at times where no instrument in the metadata is located there.
Default: []
ggg_profiles_downloader.ggg2020_standard_sites[i]
Section titled “ggg_profiles_downloader.ggg2020_standard_sites[i]”Identifier of the standard site used on the ginput server.
Min. Length: 1
Minimum: -90
Maximum: 90
Minimum: -180
Maximum: 180
Date in format `YYYY-MM-DD` from which this standard site is active.
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
Date in format `YYYY-MM-DD` until which this standard site is active. Default is yesterday.
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
Retrieval
Section titled “Retrieval”retrieval.general
Section titled “retrieval.general”How many parallel processes to dispatch. There will be one process per sensor-day. With hyper-threaded CPUs, this can be higher than the number of physical cores.
Default: 1
Minimum: 1
Maximum: 1024
How much information the retrieval queue should print out. In `verbose` mode it will print out the full list of sensor-days for each step of the filtering process. This can help when figuring out why a certain sensor-day is not processed.
Default: "compact"
Directory to store the containers in. If not set, it will use `./data/containers` inside the pipeline directory. If your system has enough memory, you could also use `/dev/shm` which is a memory-based file system where files are stored in memory and never written to disk.
Default: null
retrieval.jobs.i
Section titled “retrieval.jobs.i”Which retrieval algorithms to use. Proffast 2.X uses the Proffast Pylot under the hood to dispatch it. Proffast 1.0 uses a custom implementation by us similar to the Proffast Pylot.
Which vertical profiles to use for the retrieval.
Sensor ids to consider in the retrieval.
Min. Items: 1
Date string in format `YYYY-MM-DD` from which to consider data in the storage directory.
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
Date string in format `YYYY-MM-DD` until which to consider data in the storage directory. Default is yesterday.
Regex Pattern: "^\d{4}-\d{2}-\d{2}$"
Whether to store the binary spectra files. These are the files that are used by the retrieval algorithm. They are not needed for the output files, but can be useful for debugging.
Default: false
Value used for the `DC_min` threshold in Proffast. If not set, defaults to the Proffast default.
Default: 0.05
Minimum: 0.001
Maximum: 0.999
Value used for the `DC_var` threshold in Proffast. If not set, defaults to the Proffast default.
Default: 0.1
Minimum: 0.001
Maximum: 0.999
Whether to use the local pressure in the pcxs files. If not used, it will tell PCXS to use the pressure from the atmospheric profiles (set the input value in the `.inp` file to `9999.9`). If used, the pipeline computes the solar noon time using `skyfield` and averages the local pressure over the time period noon-2h to noon+2h.
Default: false
Whether to use the ifg corruption filter. This filter is a program based on `preprocess4` and is part of the `tum-esm-utils` library: https://tum-esm-utils.netlify.app/api-reference#tum_esm_utilsinterferograms. If activated, we will only pass the interferograms to the retrieval algorithm that pass the filter - i.e. that won't cause it to crash.
Default: true
Maps sensor IDs to ILS correction values. If not set, the pipeline will use the values published inside the Proffast Pylot codebase (https://github.com/coccon/proffastpylot).
Default: {}
Suffix to append to the output folders. If not set, the pipeline output folders are named `sensorid/YYYYMMDD/`. If set, the folders are named `sensorid/YYYYMMDD_suffix/`. This is useful when having multiple retrieval jobs processing the same sensor dates with different settings.
Default: null
Maps sensor IDS to pressure calibration factors. If not set, it is set to 1 for each sensor. `corrected_pressure = input_pressure * calibration_factor + calibration_offset`
Default: {}
Maps sensor IDS to pressure calibration offsets. If not set, it is set to 0 for each sensor. `corrected_pressure = input_pressure * calibration_factor + calibration_offset`
Default: {}
Exports
Section titled “Exports”The pipeline can bundle a set of retrieval results into GEOMS files (one file per sensor and day) or into bundle files (one file per sensor). You can define a list of bundle/geoms targets to be generated. It uses the metadata and data paths defined in the input/output section above.
bundle_exports[i]
Section titled “bundle_exports[i]”Directory to write the bundeled outputs to. You should use absolute paths, but if you need relative paths, then this is relative to the caller of the CLI or the pipeline's entrypoint.
List of output formats to write the merged output files in. Allowed values are `csv` and `parquet`.
UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` from which to bundle data
Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"
UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` to which to bundle data
Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"
The retrieval algorithms for which to bundle the outputs
The atmospheric profile models for which to bundle the outputs
The sensor ids for which to bundle the outputs
Suffix to append to the output bundles.
Default: null
Min. Length: 1
When you ran the retrieval with a custom suffix, you can specify it here to only bundle the outputs of this suffix. Use the same value here as in the field `config.retrieval.jobs[i].settings.output_suffix`.
Default: null
Whether to parse the DC timeseries from the results directories. This is an output only available in this Pipeline for Proffast2.4. We adapted the preprocessor to output the DC min/mean/max/variation values for each record of data. If you having issues with a low signal intensity on one or both channels, you can run the retrieval with a very low DC_min threshold and filter the data afterwards instead of having to rerun the retrieval.
Default: false
Whether to parse the retrieval diagnostics from the results directories - `niter`, `rms`, and `scl` for each retrieval job.
Default: false
geoms_exports[i]
Section titled “geoms_exports[i]”The sensor ids for which to generate the GEOMS outputs
The retrieval algorithms for which to generate the GEOMS outputs
The atmospheric profile models for which to generate the GEOMS outputs
UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` from which to generate GEOMS data
Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"
UTC datetime in format `YYYY-MM-DDTHH:MM:SSZ` or `YYYY-MM-DDTHH:MM:SS+0000` to which to generate GEOMS data
Regex Pattern: "^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:Z|\+0000)$"
Whether to parse the DC timeseries from the results directories. This is an output only available in this Pipeline for Proffast2.4. We adapted the preprocessor to output the DC min/mean/max/variation values for each record of data. If you having issues with a low signal intensity on one or both channels, you can run the retrieval with a very low DC_min threshold and filter the data afterwards instead of having to rerun the retrieval.
Default: false
Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCO2 records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).
Default: 0.05
Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCH4 records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).
Default: 0.05
Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XH2O records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).
Default: 0.05
Only considered if `parse_dc_timeseries` is set. Minimum DC value to consider for XCO records in the GEOMS outputs. It not set, it uses the default value of Proffast (0.05).
Default: 0.05
Maximum solar zenith angle to consider in the GEOMS outputs. If not set, it will consider all solar zenith angles.
Default: null
Minimum XAIR required to consider in the GEOMS outputs. If not set, it will consider all XAIR values.
Default: null
Maximum XAIR required to consider in the GEOMS outputs. If not set, it will consider all XAIR values.
Default: null
What to do if an output file already exist.
Default: "replace"
Minimum number of data points per day required to generate a GEOMS file for that day. If not enough data points are available, no GEOMS file will be generated for that day.
Default: 11
Minimum: 1