Configuration

Module Configuration provides an abstract configuration reader.

It supports any configuration file syntax, which provides:

  • scalar elements (integer, string, …),

  • sequences (ordered lists), and

  • dictionaries (key-value-pairs).

The abstracted data model is based on a common Node class, which is derived to a Sequence, Dictionary and Configuration class.

Inheritance diagram:

Inheritance diagram of pyTooling.Configuration

Dictionary

A Dictionary represents key-value-pairs of information.

// one-liner style
{"key1": "item1", "key2": "item2", "key3": "item3"}

// multi-line style
{
  "key1": "item1",
  "key2": "item2",
  "key3": "item3"
}
# one-liner style
section_1 = {key1 = "item1", key2 = "item2", key3 = "item3"}

# section style
[section_2]
key1 = "item1"
key2 = "item2"
key3 = "item3"
# one-liner style
{key1: item1, key2: item2, key3: item3}

# multi-line style
key1: item1
key2: item2
key3: item3
<items>
  <item key="key1">item1</item>
  <item key="key2">item2</item>
  <item key="key3">item3</item>
</items>

A Dictionary node is read like a dict, and it keeps the order the document states its keys in - the abstract node stores those keys in a list, so a round-trip through the reader never reorders a file.

settings = config["settings"]

hasKey1 =   "key1" in settings        # membership
item1 =     settings["key1"]          # by key, raises KeyNotFoundError when absent
item3 =     settings.get("key3", "")  # by key, with a default
pairCount = len(settings)             # number of key-value pairs

A key the document states without a value - key: in YAML, null in JSON - reads as None, and get() returns its default only for a key that is absent. A variable ${...} referencing such a key raises PathExpressionError.

Three iterators are named alike, so nothing has to be remembered about which one plain iteration gives: IterateKeys(), IterateValues() and IterateItems(). Iterating the node itself yields its values, which is what IterateValues yields.

for key in settings.IterateKeys():
  print(key)

for value in settings:                     # the node itself yields its values
  print(value)

for key, value in settings.IterateItems():
  print(f"{key}: {value}")

Their materialized counterparts are spelled the way dict spells them - keys(), values() and items() - and return tuples. The names are deliberate: dict looks for a keys method to decide whether an object is a mapping, so dict(node) and {**node} work because they exist.

Attention

A scalar is returned as a str, whatever the document writes: 42 is "42", 1.5 is "1.5" and true is "True". Only a key stated without a value reads as None. When the document states a sub-mapping or a list, the value is another Dictionary or Sequence, not a dict or list.

Sequences

A Sequence represents ordered information items.

// one-liner style
["item1", "item2", "item3"]

// multi-line style
[
  "item1",
  "item2",
  "item3"
]
# one-liner style
section_1 = ["item1", "item2", "item3"]

# multi-line style
section_2 = [
  "item1",
  "item2",
  "item3"
]
# one-liner style
[item1, item2, item3]

# multi-line style
- item1
- item2
- item3
<items>
  <item>item1</item>
  <item>item2</item>
  <item>item3</item>
</items>

A Sequence node is read like a list: by index, with len(), and by iterating it. An index is the position the document states an item at, and a negative index counts from the end. An element is a scalar or another node, just as a dictionary node’s value is, so a sequence of mappings is iterated and each element indexed by key. Because iteration yields the elements themselves, in asks whether an element is in the sequence - a sequence has no keys to ask for.

files = config["files"]

firstFile = files[0]                 # by index
lastFile =  files[-1]                # negative indices count from the end
fileCount = len(files)               # number of elements
hasFile1 =  "path/to/file1.ext" in files

for file in files:
  print(file)

index() and count() are provided with list’s signatures, so code written against a list keeps working when it is handed a sequence node. index accepts start and stop, treats negative values as offsets from the end, and raises ValueError when nothing in the searched range matches.

Configuration

A Configuration represents the whole configuration (file) made of sequences, dictionaries and scalar information items.

{ "version": "1.0",
  "settings": {
    "key1": "item1",
    "key2": "item2"
  },
  "files": [
    "path/to/file1.ext",
    "path/to/file2.ext",
    "path/to/file3.ext"
  ]
}

Attention

Not yet implemented.

version = "1.0"

[settings]
key1 = "item1"
key2 = "item2"

files = [
  "path/to/file1.ext",
  "path/to/file2.ext",
  "path/to/file3.ext"
]
version: "1.0"
settings:
  key1: item1
  key2: item2
files:
  - path/to/file1.ext
  - path/to/file2.ext
  - path/to/file3.ext

Attention

Not yet implemented.

<?xml version="1.0" encoding="UTF-8" standalone="yes" ?>
<configuration version="1.0">
  <settings>
    <setting key="key1">item1</setting>
    <setting key="key2">item2</setting>
  </settings>
  <files>
    <file>path/to/file1.ext</file>
    <file>path/to/file2.ext</file>
    <file>path/to/file3.ext</file>
  </files>
</configuration>

A Configuration is the root node of a document, and it is itself a dictionary node - so a file is read by indexing the configuration object directly. It adds the one thing a root has that an inner node doesn’t: ConfigFile, the path it was read from.

from pathlib                      import Path
from pyTooling.Configuration.YAML import Configuration

config =  Configuration(Path("settings.yml"))
version = config["version"]

Reaching deep into a document by chained indexing is verbose, so every node also answers a path expression through QueryPath(). Its elements are separated by a colon and it is resolved relative to the node it is asked of:

item1 = config.QueryPath("settings:key1")

A key that names no element raises KeyNotFoundError, which derives from KeyError so an ordinary except KeyError still catches it. A malformed expression raises PathExpressionError, and a value the format cannot map onto the data model raises UnsupportedValueTypeError.

Attention

A configuration is read-only. Assigning to a node - config["key"] = value - raises NotImplementedError, as does renaming a key through Key. Writing a configuration file back is not implemented for any format.

Data Model

The data model is a tree of three node kinds, and the diagram below is the whole grammar: a configuration contains dictionaries and sequences, and each of those contains dictionaries and sequences again, to any depth. The leaves are the scalars the file format supports.

Every node derives from Node and knows two neighbours - its _root and its _parent - so a node handed to a function on its own can still resolve a path against the document it came from.

The abstract classes are mixins, and a concrete format supplies the parsing. That is why Node carries the two class variables DICT_TYPE and SEQ_TYPE: the abstract code has to instantiate a format’s dictionary or sequence when it descends into a document, and these are what it instantiates. A concrete implementation sets them, which is step 5 below.

Two implementations ship with pyTooling - pyTooling.Configuration.JSON and pyTooling.Configuration.YAML - and both interpolate ${...} variable references in scalar values, raising InterpolationError for a dangling $ or an unclosed reference.

        flowchart TD
  Configuration --> Dictionary
  Configuration --> Sequence
  Dictionary --> Dictionary
  Sequence --> Sequence
  Dictionary --> Sequence
  Sequence --> Dictionary
    

Creating a Concrete Implementation

Follow these steps to derive a concrete implementation of the abstract configuration data model.

  1. Import classes from abstract data model

    from . import (
      Node as Abstract_Node,
      Dictionary as Abstract_Dict,
      Sequence as Abstract_Seq,
      Configuration as Abstract_Configuration,
      KeyT, NodeT, ValueT
    )
    
  2. Derive a node, which might hold references to nodes in the source file’s parser for later usage.

    @export
    class Node(Abstract_Node):
      _configNode: Union[CommentedMap, CommentedSeq]
      # further local fields
    
      def __init__(self, root: "Configuration", parent: NodeT, key: KeyT, configNode: Union[CommentedMap, CommentedSeq]) -> None:
        Abstract_Node.__init__(self, root, parent)
    
        self._configNode = configNode
    
      # Implement mandatory methods and properties
    
  3. Derive a dictionary class:

    @export
    class Dictionary(Node, Abstract_Dict):
      def __init__(self, root: "Configuration", parent: NodeT, key: KeyT, configNode: CommentedMap) -> None:
        Node.__init__(self, root, parent, key, configNode)
    
      # Implement mandatory methods and properties
    
  4. Derive a sequence class:

    @export
    class Sequence(Node, Abstract_Seq):
      def __init__(self, root: "Configuration", parent: NodeT, key: KeyT, configNode: CommentedSeq) -> None:
        Node.__init__(self, root, parent, key, configNode)
    
      # Implement mandatory methods and properties
    
  5. Set new dictionary and sequence classes as types in the abstract node class.

    setattr(Abstract_Node, "DICT_TYPE", Dictionary)
    setattr(Abstract_Node, "SEQ_TYPE", Sequence)
    
  6. Derive a configuration class:

    @export
    class Configuration(Dictionary, Abstract_Configuration):
      def __init__(self, configFile: Path) -> None:
        with configFile.open() as file:
          self._config = ...
    
        Dictionary.__init__(self, self, self, None, self._config)
    
      # Implement mandatory methods and properties
    

Competing Solutions

pyTooling.Configuration reads a JSON or YAML file into one read-only tree of nodes, addressed by path expressions with ${...} references between values. The packages below manage the settings of an application - layering, merging, overriding, validating and writing them - which this package doesn’t.

OmegaConf

Source: omegaconf, on PyPI as omegaconf.

Disadvantages

  • Files are read and written as YAML only.

  • Depends on PyYAML and on the ANTLR runtime, pinned to version 4.9.

Standoff

  • Both resolve ${...} references to other values. OmegaConf resolves ${a.b} from the root and ${..b} from the parent; pyTooling resolves every reference from the node holding the value, ${..:b} from its parent.

Advantages

  • Configurations are merged, can be changed and set read-only on demand, and are saved back to a file.

  • A configuration can be typed by a dataclass (“structured config”), and created from key=value arguments of a command line.

  • A scalar keeps its type; pyTooling returns a number as str.

Hydra

Source: hydra, on PyPI as hydra-core.

Disadvantages

  • A framework around an application’s main function, built on OmegaConf, rather than a reader for a document.

Advantages

  • Composes a configuration from groups of files, overrides any value from the command line, and runs an application once per combination of values.

Dynaconf

Source: dynaconf, on PyPI as dynaconf.

Disadvantages

  • It manages the settings of an application - files and sources merged into one settings object - rather than reading a given document as a tree.

Advantages

  • Reads TOML, YAML, JSON, INI and Python files, and lets environment variables override every value.

  • Switches between environments, e.g. development and production, validates settings, and loads them from Vault or Redis.

  • Has no third-party dependency.

python-box

Source: Box, on PyPI as python-box.

Standoff

  • A dictionary with attribute access - box.a.b - rather than a configuration reader. With box_dots, a dotted key box["a.b"] is a path.

Advantages

  • Converts from and to JSON, YAML and TOML, and can be frozen to stay unchanged.

Confuse

Source: confuse, on PyPI as confuse.

Disadvantages

  • Reads YAML only.

Advantages

  • Layers a default file, the user’s file from the platform’s configuration directory, environment variables and command-line arguments, and checks a value’s type when it is read.