Pipeline

pyTooling.CI models a CI pipeline independently of the service running it - the tree of workflows, matrices, jobs and steps, and the dependencies between them:

from pyTooling.CI import Pipeline, Job, Matrix, MatrixJob, Workflow

pipeline = Pipeline("Pipeline")
prepare =  Job("Prepare", parent=pipeline)
test =     Matrix("Test", parent=pipeline)
package =  Workflow("Package", reference="./.github/workflows/Package.yml", parent=pipeline)
for version in ("3.13", "3.14"):
  MatrixJob("Test", {"python": version}, parent=test)

test.AddNeed(prepare)
package.AddNeed(test)

graph = pipeline.ToGraph()        # a pyTooling.Graph.Graph of the elements and their dependencies

A service’s model derives from these classes and adds what only that service has - identifiers, URLs, its own status values, the interface of a reusable workflow.

The Tree

PipelineGroup            the pipelines started for one commit
+-- Pipeline             a pipeline
    +-- Workflow         a called workflow or a child pipeline
    |   +-- ...          the same elements a pipeline contains
    +-- Matrix           a matrix
    |   +-- MatrixJob        one job instance it produced
    |   +-- MatrixWorkflow   one instance of a called workflow it produced
    +-- Job              a job
        +-- Step         a step of that job

An element is created with its parent, which adds it to its elements: a group’s Elements, a job’s Steps or a pipeline group’s Pipelines. A group keeps one sequence of elements of every kind, in the order they were added - for a definition, the order of its file. Jobs, Workflows and Matrices select one kind from it. Every element knows its Parent and the Pipeline it belongs to. A wrong parent - a step below a workflow - is a TypeError.

IterateElements() yields what a group holds one level down, ordered by creation time; elements without a time keep the order they were added in. IterateJobs() reaches every job below a group. An element is asked for and looked up by its name, str(element) - pipeline.ContainsElement("Build"), pipeline.GetElement("Build"), matrix.GetElement("Test (3.14)") - which is how a reader resolves the names a definition refers to. A job offers the same for its steps, a pipeline group for its pipelines:

Class

Count

Check

Look up

Iterate

PipelineGroup

PipelineCount

ContainsPipeline()

IteratePipelines()

JobGroup

ElementCount

ContainsElement()

GetElement()

IterateElements()

Job

StepCount

ContainsStep()

IterateSteps()

QualifiedName names an element by the workflows containing it - Package / Build, or Test (3.14) for a matrix instance, whose name carries its matrix’ name already.

Definition and Run

The model holds what a pipeline’s definition says and what a run reports, as far as every service has it:

A group the service reports as an element of its own - a pipeline, a GitLab child pipeline; one given a time or an outcome - keeps the times and the outcome it was given. A group it doesn’t report - a GitHub called workflow, a matrix - spans what it holds: it begins with its earliest element and ends with its latest, and has no end while an element below it is still running. The span is available for every group as ContentsCreatedAt, ContentsStartedAt, ContentsCompletedAt and ContentsOutcome; Outcome.Combine lets the worst outcome win.

Key-Value Pairs

Every element - a pipeline group included - carries a dictionary of arbitrary key-value-pairs, like the elements of pyTooling.Graph. A consumer attaches what the model has no field for, without deriving its own classes. The pairs are given when an element is created, with keyValuePairs, and accessed with the element’s dictionary operators:

from pyTooling.CI import Job

job = Job("Build", keyValuePairs={"runner.os": "Linux"})
job["runner.arch"] = "x64"

"runner.os" in job           # True
job["runner.os"]             # "Linux"
len(job)                     # 2
list(job)                    # ["runner.os", "runner.arch"]
del job["runner.arch"]

The operators address the key-value-pairs only. What an element contains is reached by name - ContainsElement, GetElement, IterateElements - see The Tree.

Dependencies

The elements one level below a workflow - jobs, matrices and called workflows - can need each other, and so can the pipelines of a group. AddNeed() links two of them and records the reverse link, so Needs and Dependents are always consistent:

release.AddNeed(package)

package in release.Needs          # True
release in package.Dependents     # True

Needs and dependents can also be given when an element is created, e.g. Job("Release", parent=pipeline, needs=[package]). A group needing another group needs everything that group contains.

  • A need is a sibling - an element of the same group. A job can’t need a job inside a called workflow; it needs the workflow. Anything else raises NeedDependencyError when the need is added.

  • Cycles are found once the pipeline is complete. JobGroup.Validate searches a group and every group it contains, PipelineGroup.Validate also the pipelines of a group, each element and need once. A cycle raises NeedDependencyCycleError, whose note names it - Cycle: A -> D -> C -> A.

Conversion to a Graph

ToGraph() converts a pipeline or a called workflow into a pyTooling.Graph.Graph:

  • Every element one level below becomes a vertex. Its ID and its Value are the element, so graph.GetVertexByID(job) finds a job’s vertex. A vertex has no name; label it by vertex.Value.QualifiedName.

  • Every dependency becomes an edge from the element needing to the element it needs: an edge reads needs. IterateTopologically() therefore yields the elements in an order they can run in.

  • A called workflow or a matrix holding elements is expanded into a Subgraph, named by its qualified name and built the same way. The group’s vertex has a Link to each vertex of its subgraph. depth limits how many levels are expanded; 0 expands none.

Dependencies only link siblings, so every edge lies within one graph or subgraph, and the graph algorithms of pyTooling.Graph apply to each of them. reduce - on by default - applies the transitive reduction (RemoveTransitiveEdges()) to the graph and every subgraph: a dependency a longer path already implies - Release needing Prepare although it needs Test, which needs Prepare - has no edge. The model itself keeps every dependency; reduce=False gives each of them an edge.

graph = pipeline.ToGraph(depth=1)                 # reduced
every = pipeline.ToGraph(depth=1, reduce=False)   # an edge per dependency

for vertex in graph.IterateTopologically():
  print(f"{vertex.Value.QualifiedName}: {type(vertex.Value).__name__}")

Note

pyTooling.Graph registers a subgraph’s vertices and edges on the subgraph, so the graph’s own VertexCount counts the top level only.

Services

A service’s model derives its classes from these and adds what only the service reports. Its matrix instance derives from its own job class and mixes in MatrixInstanceMixin, which carries the dimensions - as MatrixJob does with Job, and MatrixWorkflow with Workflow.

A matrix instance is one combination of the matrix’ variables: its Dimensions maps each dimension’s name to the value it ran with, in the matrix’ order - {"os": "ubuntu-26.04", "python": "3.14"}. Its name prints the values only, as a service does: Test (ubuntu-26.04, 3.14).

GitHub Actions

GitHub Actions

pyTooling.CI

Runs of one commit

PipelineGroup

Workflow run (entry-point workflow file)

Pipeline

Job with uses: (calls a reusable workflow)

Workflow, uses: as Reference; without contents, if the called file isn’t read

Job with strategy.matrix

Matrix, an instance per combination as MatrixJob - or MatrixWorkflow, if the job calls a reusable workflow

Job with steps:

Job, its steps as Step

needs:

AddNeed(); needing a matrix job or a calling job needs the group

if:

Condition

conclusion

Outcome (e.g. timed_out → Timeout, startup_failure → Error)

GitLab CI

GitLab CI

pyTooling.CI

Pipelines of one commit (branch, merge request)

PipelineGroup

Pipeline (.gitlab-ci.yml)

Pipeline

stages:

No element. A job of stage N without needs: needs every job of the nearest earlier stage holding jobs; the reader adds these dependencies. The stage stays a fact of the GitLab job.

needs: (DAG)

AddNeed(), replacing the stage’s implicit dependencies; needs: [] needs nothing

needs:parallel:matrix (some instances of a matrix)

A need of the whole Matrix

needs:pipeline, needs:project (artifacts of another pipeline)

No dependency; they cross the pipeline’s boundary

Parent-child pipeline (trigger:include)

Workflow named after the trigger job, the child pipeline’s jobs below it; it reports its own times

Multi-project pipeline (trigger:project)

Workflow with the project as Reference and no contents

parallel:matrix

Matrix, an instance per combination (test: [3.14, linux]); MatrixWorkflow instances for a trigger job

parallel: N

Matrix with N instances (test 1/3)

rules:if

Condition

Job status

Outcome (e.g. canceled → Cancellation); manual and created haven’t ended

Competing Solutions

No package on PyPI models a CI pipeline independently of the service running it. The solutions below are a vocabulary or an event format shared by services, a client of one service, or a pipeline to run on one engine.

OpenTelemetry CI/CD Semantic Conventions

Source: Semantic conventions for CI/CD, in Python as opentelemetry-semantic-conventions.

Disadvantages

  • Names of attributes - cicd.pipeline.name, cicd.pipeline.task.run.result - for spans and metrics, not a tree of elements. There are no dependencies and no matrices.

Standoff

  • pyTooling.Tracing.CI writes a pipeline as spans with these attributes, and Outcome takes its values from cicd.pipeline.task.run.result.

CDEvents

Source: CDEvents of the CD Foundation, in Python as sdk-python.

Disadvantages

  • Events about a pipeline run or a task run - queued, started, finished - for one service to tell another. Neither nested workflows, matrices nor dependencies are described.

  • The Python SDK isn’t on PyPI.

Advantages

  • An event format several services and tools send already.

Service Clients

Source: PyGithub (WorkflowRun, WorkflowJob), python-gitlab (ProjectPipeline, ProjectPipelineJob, bridges).

Disadvantages

  • Each client models its own service’s REST API, so code reading a pipeline is written once per service.

  • GitHub’s REST API doesn’t report a job’s needs, so a workflow run has no dependencies.

Advantages

  • Every field of the service’s API is available, and the client can act - cancel a run, rerun a job.

Pipeline Engines

Source: Hera for Argo Workflows, Tekton with its Python SDK tekton-pipeline (last release 2021).

Disadvantages

  • A pipeline is defined in Python to be run on one engine. A run on another service is not read into it.

Advantages

  • Tasks depend on each other - Hera’s >>, Tekton’s runAfter - like the dependencies of this model.