Pipeline
pyTooling.CI models a CI pipeline independently of the service running it - the tree of
workflows, matrices, jobs and steps, and the dependencies between them:
from pyTooling.CI import Pipeline, Job, Matrix, MatrixJob, Workflow
pipeline = Pipeline("Pipeline")
prepare = Job("Prepare", parent=pipeline)
test = Matrix("Test", parent=pipeline)
package = Workflow("Package", reference="./.github/workflows/Package.yml", parent=pipeline)
for version in ("3.13", "3.14"):
MatrixJob("Test", {"python": version}, parent=test)
test.AddNeed(prepare)
package.AddNeed(test)
graph = pipeline.ToGraph() # a pyTooling.Graph.Graph of the elements and their dependencies
A service’s model derives from these classes and adds what only that service has - identifiers, URLs, its own status values, the interface of a reusable workflow.
The Tree
PipelineGroup the pipelines started for one commit
+-- Pipeline a pipeline
+-- Workflow a called workflow or a child pipeline
| +-- ... the same elements a pipeline contains
+-- Matrix a matrix
| +-- MatrixJob one job instance it produced
| +-- MatrixWorkflow one instance of a called workflow it produced
+-- Job a job
+-- Step a step of that job
An element is created with its parent, which adds it to its elements: a group’s
Elements, a job’s Steps or a pipeline
group’s Pipelines. A group keeps one sequence of elements of every kind,
in the order they were added - for a definition, the order of its file. Jobs,
Workflows and Matrices select one
kind from it. Every element knows its Parent and the
Pipeline it belongs to. A wrong parent - a step below a workflow - is a
TypeError.
IterateElements() yields what a group holds one level down, ordered by creation
time; elements without a time keep the order they were added in. IterateJobs()
reaches every job below a group. An element is asked for and looked up by its name, str(element) -
pipeline.ContainsElement("Build"), pipeline.GetElement("Build"),
matrix.GetElement("Test (3.14)") - which is how a reader resolves the names a definition refers to. A job
offers the same for its steps, a pipeline group for its pipelines:
Class |
Count |
Check |
Look up |
Iterate |
|---|---|---|---|---|
QualifiedName names an element by the workflows containing it -
Package / Build, or Test (3.14) for a matrix instance, whose name carries its matrix’ name already.
Definition and Run
The model holds what a pipeline’s definition says and what a run reports, as far as every service has it:
Condition- the condition under which a workflow, a matrix, a job or a step runs, as written (GitHubif:, GitLabrules:if). It is not evaluated.Reference- what a called workflow calls, as written.CreatedAt,StartedAt,CompletedAtandOutcome- the times and theOutcomeof a run.
A group the service reports as an element of its own - a pipeline, a GitLab child pipeline; one given a time or an
outcome - keeps the times and the outcome it was given. A group it doesn’t report - a GitHub called workflow, a
matrix - spans what it holds: it begins with its earliest element and ends with its latest, and has no end while an
element below it is still running. The span is available for every group as
ContentsCreatedAt, ContentsStartedAt,
ContentsCompletedAt and ContentsOutcome;
Outcome.Combine lets the worst outcome win.
Key-Value Pairs
Every element - a pipeline group included - carries a dictionary of arbitrary key-value-pairs, like the elements
of pyTooling.Graph. A consumer attaches what the model has no field for, without deriving its own classes. The
pairs are given when an element is created, with keyValuePairs, and accessed with the element’s dictionary
operators:
from pyTooling.CI import Job
job = Job("Build", keyValuePairs={"runner.os": "Linux"})
job["runner.arch"] = "x64"
"runner.os" in job # True
job["runner.os"] # "Linux"
len(job) # 2
list(job) # ["runner.os", "runner.arch"]
del job["runner.arch"]
The operators address the key-value-pairs only. What an element contains is reached by name - ContainsElement,
GetElement, IterateElements - see The Tree.
Dependencies
The elements one level below a workflow - jobs, matrices and called workflows - can need each other, and so can
the pipelines of a group. AddNeed() links two of them and records the
reverse link, so Needs and
Dependents are always consistent:
release.AddNeed(package)
package in release.Needs # True
release in package.Dependents # True
Needs and dependents can also be given when an element is created, e.g. Job("Release", parent=pipeline,
needs=[package]). A group needing another group needs everything that group contains.
A need is a sibling - an element of the same group. A job can’t need a job inside a called workflow; it needs the workflow. Anything else raises
NeedDependencyErrorwhen the need is added.Cycles are found once the pipeline is complete.
JobGroup.Validatesearches a group and every group it contains,PipelineGroup.Validatealso the pipelines of a group, each element and need once. A cycle raisesNeedDependencyCycleError, whose note names it -Cycle: A -> D -> C -> A.
Conversion to a Graph
ToGraph() converts a pipeline or a called workflow into a
pyTooling.Graph.Graph:
Every element one level below becomes a vertex. Its
IDand itsValueare the element, sograph.GetVertexByID(job)finds a job’s vertex. A vertex has no name; label it byvertex.Value.QualifiedName.Every dependency becomes an edge from the element needing to the element it needs: an edge reads needs.
IterateTopologically()therefore yields the elements in an order they can run in.A called workflow or a matrix holding elements is expanded into a
Subgraph, named by its qualified name and built the same way. The group’s vertex has aLinkto each vertex of its subgraph.depthlimits how many levels are expanded;0expands none.
Dependencies only link siblings, so every edge lies within one graph or subgraph, and the graph algorithms of
pyTooling.Graph apply to each of them. reduce - on by default - applies the transitive reduction
(RemoveTransitiveEdges()) to the graph and every subgraph: a dependency a longer path
already implies - Release needing Prepare although it needs Test, which needs Prepare - has no edge.
The model itself keeps every dependency; reduce=False gives each of them an edge.
graph = pipeline.ToGraph(depth=1) # reduced
every = pipeline.ToGraph(depth=1, reduce=False) # an edge per dependency
for vertex in graph.IterateTopologically():
print(f"{vertex.Value.QualifiedName}: {type(vertex.Value).__name__}")
Note
pyTooling.Graph registers a subgraph’s vertices and edges on the subgraph, so the graph’s own
VertexCount counts the top level only.
Services
A service’s model derives its classes from these and adds what only the service reports. Its matrix instance
derives from its own job class and mixes in MatrixInstanceMixin, which carries the
dimensions - as MatrixJob does with Job, and
MatrixWorkflow with Workflow.
A matrix instance is one combination of the matrix’ variables: its
Dimensions maps each dimension’s name to the value it ran with,
in the matrix’ order - {"os": "ubuntu-26.04", "python": "3.14"}. Its name prints the values only, as a service
does: Test (ubuntu-26.04, 3.14).
GitHub Actions
GitHub Actions |
|
|---|---|
Runs of one commit |
|
Workflow run (entry-point workflow file) |
|
Job with |
|
Job with |
|
Job with |
|
|
|
|
|
|
|
GitLab CI
GitLab CI |
|
|---|---|
Pipelines of one commit (branch, merge request) |
|
Pipeline ( |
|
|
No element. A job of stage N without |
|
|
|
A need of the whole |
|
No dependency; they cross the pipeline’s boundary |
Parent-child pipeline ( |
|
Multi-project pipeline ( |
|
|
|
|
|
|
|
Job |
|
Competing Solutions
No package on PyPI models a CI pipeline independently of the service running it. The solutions below are a vocabulary or an event format shared by services, a client of one service, or a pipeline to run on one engine.
OpenTelemetry CI/CD Semantic Conventions
Source: Semantic conventions for CI/CD, in Python as opentelemetry-semantic-conventions.
Disadvantages
Names of attributes -
cicd.pipeline.name,cicd.pipeline.task.run.result- for spans and metrics, not a tree of elements. There are no dependencies and no matrices.
Standoff
pyTooling.Tracing.CIwrites a pipeline as spans with these attributes, andOutcometakes its values fromcicd.pipeline.task.run.result.
CDEvents
Source: CDEvents of the CD Foundation, in Python as sdk-python.
Disadvantages
Events about a pipeline run or a task run - queued, started, finished - for one service to tell another. Neither nested workflows, matrices nor dependencies are described.
The Python SDK isn’t on PyPI.
Advantages
An event format several services and tools send already.
Service Clients
Source: PyGithub (WorkflowRun, WorkflowJob),
python-gitlab (ProjectPipeline, ProjectPipelineJob, bridges).
Disadvantages
Each client models its own service’s REST API, so code reading a pipeline is written once per service.
GitHub’s REST API doesn’t report a job’s
needs, so a workflow run has no dependencies.
Advantages
Every field of the service’s API is available, and the client can act - cancel a run, rerun a job.
Pipeline Engines
Source: Hera for Argo Workflows, Tekton with its Python SDK tekton-pipeline (last release 2021).
Disadvantages
A pipeline is defined in Python to be run on one engine. A run on another service is not read into it.
Advantages
Tasks depend on each other - Hera’s
>>, Tekton’srunAfter- like the dependencies of this model.