Pipeline Trace Diagram

A pipeline run read into a trace - a timespan for the run, every called workflow, matrix, job and step - is written as OpenTelemetry’s OTLP/JSON, which trace viewers read, or drawn as a Gantt chart.

Reading a Run

WorkflowRunReader reads a workflow run through the GitHub REST API, using the standard library only:

from os import getenv
from pathlib import Path
from pyTooling.GitHub.Tracing import WorkflowRunReader

reader = WorkflowRunReader("pyTooling/Actions", token=getenv("GITHUB_TOKEN"))
trace = reader.ReadRun(34937615362)      # optionally: attempt=2
trace.WriteJSONFile(Path("report/Pipeline.otlp.json"))

Inside a workflow, GITHUB_TOKEN with the actions: read permission suffices. A job can’t see itself: it is still running when it reads the run, so a timing job depends on every other job and runs last.

A request failing transiently - HTTP 429, 500, 502, 503 or 504, a timeout, or an unreachable API - is tried again, retries times (default: 3), after a pause of retryDelay seconds (default: 2), which doubles with every attempt or lasts as long as a Retry-After header demands, up to a minute. HTTP 401, 403 and 404 fail at once.

WorkflowRunTrace.FromJSON does the conversion alone, for a run and jobs that were fetched another way. It is a class method, so converting needs no reader and therefore no token. It reads both payloads into a Pipeline - see Pipeline Run - and hands that to FromPipeline(), which is the entry point when the model was built elsewhere. Reading the payloads is therefore the model’s job, and a field GitHub doesn’t document raises GitHubError - as does an answer the reader itself can’t read, so everything GitHub says that can’t be made sense of is one exception type. A request that fails is a RESTError, because nothing about GitHub’s answer was wrong - there wasn’t one.

Timespans and Attributes

The run becomes the trace, and every timespan below it is marked by Kind with a member of SpanKind. Each kind is a class of its own - a JobSpan sets ci.span.kind to job because that is what it is, and takes the attributes of a job as parameters - so a reader states values and never a key, and a reader of another service builds the same classes:

Kind

Timespan

pipeline

The workflow run, from its start to its last update once it completed.

workflow

A called workflow: the jobs named Caller / Job are grouped below a timespan Caller.

matrix

A matrix: the jobs named Job (ubuntu-26.04, 3.14) are grouped below a timespan Job, and the called workflows of Caller (3.14) / Job - a workflow each - below a timespan Caller.

queued

<job> (queued), the time a job waited for a runner, in front of the job.

job

A job, from its start to its completion.

step

A step that started, below its job.

Every timespan also carries the attributes of OpenTelemetry’s semantic conventions for CI/CD, which OTLP names as a namespace nested the way the keys are - so OTLP.CICD.Pipeline.Task.Run.ID spells cicd.pipeline.task.run.id and the path can be read to check the key. The values a result may take are Result, which are those of the model’s Outcome, so a result is the element’s outcome. What only GitHub reports is named the same way by GitHub, e.g. github.conclusion beside the result it was mapped to. A job’s timespan names its runner and the labels it was requested by, so a renderer can group waiting times per operating system, and a matrix instance additionally lists the values it was produced for in github.matrix.dimensions.

A task is named the way GitHub reports it - Caller / Build (ubuntu-26.04) - while the timespan itself is named by the part the model holds, so a timespan reads in the context its parents already give.

Each flavour builds itself from the model: JobSpan.FromJob takes a Job and produces the job’s timespan, the waiting timespan in front of it, and a timespan per step. So reading a service means mapping its model onto these classes, and everything else - the kinds, the attribute keys, and skipping what the service doesn’t report - is pyTooling.Tracing.CI’s.

What the payloads say is the model’s, including the two facts a timeline depends on: a group’s elements come in the order they were queued, and a job’s times contain its steps, because GitHub reports both in whole seconds and a step is sometimes reported as running outside the job holding it - see Pipeline Run.

GitHub reports timestamps in whole seconds. A step shorter than a second lasts zero seconds, and an end reported a second before its begin is moved to the begin.

Gantt Chart

pyTooling lays a trace out as a Gantt chart and draws it with matplotlib - see pyTooling’s rendering. ciSpanFilter() selects what a pipeline’s chart shows:

from pyTooling.Tracing.Render            import GanttLayout, ciSpanFilter
from pyTooling.Tracing.Render.Matplotlib import MatplotlibRenderer

layout = GanttLayout(trace, spanFilter=ciSpanFilter())
MatplotlibRenderer(layout).Write(Path("report/Pipeline.svg"))
  • A row per job, below the rows of the called workflow and the matrix it belongs to. The time a job waited for a runner is drawn in light gray in front of the time it ran, colored per runner image.

  • The legend carries the statistics per runner image: the number of jobs, and the shortest, average and longest waiting and running times.

  • The steps are left out - a pipeline of 52 jobs has hundreds of them. excludeSteps=StepExclusion.Skipped shows the steps that ran, excludeSteps=False every step. Skipped jobs are left out too, unless excludeSkippedJobs=False.

  • MatplotlibRenderer(layout, collapsible=True) writes an SVG file whose rows of called workflows and jobs can be collapsed and expanded with a click.

The drawing needs the extra diagram (see Gantt Charts (Optional)). The program draws the same chart with --gantt - see Drawing the run.

Gantt chart of a pyTooling.GitHub pipeline run

A run of pyTooling.GitHub’s own pipeline (run 37187783003), drawn as above.