CPUMaxxing our CI Runners

by Brendan O'LearySep 30, 20266 min read
Engineering

Two processors connected by blue streams, with a few active cores on the left and all cores active on the right.

I spent a year as the product manager for GitLab CI in 2017/2018, working on caching and runner sizing. When Abir Taheer offered to show me how he made Composio's CI faster, I asked him to walk me through the data.

Every PR and commit to our main repo used to wait around seven minutes for CI. Seven minutes was enough to break your concentration without giving you time to do anything useful. Now our Vitest suite runs in three shards that each finish in about two minutes, even though the number of tests has roughly doubled since Abir started.

The chart that showed where the time went

We run CI on Depot's GitHub Actions runners, which give us CPU and memory graphs for every job.

Original Vitest run taking 6 minutes 41 seconds, with CPU near idle for more than half the job.
Before: a 6m 41s Vitest run with a long stretch of idle CPU.

For more than half the Vitest run, CPU sat near zero. We were paying for sixteen cores while the job waited on network and disk I/O.

A bigger runner is cost-effective only when the job can use the additional cores. Otherwise, you pay a higher per-minute rate without a proportional reduction in runtime.

Abir looked for flat stretches after each run and moved more setup work earlier so the tests could start sooner.

Optimized Vitest run taking 3 minutes 17 seconds, with CPU active through most of the job.
After: a 3m 17s run with CPU active through most of the job.

Start everything early and wait only when you have to

The original job ran as a straight line. Install npm packages, install Go dependencies, seed the database, start the app, run the tests. Each step waited for the one before it, even when the two had nothing to do with each other.

Abir's first change was to start every install at the top of the job using shell backgrounding. Database seeding and the rest of the setup happened while the installs were still going. When a step actually needed node_modules, it checked whether the install had finished, and by then it usually had.

That change alone took a typical PR run from about seven minutes to about four and a half.

GitHub announced native background steps on June 25, 2026. We've since moved most of this work to background: true and wait. This trimmed excerpt from apollo-vitest-e2e.yaml fetches the shared Docker image for Thermos, our Go orchestration service, while toolchain and database setup continue. Checkout, conditions, and intervening steps are omitted.

# Poll the sibling builder without a needs dependency so image retrieval
# overlaps this shard's toolchain and database setup.
- name: Fetch Thermos image
  id: thermos-build
  timeout-minutes: 16
  env:
    GH_TOKEN: ${{ github.token }}
  run: |
    set -euo pipefail
    bash .github/scripts/fetch-run-artifact.sh thermos-ci-image-vitest "Build Thermos (shared)" /tmp/thermos-artifact
    zstd -dc /tmp/thermos-artifact/thermos-ci.tar.zst | docker load
    docker image inspect thermos:ci --format 'Loaded thermos:ci ({{.Id}})'
  background: true

# Toolchain setup, database migrations, and seeding run here.

- name: Wait for Thermos image
  wait: thermos-build

background: true starts the step and lets subsequent setup steps continue. Our fetch-run-artifact.sh helper polls the sibling builder without a needs dependency that would hold up the entire shard. It downloads the image when it becomes available, then loads it into Docker.

Just before a step needs Thermos, wait: thermos-build blocks until the fetch and load finish. If the background step fails, the wait step fails too. GitHub handles that coordination, so we no longer need the old shell approach's exit-status file and polling loop.

Give your dependency cache a fallback

A lockfile change invalidates the exact dependency-cache key, even if the PR adds only one package. A fallback lets the install reuse packages from an older cache.

Our reusable Playwright workflow caches the pnpm store, where downloaded packages live. After pnpm setup, it runs these steps:

- name: Get pnpm store directory
  id: pnpm-cache
  shell: bash
  run: |
    echo "STORE_PATH=$(pnpm store path)" >> $GITHUB_OUTPUT
- name: Setup pnpm cache
  uses: actions/cache@668228422ae6a00e4ad889ee87cd7109ec5666a7 # v5.0.4
  with:
    path: ${{ steps.pnpm-cache.outputs.STORE_PATH }}
    key: ${{ runner.os }}-pnpm-store-${{ hashFiles('**/pnpm-lock.yaml') }}
    restore-keys: |
      ${{ runner.os }}-pnpm-store-

The fallback prefix omits the lockfile hash. GitHub searches the current branch first, then the default branch, as described in its cache lookup rules. There is no explicit master key. The install still runs, but can reuse cached packages and download what's missing. A successful job saves the store under the new lockfile key.

The PR that added this cache change, along with the first round of sharding, took the run from four minutes to two and a half.

Sharding Vitest and overcommitting workers

After backgrounding setup and improving the cache, further single-runner optimizations produced little benefit. The Vitest suite still ran as one job, and all of it was CPU-bound.

The integration workflow splits the suite into three shards on 16-vCPU runners. This example shows the matrix and test command with Abir's recommended fail-fast: true setting. Setup steps, environment variables, and the Doppler wrapper are omitted:

jobs:
  run-vitest-shards:
    name: run-vitest-tests (${{ matrix.shard }}/3)
    runs-on: depot-ubuntu-24.04-16
    strategy:
      fail-fast: true
      matrix:
        shard: [1, 2, 3]
    steps:
      # Setup omitted.
      - name: Run Vitest tests
        run: pnpm --filter @composio/apollo test:unit --shard=${{ matrix.shard }}/3 --maxWorkers=175%

The test:unit script invokes Vitest. Each shard gets a share of the test files. With fail-fast: true, GitHub cancels the remaining queued or running matrix jobs when one fails, avoiding more work on an already-failing run. Use fail-fast: false when you want every shard's diagnostics instead.

Abir also increased the worker count. Some tests spend part of their time waiting on database and service I/O, leaving room for other tests to run. The current workflow sets --maxWorkers=175%, or 1.75 workers per vCPU. The right ratio depends on your suite, so measure it on yours before copying ours.

With sharding and overcommit together, the Vitest job went from 6 minutes 29 seconds to three shards that each finish in about two minutes, running side by side. That's roughly the duration of setup alone in some CI pipelines.

GitHub Actions showing three successful parallel Vitest shards, with the selected shard finishing in 2 minutes 11 seconds.
Three Vitest shards run in parallel. The selected shard finished in 2m 11s.

Build once and share the artifact

Several test jobs each built Apollo from scratch, which meant doing the same heavy compile once per job.

Abir moved the Apollo build into a shared workflow that runs once at the start. It backgrounds its own setup, restores the Turbo cache, builds Apollo in about 1 minute 27 seconds, and uploads the result as an artifact. The matrix jobs download that artifact in the background as soon as they start, so Apollo is already built by the time a test needs it.

The Thermos image example above follows the same approach: a shared builder produces the image once, and each shard fetches it while its own setup continues.

For a few weeks, Abir kept a Depot tab open and checked each new run for flat stretches in the CPU line.

Get new posts in your inbox

Subscribe for the latest from the Composio blog.

Blog and newsletter updates. Privacy policy.

Share