CI Throughput

CI at Agent Speed: Cut the Setup Before You Shard

Coding agents write code faster than CI was sized to check it. Linear cut its PR wait to just over five minutes while its test suite almost quadrupled. Measure the wait, cut the setup every job repeats, and only then add shards.

Cloud X Ops TeamDevOps & AI Integration
September 27, 2026
12 min read

Linear's test suites almost quadrupled this year, yet its pull request wait fell from more than six minutes to just over five, and runner time per test roughly halved. Linear's Mufeez Amjad wrote up how on 21 September: “Agents have made it exponentially faster to ship code, but validating those changes hasn't quite kept up at the same rate.” The post drew 316 points and 407 comments on Hacker News.

Anthropic had already reported a 25x rise in CI jobs over six months on 14 September, and Dagger's Solomon Hykes published “The Great CI Bottleneck of 2026” a day later. Linear's visible gain is about a minute; without the work, it says, today's suite would take roughly 11 minutes, close to double the wait now.

Measure the wait, not the bill

Linear optimised two numbers: PR wait and runner time. You can compute both from run and job timestamps in the Actions REST API, plus the Checks API for required checks that other apps post.

  • Wait per PR. From the push to the last required check going green.
  • Critical path. The jobs that wait depends on. Linear's eight API test shards wait for small gate jobs that check which paths changed and whether the same inputs already passed, so seconds there cost every PR.
  • Runner-minutes per merged PR. The cost side of the trade: every extra shard adds its setup to it.

Then judge each change like for like. Linear compared the two days either side of its move to third-party runners, whose provider it does not name: jobs ran 34% faster on average and tsc time fell 52%. Switching to tsgo, the native TypeScript compiler, cut the weekly median of the tsc check by 73%. Rewriting a handful of custom lint rules to work without type information let ESLint drop TypeScript, cutting API lint time by 68% and full-repository lint by 55%.

Faster runners can sit further from GitHub

After the switch, checkouts slowed and sometimes hung; the provider traced it to intermittent degradation on the direct IP link its runners use to reach GitHub. Linear's own checkout action now retries with backoff, uses GIT_HTTP_LOW_SPEED_LIMIT and GIT_HTTP_LOW_SPEED_TIME to abort a stalled fetch after about 30 seconds, and reads from a git mirror on a sticky disk.

Cut the setup every job repeats

Every job pays again for booting, checkout and installs. Seconds of real work can cost minutes of runner time. Linear's cuts:

  • Check out only what the job reads. Capping fetch depth took the slowest gate from 94 seconds to 20, and jobs that never needed a working tree lost checkout, from 27 seconds to 7.
  • Bake tools into the image. Each test shard spent 7-8 seconds installing the Postgres client with apt; now it ships in a small base image with Node.
  • Install one package. Filtering pnpm install to the API package cut it from 44-73 seconds to 16-18. Linear also tried a node_modules cache: a hit took about 28 seconds to restore, against 7.5 for a filtered install.
  • Load a schema snapshot. For PRs that leave the schema alone, Linear loads a generated snapshot instead of replaying every migration, and database setup fell from about 12 seconds per container to 1-2.
  • Batch small checks. Seven checks that each paid for a runner, checkout and install became two jobs, saving roughly 87,000 runner-minutes a month on June usage, 11.8% of the total.

The base image, the filtered install and going without a node_modules cache cut per-shard setup by about 44%, from 110-140 seconds to 67-73. The same pattern in GitHub Actions:

.github/workflows/api-tests.yml
on: pull_request
jobs:
  changes:   # GATE: every shard waits for this job
    runs-on: ubuntu-latest
    permissions: { pull-requests: read }   # no checkout: files via the API
    outputs: { api: "${{ steps.f.outputs.api }}" }
    steps:
      - id: f
        uses: dorny/paths-filter@ceb8a2b8f2d89434be7ff52d3de7ec3738c5cc9d # v4.0.3
        with: { filters: "api: ['packages/api/**', 'pnpm-lock.yaml']" }
  test:
    needs: changes
    if: needs.changes.outputs.api == 'true'   # skipped work never takes a runner
    runs-on: ubuntu-latest
    container: ghcr.io/your-org/ci-base:node24   # SETUP: git, pnpm, psql baked in
    services:
      postgres:
        image: postgres:18
        env: { POSTGRES_PASSWORD: ci }
        options: --health-cmd pg_isready --health-interval 2s --health-retries 15
    strategy:
      matrix: { shard: [1, 2, 3, 4, 5, 6, 7, 8] }   # SHARD: once setup is small
    env:
      DATABASE_URL: "postgres://postgres:ci@postgres/postgres"
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
      - run: pnpm install -F "@your-org/api..."   # SETUP: one package, no cache
      # SETUP: load the schema snapshot instead of replaying every migration
      - run: psql "$DATABASE_URL" -v ON_ERROR_STOP=1 -qf packages/api/schema.sql
      # TEST: Vitest shards by file, so split oversized files first
      - run: pnpm -F @your-org/api exec vitest run --shard=${{ matrix.shard }}/8
  tests-passed:   # REQUIRED: the one check branch protection names
    needs: [changes, test]
    if: always()   # run even when a shard fails or the matrix is skipped
    runs-on: ubuntu-latest
    steps:
      - run: exit 1
        if: contains(needs.*.result, 'failure') ||
          contains(needs.*.result, 'cancelled')

Linear loads the snapshot only when a PR leaves the schema alone. This workflow loads it on every PR, which is simpler, but then the snapshot must track the migrations: regenerate it in any PR that adds one, and have a separate job replay them into an empty database and fail on any difference.

Do not make the shards required checks; require tests-passed instead. A job-level if skips a matrix before it expands, so no test (1) check ever arrives and the PR waits forever. A plain needs job does not fix that: it is skipped when a shard fails, and GitHub counts a skipped required check as a pass. So tests-passed runs with if: always() and fails when any job it needs failed or was cancelled.

Shard once setup is cheap

Vitest splits work by file, not by duration, so Linear first broke up the oversized files that held shards back. It had gone from three shards to four earlier in the year. Eight made the critical job roughly 19% faster and 19% cheaper in its initial benchmark, and a week later the slowest shard had dropped from 5.25 minutes to 4.33.

That worked because setup had shrunk. Every shard pays it in full, so doubling the shards doubles the setup. At 110-140 seconds a shard, Linear says, eight shards would have spent 15-19 minutes of runner time on setup alone, more than the tests themselves. With the later cuts, Linear now puts setup at around 40 seconds, so eight spend less on it than four did before.

Linear's largest single gain was an opt-in Vitest project with isolate: false, letting safe files share a module registry within a worker, worth roughly 17% in monthly savings. It also carried the highest correctness risk, so files opt in with a comment, shared state gets a teardown, and a few files with fake timers or tangled shared state stay isolated. Agents now write most of Linear's tests, and the skill files they read carry the same isolation rules.

Watch setup multiply with the shards

Press Run shards below. The numbers are our model, sized on Linear's figures rather than taken from them. The same 16 minutes of tests run three ways: four shards with 125 seconds of setup each, eight with the same setup, then eight once setup is down to 40 seconds. Each lane fills with setup, then tests; the clock stops at the slowest shard, and a meter totals the runner-minutes.

Shard Planner · api-tests.yml
setuptests
wall0:00
shard 1/4
shard 2/4
shard 3/4
shard 4/4
runner time 0.0runner-min setup 0.0 min
modelwall = setup + 960 s / N · runner time = N × setup + 960 s
# api-tests.yml · 960 s of vitest work, split evenly across the shards
Our model, sized on Linear's figures: at 125 s of setup a shard, doubling the shards buys 2 minutes for 8.3 more runner-minutes; at 40 s, eight shards beat four on wait and on cost.

Faster tests are half the answer

The sharpest reply on Hacker News was a question: “Did the tests produce four times as much value, though?” Others suspected piles of trivial agent-written tests, and one noted that the post reports no pass or fail comparison between the old checks and the faster ones.

  • Select tests per change. Anthropic already does, with test impact analysis: a listener records the results of every CI run, and a selector picks each PR's tests from that history and package relevance. Its 25x growth overloaded that service until Anthropic rebuilt it to scale out. Selection is infrastructure that has to keep up too. Keep a full run after merge to catch what selection misses.
  • Prune what agents write. The skills that teach agents your isolation rules can also ban tests that only exercise the framework.
  • Prove the faster tool agrees. Run tsgo beside tsc, and new lint rules beside old, on the same commits and compare the output.
Every shard pays for the setup again.

Takeaways

  • Track wait per PR and runner-minutes per merged PR together. A monthly bill hides what the extra minutes bought.
  • Every job and every shard pays setup in full: bake tools into the image, install only the package under test and load a schema snapshot.
  • At 125 s of setup, eight shards cost 8.3 more runner-minutes than four for 2 minutes off the wait. At 40 s they beat four on both.
  • Third-party runners can sit outside GitHub's network. Give checkout retries, a stall timeout and a local mirror.
  • Select tests per change from CI history, keep a full run after merge, and prune agent-written tests that only exercise the framework.

Is CI the slowest reviewer on your team?

We measure wait per PR and runner-minutes per merge, strip out the setup every job repeats, and size your shards to the setup that is left.

Profile your CI
ci.sh
SECURE
cloudxops@ci:~$ ./ci-profile.sh --last 30d
# Timing every job on the critical path...
[OK] setup per job measured
[INFO] runner-minutes per merged PR tracked
[READY] shards sized to the setup left
$ ▋
CI throughput