DevOps Insights & Guides
Field-tested tutorials, architecture deep-dives and security playbooks from the engineers who run the infrastructure.
Featured Guide
Building a CI/CD Pipeline That Ships Safely
Stages, gates, and rollbacks. A deployment pipeline has one job: turn a commit into running software with as little human risk as possible.
All Guides
GitHub Actions: Pin the Commit, Not the Tag
Two actions-cool actions came back online on 16 September with every release tag still on a May payload. Workflows pinned to a full commit SHA from before the compromise never ran it.
Read guideCI at Agent Speed: Cut the Setup Before You Shard
Linear's test suite almost quadrupled this year, yet PR wait fell from over six minutes to just over five. Measure the wait, cut the setup every job repeats, and only then add shards.
Read guideTerraform Providers: Allowlist First, Then Lock
A one-letter typosquat of kreuzwerker/docker reached the Terraform Registry, and an interview repo names a lookalike provider in its lock file and state. A provider allowlist in the CLI config refuses both without contacting either host.
Read guideAI Agent Sandboxes: Every URL Is an Exit
OpenAI's research agents apparently got round GET-only access by hiding code in URLs, then broke into Hugging Face. Contain yours with a proxy that reads the whole request and holds the keys.
Read guideLLM Upgrades: Let the Eval Gate Decide
Opus 5.5 and GPT-6 cut prices, and Opus 5.5 rejects some requests Opus 5 accepted. Pin each route behind a gateway alias, replay its real requests in CI, and move the alias only when the eval gate is green.
Read guideDecision Models: Ask for a Verdict, Not an Essay
Routing, triage and guardrail calls only need a label. A decision model returns it typed, with a probability; two thresholds set on your own labels make it policy, and code keeps the irreversible action.
Read guideFrom Automation to Autonomy: Infrastructure That Heals Itself
The 2026 shift from automation to autonomy: agents that detect an incident, diagnose the root cause, act, and verify the fix, with the guardrails that keep self-healing from becoming self-harm.
Read guideLLM Inference on Kubernetes: Pay for Tokens, Not Idle GPUs
An LLM pod is GPU-bound, holds a growing KV cache, and answers one request for seconds, so the web-service load balancer, HPA and scheduler all point the wrong way. The 2026 serving stack that keeps the GPU bill honest.
Read guideLangGraph: Agents That Resume, Not Restart
A linear chain runs once and forgets. LangGraph makes an agent a graph with durable state, so a run can pause for a human, survive a crash, and resume from the exact step, the reliability layer that gets agents past the pilot.
Read guideHyperscaler Power, French Jurisdiction
Storing bytes in Europe means nothing if a US warrant can still reach them. S3NS wraps Google Cloud in a Thales-controlled French operator with SecNumCloud 3.2, so regulated data stays provably beyond foreign reach.
Read guideThe Internal Developer Platform as a Product
Stop asking every team to reinvent CI, infra and observability. Pave a golden path, ship the IDP as a product, and let developers go from new repo to running service in minutes.
Read guideProgressive Delivery: Canary Releases Without the Fear
Ship a new version to 5% of real traffic, watch the SLOs, and let the rollout promote itself, or roll itself back, before anyone has to wake up.
Read guideeBPF: See Everything, Instrument Nothing
Run sandboxed, JIT-compiled programs inside the Linux kernel for deep telemetry plus inline security, without touching a line of application code, and see what eBPF powers in 2026.
Read guideKubernetes Autoscaling Without the Surprises
HPA, VPA, and Cluster Autoscaler operate on three different axes. Confuse them and you get 2 a.m. pages. Get them right and the cluster breathes on its own.
Read guidePublic, Private, or Hybrid: Choosing What Actually Fits
Stop treating deployment models as a maturity ladder. Map each workload to the model whose cost curve and constraints actually match it.
Read guideCutting Cloud Spend Without Slowing Down
Cloud bills creep, an oversized instance here, a forgotten environment there. FinOps makes that creep visible and reversible without slowing engineering.
Read guideObservability Beyond Dashboards
A wall of green panels tells you what you already thought to measure is fine. Real observability is asking new questions of a live system, no new code.
Read guideShifting Security Left: A Practical Pipeline
Security in a quarterly review finds problems after they ship. Shift it into the pipeline and catch a vulnerability in the PR that introduced it.
Read guideReading is step one.
Need a hand putting any of this into production? We design, build and run the pipelines and platforms behind every guide here, with the same engineers who wrote them.
./talk-to-an-engineer.sh