LLMOps

LLM Upgrades: Let the Eval Gate Decide

Anthropic and OpenAI both cut prices on 22 September, but Opus 5.5 also turns some Opus 5 requests into 400s and runs at a lower default effort. Pin every route to an exact version, replay its real requests against the candidate, and move the alias only when the eval gate is green.

Cloud X Ops TeamDevOps & AI Integration
September 26, 2026
14 min read

Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol both launched on 22 September. Opus 5.5 lists at $4/$20 per million input and output tokens, down from $5/$25 for Opus 5, and GPT-6 Sol at $2/$10, half GPT-5.6 Sol's price. Claude Code 2.1.280 made Opus 5.5 its default Opus on launch day, and by 25 September both were on Databricks' Unity Gateway and in GitHub Copilot.

So for anyone on Claude Code's default Opus, the model changed before they chose to change it, and in Copilot or on Unity Gateway it is one setting away. Treat a model like any other dependency, one that moves only when a test on your own traffic passes.

A cheaper model can still return 400

Anthropic's Opus 5.5 migration guide lists four changes that turn working Opus 5 requests into errors, two of them only in some setups.

  • Thinking cannot be turned off. thinking: {"type": "disabled"} and a manual budget_tokens both return 400 at every effort level.
  • Forced tool use is gone. A tool_choice of any or a named tool returns 400. Use auto, add strict tool use (or structured outputs) where you need schema-valid input, and say in the prompt when the tool applies.
  • Computer use moved to a toolset. computer_20251124 is rejected on the Claude API and Google Cloud in favour of computer_toolset_20260801. Amazon Bedrock still accepts it.
  • Thinking blocks are tied to the conversation. For accounts created on or after 31 August 2026, replaying one after editing the system prompt, the tools or an earlier message returns 400 by default. Append-only code needs no change.

OpenAI has one of its own: the GPT-6 model pages say Chat Completions supports function calling only with reasoning_effort at none.

Two changes fail nothing, which makes them easy to miss. Default effort is now medium where Opus 5 used high, so a request that never set effort runs one level lower. And notes written between tool calls now arrive as thinking blocks, empty by default: an agent UI that streamed them goes quiet.

Pin the version, move the alias

Application code should never name a model. Each route calls the gateway by an alias such as support-triage, and the gateway maps it to one exact model ID. Upgrade and rollback are then one reviewed line each.

Coding agents need pins too. This week's controls:

  • Exact allow lists. Claude Code 2.1.283 added availableModelsMatch. Set to "exact", an availableModels entry allows only the version it names, so new releases stay blocked until listed. deniedModels blocks a model even if the allow list includes it.
  • Platform-owned telemetry. Since 2.1.282, project and local settings ignore the OpenTelemetry variables that switch export on or set its endpoint, so the platform team, not a settings file in a repo, decides where the record of which model ran goes. GitHub now exports Copilot agent activity to OpenTelemetry through enterprise-managed settings.

Then prove the config loaded. In a post this week on blog.szypowi.cz, Przemek showed that Claude Code's new AGENTS.md loader was skipped with telemetry off, which hit Bedrock, Vertex and gateway users. He caught it with a canary word, and 2.1.281 fixed it the same day. A canary instruction checked in CI catches the next one.

Replay real traffic before the alias moves

Keep a golden set per route, real requests recorded with their full shape and what a correct answer must contain, and replay it against the candidate in CI. That one replay is both the contract test, because a rejected shape fails outright, and the eval, because every answer is graded. Run each case three times and count it only if all three pass.

ci/model_gate.py
import json, statistics, sys, anthropic  # runs in CI against each candidate

CANDIDATE = "claude-opus-5-5"  # an exact model ID, never "latest"
EFFORT = "medium"              # set explicitly, not inherited
IN_USD, OUT_USD = 4.00, 20.00  # per 1M tokens; thinking bills as output
CW_USD, CR_USD = 5.00, 0.20    # 5-minute cache writes, cache reads
K, GATE = 3, 0.95              # pass^3: a case needs all three runs
client = anthropic.Anthropic(timeout=1800)  # else 64k max_tokens raises
cases = [json.loads(line) for line in open("golden/support-triage.jsonl")]
passed, done, cost, thinking, broken = 0, 0, 0.0, [], []
for case in cases:
    req = dict(case["request"], model=CANDIDATE)  # recorded prod shape
    req["output_config"] = dict(req.get("output_config", {}), effort=EFFORT)
    try:
        runs = [client.messages.create(**req) for _ in range(K)]
    except anthropic.BadRequestError as err:  # contract: shape rejected
        broken.append(f"{case['id']}: {err.body['error']['message']}")
        continue
    ok = 0
    for msg in runs:
        u = msg.usage  # what was delivered, not what was configured
        cw = u.cache_creation_input_tokens or 0
        cr = u.cache_read_input_tokens or 0
        cost += (u.input_tokens * IN_USD + cw * CW_USD + cr * CR_USD
                 + u.output_tokens * OUT_USD) / 1e6
        d = u.output_tokens_details  # None when the API omits it
        if d: thinking.append(d.thinking_tokens)
        text = "".join(b.text for b in msg.content if b.type == "text")
        ok += case["expect"] in text
    passed, done = passed + (ok == K), done + ok
if broken: sys.exit("contract broken:\n" + "\n".join(broken))
score = passed / len(cases)
print(f"pass^{K} {score:.1%}  cost/completed task ${cost / max(done, 1):.4f}")
print(f"median thinking {statistics.median(thinking or [0]):.0f} tokens")
sys.exit(0 if score >= GATE else 1)  # exit 1 blocks the upgrade

Thinking comes from usage.output_tokens_details.thinking_tokens, which counts billed reasoning even when its text is omitted. Cache writes and reads are priced on their own lines because the API reports them apart from input_tokens. Cost is divided by completed runs, so failed runs count against the model that caused them. Without the explicit timeout the Python SDK refuses a non-streaming request with max_tokens above about 21,000, and Anthropic suggests starting at 64k for Opus 5.5 at xhigh or max.

A fallback runs without the reasoning

On the Claude API only Fable 5.1 and Mythos 5.1 read Opus 5.5 thinking blocks. A gateway that falls back to Opus 5, or retries on any other model, gets a normal response, but the turns after the switch run without that reasoning. Nothing errors. Put fallback routes through the same gate.

Watch an upgrade meet the gate

Press Run gate below. The support-triage alias is pinned to claude-opus-5; the candidate is claude-opus-5-5 at medium effort. Three of the route's four request shapes, replayed on the Claude API as it sends them today, return 400, and the gate blocks. Then click each row that returned 400 to migrate it. Each click replays that shape, and once the last 400 is gone the golden set clears the 95% gate over three runs per case and the alias moves.

Model Gate · support-triage
Contract · Claude API4 shapes
c-history append-only history
Click a row to migrate its shape
Golden set
pass^3waitingon the contract96.5%193/200
cost/taskopus-5-5n/aopus-5$0.0816
0 breaks · pass^3 ≥ 95%not run
# press Run gate or click a row
Gateway alias
aliassupport-triage
claude-opus-5pinned · effort high
claude-opus-5-5candidate · effort medium
The four rows are the distinct request shapes in the route's 200 golden cases on the Claude API. Click one to migrate it. Scores, costs and token counts are illustrative.

What a finished task costs

Anthropic's tests put Opus 5.5 at 40% less than Opus 5 on typical workloads at default settings, without saying how that figure splits. List prices fell, 20% for input and output and 60% for cache reads, which Anthropic says are most of what agentic and coding work costs, and the default effort dropped from high to medium in the same release. Anthropic also says Opus 5.5 thinks more per turn than Opus 5 at the same effort level, so how much of the saving survives on your routes is a question for your gate.

With assumed, illustrative counts, an uncached request reading 6,000 tokens and writing 2,000 costs $0.080 on Opus 5. On Opus 5.5, writing the same 2,000, it costs $0.064, the plain 20% cut. If medium effort writes 1,500, it costs $0.054, about a third less. Thinking bills as output, so the effort level moves the bill alongside the list price. Agent loops that reread a long cached context gain most: 50,000 cached tokens cost $0.025 to read on Opus 5 and $0.010 on Opus 5.5.

OpenAI's halving has fine print. Its changelog listed GPT-5.6 Sol's $4/$20 as a promotional price, so the cut starts from a discount; a spokesperson told VentureBeat the new prices are permanent. Prompts over 272K input tokens pay double for input and cached input and one and a half times for output, on the whole request.

Watch the cache when you tune effort. On Anthropic's API a top-level effort change invalidates cache breakpoints, and Opus 5.5 keeps the cache through per-message effort and mid-conversation tool changes, both in beta. VentureBeat reports GPT-6 can change effort or tools without invalidating cached context.

Last, track the reasoning you get. Anthropic's docs say tool-result follow-ups can skip thinking even at xhigh, so one empty turn means nothing: alert on shifts in the thinking_tokens distribution per route.

The price tells you what a token costs. Only your own traffic tells you what the model will do with it.

Takeaways

  • Opus 5.5 turns four Opus 5 request shapes into 400s, two of them only in some setups.
  • Call models through a gateway alias pinned to an exact ID, so an upgrade and its rollback are each one reviewed line.
  • Replay each route's recorded requests against the candidate in CI, fail the gate on any 400, and count a case only when all three runs pass.
  • Compare cost per completed task at the effort you will ship. Thinking bills as output, so test Anthropic's 40% on your own routes before you budget with it.
  • Pin coding agents too, with exact allow lists, platform-owned telemetry and a canary that proves instruction files loaded.

Moving production traffic to a new model?

We put the gateway, the pinned aliases, the contract tests and the eval gates in place, so the next model upgrade ships like any other dependency bump and rolls back like one too.

Plan your next upgrade
upgrade.sh
SECURE
cloudxops@llmops:~$ python ci/model_gate.py
# Replaying 200 recorded requests...
[OK] 0 contract breaks · 4 request shapes
[INFO] pass^3 96.5% · gate 95% · effort medium
[READY] support-triage -> claude-opus-5-5
$ ▋
Upgrade safety