GPT-6 got the headlines in September. If you have OpenAI calls in production, read the API changelog as well: between July and September 2026 it recorded two model generations, one retired API, new error codes and several account controls.
Jul 9
GPT-5.6 family: Sol (frontier), Terra (balanced), Luna (efficient). Programmatic Tool Calling, explicit prompt caching controls and a multi-agent orchestration beta in the Responses API.Jul 22
Hard spend limits for organizations and projects. Requests return 429 once the monthly cap is reached.Jul 30
Price cuts on GPT-5.6 Luna and Terra. "Fast mode" replaces Priority Processing.Aug 13
Ultrafast mode for GPT-5.6 Sol in limited preview, "up to 14x faster than Standard processing".Aug 26
The Assistants API shuts down. Older transcription models get a shutdown date of Feb 26, 2027.Sep 2
New error codes: 429slow_downfor traffic spikes, 503server_is_overloadedfor model overload.Sep 3
GPT-6 Astra, the most capable model. Async tool calling and mid-turn steering arrive in the Responses API with it.Sep 10
Agents API public beta with managed Codex and MCP support. Project API keys can now expire.Sep 22
GPT-6 Sol and GPT-6 Luna in the Responses and Chat Completions APIs.
The GPT-6 lineup
| Model | APIs | Input / output per 1M tokens |
|---|---|---|
| GPT-6 Astra | Responses, Chat Completions (tools: Responses only) | not listed in the changelog |
| GPT-6 Sol | Responses, Chat Completions | $2 / $10 |
| GPT-6 Luna | Responses, Chat Completions | $0.10 / $0.50 |
Astra runs on Chat Completions too, but its tool calls need the Responses API, so an agent loop built on Chat Completions has to migrate before it can use Astra's tools. On September 25 OpenAI fixed an image-encoding bug in Sol and Luna that had degraded image understanding. If you evaluated either model on images before then, run those evals again.
Changes for your codebase
Migrate off the Assistants API if anything still calls it. It shut down on August 26, and OpenAI's migration guide points to the Responses and Conversations APIs.
Handle 429s by their code. A 429 can mean an ordinary rate limit, a traffic spike (slow_down) or your own spend cap, and the cap won't clear by waiting. A 503 server_is_overloaded means the model is busy. The SDK retries 429s and 5xx responses twice by default, so if you write your own loop, switch those retries off to avoid stacking them:
import OpenAI from "openai";
// Retries live in withRetry below, so the SDK's own are off.
const client = new OpenAI({ maxRetries: 0 });
const TRANSIENT = new Set(["rate_limit_exceeded", "slow_down", "server_is_overloaded"]);
async function withRetry<T>(call: () => Promise<T>, attempt = 0): Promise<T> {
try {
return await call();
} catch (error) {
if (!(error instanceof OpenAI.APIError) || attempt >= 4) throw error;
// Anything else, a spend cap included, fails fast: waiting won't fix it.
if (!TRANSIENT.has(String(error.code))) throw error;
await new Promise((r) => setTimeout(r, 2 ** attempt * 500));
return withRetry(call, attempt + 1);
}
}
Set key expiry and spend caps. Both arrived this quarter. If a key leaks through a front-end bundle, an expiry date and a hard cap on its project limit what an attacker can spend.
Watch your cache hit rate. The August prompt caching dashboard tracks it, and the September Prompt Cache Diagnostics tell you why an individual request missed. A timestamp or request id near the top of a system prompt is a common reason for misses.
DevDay
OpenAI holds DevDay 2026 on September 29 at Fort Mason in San Francisco, with the keynote streamed live. The Agents API beta from September 10 is the thread I'll follow there.
Sources
- OpenAI API changelog, July to September 2026
- OpenAI DevDay 2026, OpenAI
- OpenAI announces rollout of GPT-6 Astra model, CNBC, September 3, 2026
Volodymyr Chornous

