“Selected model is at capacity” is an error Codex classifies as server overload. We verified how the message maps to OpenAI’s public code and found the same full message with serverOverloaded in this computer’s execution history. Treat it separately from exhausted usage allowances or a PC malfunction. However, neither public information nor the computer’s records established what caused the overload.

Selected model is at capacity. Please try a different model.

What to check when work stops

01 Save your request and wait

Keep your prompt and the last completion report. Pause before retrying instead of repeatedly sending it.

02 Check incidents and usage

Check OpenAI Status and your usage separately. Incidents can occur even when you have allowance left.

03 Consider an alternative if urgent

Choose another available model yourself. Switching may not help during an incident affecting multiple models.

The order recommended in this article. Switching models is an option, not a guaranteed fix.

What is happening? What official incident records show

The message says the selected model has reached capacity and asks you to try another model. There is no basis for interpreting “capacity” here as your PC’s memory or disk space. OpenAI’s official status page records service incidents displaying this message.

June 16, 2026: Codex capacity errors

OpenAI Status lists Desktop, Web, API, CLI and the VS Code extension as affected. It reports that mitigations were applied and the incident was resolved (official incident record).

July 9, 2026: the same message across multiple models

The OpenAI Status report contains the same full error message shown at the beginning of this article. It explicitly says multiple models were affected and subsequently reports recovery (official incident record).

These two records establish that the same message can appear during service incidents and that switching models does not always help. The published incident count cannot establish how frequently the error occurs, nor does it show that your error had the same cause as a previous incident. Dates follow the official pages; we have not converted their times to Japan Standard Time.

Neither record publishes a root cause such as a particular hardware shortage or an account-switching bug. Claims such as “there are not enough GPUs” or “Pro authentication is broken” go beyond the information verified. We checked the original official sources on October 1, 2026, and distinguish established facts from our recommendations.

How it differs from usage limits, 401 and thread not found

These problems can all look like stalled work, but they require different checks. For a capacity error, start by checking both the full message on screen and your usage status. Several types of problem can occur on the same account.

Message or situationMain checksHow to interpret it
Selected model is at capacityOfficial incident reports, time of occurrence and selected modelThe message alone does not mean your usage allowance is exhausted.
A usage-limit or reset-wait noticeRemaining allowance and reset time on the usage screenCheck the remaining allowance before deciding whether to wait or how to continue.
401 Unauthorized
Incorrect API key provided
Sign-in method and active accountAn authentication problem requires a different response from a capacity error.
thread not foundWhether the affected chat loads or resumesThis means the chat was not found; do not equate it with model congestion.

Check remaining allowance on the dedicated usage screen

OpenAI’s pricing and usage guide directs users to the usage dashboard to check current limits. During a Codex CLI session, you can also use /status. Checking immediately after the error makes it easier to compare the error with your remaining allowance at that moment.

The official guide also explains that work already in progress can continue under fair-use constraints even if a usage limit is reached during processing. Therefore, an error followed by a completion report does not, by itself, establish whether a limit was reached. Conversely, a service incident can occur with plenty of allowance remaining. See our ChatGPT Pro plan comparison for included allowances and additional credits.

For a 401 error, first check how you are signed in

The Codex authentication guide distinguishes signing in with ChatGPT from signing in with an API key. In the desktop app, the profile menu shows the active account or API-key status. In the CLI, use codex login status.

Even if “Incorrect API key” appears while you are signed in with ChatGPT, the screen alone does not establish that you configured an incorrect key. Identifying which credentials were rejected requires further investigation. A capacity error is not a reason to immediately sign out or delete authentication files. Using an API key incurs standard API charges separately from your ChatGPT subscription.

If the message is thread not found, see our investigation and troubleshooting guide for Codex’s “thread not found” error. Previous authentication or chat errors on the same PC do not establish a causal connection to the current capacity error.

Distinguish API HTTP 503 errors from the app’s message

The OpenAI API error-code guide explains HTTP 503, service_unavailable_error and server_is_overloaded as indicators of temporary model overload. HTTP 429 includes errors concerning request frequency or usage limits, while HTTP 401 concerns authentication.

Do not infer an HTTP code from the app’s wording
The API documentation helps distinguish error types, but it does not establish that Codex’s “at capacity” message always means HTTP 503. Our investigation of this computer did not obtain the HTTP code for the affected requests either.

Steps to resume your work

The following recommendations draw on official incident records, usage guides and general troubleshooting documentation. They are not published as a guaranteed fix specifically for this error.

1. Save your prompt and the last completed work

If your unsent prompt remains visible, copy it, and preserve the last completion report or the state of modified files. Before closing or replacing the chat, record what you requested and how much was finished. This makes resuming easier.

An error message alone does not prove that no files were edited or commands executed. For development using Git, inspect the diff or run git status before requesting the same change again. For work involving publication, sending messages or purchases, do not repeat the instruction without first checking its outcome.

2. Wait briefly and check the official status page

Open OpenAI Status and check for incidents affecting Codex or model selection. Compare the time of your error with the incident’s time window, rather than looking only at its current status. A past incident with the same name does not mean that incident is still happening.

For API overload, official guidance says to wait at least as long as Retry-After specifies, if present, or increase the interval between retries if it is absent. We could not verify a fixed number of seconds app users must wait when no delay is displayed. Our recommendation is to wait briefly before retrying instead of repeatedly sending requests in quick succession. During a listed incident, consult the official recovery updates.

The status page presents aggregate information. It may not fully match the situation of a particular account or model. An incident not being listed does not establish that your PC is broken.

3. Check usage and consider another model if the work is urgent

Check your remaining allowance and reset time. If a usage limit is explicitly reported, follow that notice when deciding what to do. Do not treat buying additional credits or a paid reset as a recovery method when a capacity error is the only message. Restoring your own usage allowance and whether the model can accept requests are separate issues.

Prioritize quality and continuity

Wait for recovery if you want to resume with the same model. Use the time to review requirements, changes and outstanding checks.

Prioritize making progress now

Choose another available model and resume with a small task. An incident affecting multiple models can still stop work after switching.

The official model-selection guide explains that the desktop app’s model and reasoning-effort controls are below the prompt field. In the interactive CLI, use /model. Available models vary by account, client and other factors, so do not assume a model absent from your interface is available.

Changing models can change the nature of responses and usage consumption. There is no basis for always choosing the highest-tier model to avoid capacity errors. When handing over complex implementation or design work, review the new output through diffs and tests. Our GPT Sol generation comparison and selection guide also offers context.

4. Investigate app responsiveness separately if the app itself is frozen

Distinguish a capacity error from unresponsive input, a terminal or the app interface. For chats that appear stuck, official troubleshooting guidance recommends checking for pending approvals, checking the terminal with a basic command, and trying a small request in a new chat.

If the terminal remains stuck, the guidance recommends waiting for active chats to finish before restarting the app. This addresses a general unresponsive state; it does not say that restarting resolves a shortage of server-side capacity. First check the state of any other chats that are still running.

If you resume in a new chat, briefly provide the goal, working folder, completed changes and unfinished tasks. Do not send only “continue” on the assumption that the entire previous conversation is available. A new chat uses the same service, so it is not guaranteed to avoid the capacity error.

What to record if it happens again

If the error recurs, preserve evidence that supports later investigation. Recording it before repeatedly changing settings makes the conditions of the failure clearer.

Checklist for support requests and recurring errors

  • Date, time and time zone of the error
  • Where you used Codex—app, CLI, IDE—and its version
  • Selected model, reasoning effort and speed setting
  • Full error message and the action immediately before it
  • Remaining allowance, reset time and official incident information at that moment
  • What happened after waiting and retrying, and whether another model showed it too

Depending on the available records, you may not be able to confirm that the model selected in settings was the model actually used for the failed request. If you recorded only the visible setting, describe only that evidence. Likewise, recovery immediately after switching models does not prove the switch fixed the problem: the service may have recovered at the same time.

The official guide explains how to submit feedback by entering / in the message field. When submitting from an existing chat, you can choose whether to share the conversation. Remove API keys, email addresses, private conversations and internal company information from logs or screenshots you send. You do not need to publish key values or authentication files themselves.

A report giving the date, model, exact message and displayed remaining allowance is more useful than simply saying “it keeps breaking.” An HTTP code or request ID, if obtained, can help support investigate, but do not fill gaps with guesses.

What we found on this computer—and what remains unknown

After a user reported repeated occurrences with screenshots, AI Arte had Codex on this computer perform a read-only investigation of usage and local logs on October 1, 2026. The Windows app package version was 26.928.3736.0. We did not verify that the same version was installed when the errors occurred.

Verified

Execution history stored the same full message with serverOverloaded. The public code also classifies it as server overload.

Not established

What caused the overload, the allowance remaining then, the HTTP code and the actual model used by the failed request.

The initial search did not find the error in 32,737 records in the regular log database, 15 desktop-log files or 143 updated session-record files. After a reader questioned the findings, we checked other storage locations and found failure records in a separate execution-history database. The initial search scope was insufficient.

At the time of the follow-up investigation, 35 of the 6,781 turns in the execution history were failures. Across the entire recorded period, 13 contained the same full error message and codexErrorInfo: serverOverloaded. Of these, 12 occurred across five chats between September 30, 2026, at 22:10:01 and October 1 at 00:24:55 (JST). The chat where the user attached error screenshots also contained four failures on September 30, at 22:15:06, 22:57:58, 23:12:31 and 23:24:38, with matching messages and classifications. These times are recorded completion times of failed turns, not screenshot timestamps. They describe this computer’s records, not an error rate for all users.

The Codex CLI bundled with the app was version 0.159.2. In the error definitions under the matching public tag, the full message corresponds to ServerOverloaded, classified separately from exhausted usage allowances. The stream-handling code converts server_is_overloaded into an overload error, and the display-conversion code connects it to the full message at the beginning of this article. There is also a path that converts HTTP 503 responses with the same error code, but errors can arrive within a stream too. The displayed message therefore does not establish that the response was HTTP 503.

What we established is that the failures were classified and stored as server overload. The records do not show whether this involved an actual GPU shortage, request routing or capacity-control issues. Additional details for the four failures were empty, and no HTTP code was recorded. During the initial investigation, 31% of the weekly allowance remained and ordinary usage was allowed. This was not the allowance at the time of the errors, and the actual model used by those failed requests was not established either.

For authentication errors, the official root-cause report for a separate September 25 incident explains that mistakenly detecting and invalidating internal service credentials caused 401 and 502 errors in Codex signed in through ChatGPT. This does not establish the cause of the overload errors examined here. Do not conflate an officially explained internal authentication incident with the present overload errors.

This investigation did not change models or subscriptions, use paid resets or delete credentials. We also did not deliberately send rapid repeated requests to reproduce the problem, so we have no experimental comparison of which actions restore service. We distinguish officially verified incident examples from observations on this one computer.

Summary: identify the error before choosing a response

If “Selected model is at capacity” appears, save your prompt, wait briefly, and check incident reports and usage. For urgent work, you can consider another available model, but some incidents affect multiple models. The capacity error alone provides no basis for spending more, buying a paid reset or deleting credentials.

Check authentication for a 401, the chat’s state for thread not found, and usage status for a limit notice. Recording the time, full message, model and remaining allowance when the error recurs helps the next investigation more than changing settings without knowing the cause.

Frequently asked questions

Can this happen with a Pro subscription?

The user in this article reported the same message with a Pro subscription. One report cannot establish incident rates by plan, however. The official incident records do not say that Pro is exempt. Check remaining subscription allowance separately from whether the model can accept work at that moment.

Can additional credits or a paid reset fix it?

We found no official evidence that these actions resolve this capacity error. Additional credits and similar mechanisms concern your own usage allowance. First check whether you have actually reached a limit, and do not confuse restoring allowance with recovery from a capacity error.

Why does it still happen after switching models?

Official incident records describe the same message affecting multiple models. Its appearance on another model does not, by itself, prove a PC or account malfunction. Check incident reports and your remaining allowance, and record retry outcomes. Public documentation does not establish a model guaranteed to avoid the error.

Should I restart the app or open a new chat?

We could not establish that either is necessary. These actions can help investigate an unresponsive app or terminal, but are not guaranteed to resolve a service-side capacity problem. Check other ongoing work, and preserve your prompt, changes and unfinished tasks before deciding.