If you keep sending Codex “continue,” one option is to use /goal to retain completion criteria. It suits work such as reproducing a bug, fixing it, and choosing the next step based on test results. It is more than an instruction to run for a long time: it keeps the conditions for finishing in the same chat.
This guide explains when to use it, how desktop and CLI controls differ, and what to check when work stops. It also explains why a goal’s token budget should not be confused with your remaining plan allowance or a cap on your bill.
Decide these three things first
Which bug to fix or what to build
What may change, what must keep working, and which actions are allowed
Which tests or measurements will establish that the work is complete
Specifications checked: October 7, 2026. This explanation is based on official OpenAI documentation. We did not run Goal mode for this article or measure its duration, usage, or effect on outcomes.
1. What /goal does: ordinary requests, /plan, and dot
/goal sets an ongoing goal in a Codex chat. If a test fails along the way, Codex can choose its next action against the original completion criteria. OpenAI lists bug investigations, performance improvements, migrations, and research as examples where the next step depends on what the investigation reveals.
Repeat work and verification against the goal
- Work: inspect code or documents, make changes, and take measurements
- Check: assess whether evidence shows that the goal has been met
- Choose: finish if achieved, take the next step if unfinished, or report why progress is blocked
Unfinished work continues when the goal is active, within budget, and the conditions for automatic continuation are met.
The official explanation describes a goal as state saved in the current chat. It is not global memory that automatically applies to other chats or a rule for the entire repository. The necessary code, tests, and documents still need to be accessible from that chat. Source: OpenAI Cookbook: Using Goals in Codex.
| Method | Suitable requests | When to use it |
|---|---|---|
| Ordinary request | A single change, explanation, or short check | When you want a result from one request |
/plan | Clarifying what to build or how much to change | When the goal is vague and you need to settle requirements and verification criteria |
/goal | Work that repeats investigation, changes, and verification until completion criteria are met | When the endpoint is clear but the steps to reach it are not yet known |
| dot | Ongoing assistance, delegation to development agents, and coordination | When you also want to delegate coordination across multiple jobs |
Making a plan alone does not start Goal mode’s automatic continuation. Nor does /goal guarantee an independent review by another model. For the distinction from dot, see how to use dot, its pricing, and delegation to Codex. For how to supply information, see context engineering.
2. Starting on desktop, CLI, and IDE
Although they share /goal, the controls after starting differ by interface. The official guide to long-running work explains how to start on desktop, CLI, and the IDE extension. Its web section describes giving ChatGPT Work your outcome, constraints, and evaluation criteria; that does not establish that the web interface has the same command and controls.
Desktop
- Open the relevant project and chat
- Use
/goalin the composer and state the completion criteria - Check the goal progress row above the composer
Use that progress row to pause, resume, edit, or clear the goal.
Codex CLI
- Open an interactive session in the relevant working directory
- Enter
/goalfollowed by your objective - Send status questions or instructions for changes in the same session
Use the CLI commands below to check the status or pause the goal.
IDE extension
- Open the relevant workspace
- Use
/goalin the extension’s chat - Provide additional information in the same chat
Keep the workspace accessible while work is in progress.
For the CLI, the OpenAI Cookbook lists Codex 0.128.0 or later as supported. The current configuration reference describes features.goals as stable and enabled by default. Do not assume that you need to add older instructions for enabling an experimental feature to your configuration. If the feature is missing, check your interface, version, and the current official guidance.
Sources: Long-running work, Desktop slash commands, and Configuration reference. The instructions in this article explain what the controls do; they are not a list of button labels verified in each version of the app.
3. Defining completion: making vague goals testable
“Make it high quality” or “keep going until it’s done” makes it difficult to decide what counts as finished. Completion criteria should describe results you can verify. Specify screen widths, behavior after saving, test results, or the documents to compare.
“Improve this to-do app and keep going until it’s finished.”
Appearance, features in scope, checks, and stopping conditions are undefined.
“Restore one deleted item to its original position and completion state. Check persistence after restoration and prevent duplicate restoration, then report the results.”
Match the required behavior to the evidence that will establish success.
In the CLI, the goal text must be nonempty and no longer than 4,000 characters. Rather than fitting an entire lengthy specification into it, point to a specification file and keep the outcome, key constraints, and verification criteria in the goal. Providing a file does not automatically carry over the history of another chat. Source: Codex CLI slash commands.
If the specification is undecided, you can first ask: “Do not implement yet. Clarify the requirements and draft a /goal.” After working through it with /plan, read the proposed goal yourself, check the constraints and scope, and then start.
4. Example prompts for bugs, interfaces, and research
The following prompts were written for this article. They are not examples of successful executions we performed. Check first whether tests and a browser are available, and include a requirement to report unavailable checks as not performed.
Bug fixes: distinguish the reproduced issue from regression checks
/goal Fix “undo deletion” in this to-do app so that the most recently deleted item can be restored to its original position and completion state.
Preserve the existing behavior for adding items, toggling completion, deleting, and saving.
Reproduce the issue first. After fixing it, check the restored position, completion state, prevention of duplicate restoration, consecutive deletions, and persistence after restoration.
Change only related files and tests. Do not add dependencies, publish externally, push, or make purchases.
If you cannot run a check, report why and what environment is needed. In the completion report, include changed files, commands run and their results, and checks not performed.
“Pass the tests” alone can make a change that removes functionality appear to meet the goal. Including the behavior to preserve and the permitted scope of changes provides criteria for choosing how to fix it.
Interface improvements: specify display and interaction requirements
/goal Make this screen fit widths of 390px and 1280px without horizontal overflow, with completion and delete buttons that remain usable for long task names.
Do not change the existing data structure or storage format.
Check both widths in an available real browser. Test adding items, toggling completion, deleting, persistence after reloading, and keyboard interaction.
If browser testing is unavailable, do not count image or code inspection as a pass for actual interaction. Report those checks as not performed.
Do not publish, push, make purchases, or change device settings.
The widths are test conditions for this request, not screen widths guaranteed by the product. Also distinguish results from interacting with a real browser from results based only on inspecting code or images.
Research: retain evidence rather than filling unknown fields
/goal Compare the specified two services’ data retention, use for training, and deletion conditions using publicly available official documents.
Match each claim to a source URL and the conditions in the text you checked, and produce a comparison table and explanation.
For items where you find no explanation after searching, identify the documents searched and the missing information. Do not fill the table with guesses.
Do not log in, change settings, connect apps, upload, or make purchases.
At the end, submit confirmed specifications, interpretations, and items you researched but could not establish as separate categories.
Deciding to “continue indefinitely until every answer is found” removes the endpoint from research. Even when information is not public, you can define a report of the sources searched and remaining questions as the outcome.
5. Pausing, resuming, editing, and clearing a goal
On desktop, use the goal progress row above the composer. The CLI documentation lists the following commands. Do not treat this CLI table as a list of desktop button actions.
| What to enter in the CLI | Purpose | What to check |
|---|---|---|
/goal | Show the current goal | Does it reflect the completion criteria for this job? |
/goal edit | Edit the goal | Do the new criteria require fresh verification? |
/goal pause | Pause the active goal | Check the state after pausing |
/goal resume | Resume a paused goal | Have the working environment or constraints changed? |
/goal clear | Clear the current goal | Avoid carrying old completion criteria into the next job |
You can provide additional information or constraints in the same chat while work is underway. Explicitly state changes in your decisions, such as “do not publish yet” or “cancel the changes to this feature.” After editing the goal, check whether earlier test results alone establish that the new completion criteria have been met.
If code or saved data needs to be restored, check the diff and saved state separately. The CLI command /stop stops background terminals; it is not an alias for /goal pause.
Sources for the controls: CLI goal commands and Desktop goal controls.
6. Token budgets, usage allowances, and pricing
For ongoing work, distinguish when to stop work on the goal from how much of your plan allowance it consumes. Even with goal budget remaining, account usage limits or execution environment problems can prevent work from continuing.
| Item | What it covers | What it should not be equated with |
|---|---|---|
| Goal token budget | Managing the budget and usage for continuing work on that goal | Remaining plan allowance or a strict billing cap |
| Plan usage allowance and credits | A shared allowance and additional balance for processing in Work, Codex, and related services | A budget dedicated to one goal |
| Context capacity | The amount of context the model can process | Total tokens across the ongoing work or the monthly price |
Does /goal incur an extra charge?
In the official pricing material we checked, we found no separate per-start charge for /goal. However, repeated model processing consumes ordinary Codex usage. Enabling Goal mode does not make cycles of testing, fixing, and checking free and unlimited.
OpenAI explains that Work and Codex share a usage allowance, and that consumption varies with the model, task, and other factors. Continuing with additional credits after using the plan allowance must also be distinguished from separate billing with an API key. Source: Work and Codex pricing and usage limits. For plan selection and resets, see ChatGPT Pro pricing and usage allowance comparison.
What do the budget figures establish?
The official developer documentation for App Server lists the goal’s tokenBudget, usage field tokensUsed, and time measurement field timeUsedSeconds. This establishes a mechanism for recording budget and progress in internal goal state. It is not an instruction to enter these RPC field names as user-facing CLI flags.
- The budget-setting syntax for ordinary users and how to enter it in each interface
- The detailed calculation of goal counters, including cached input and delegated work
- Overshoot at the budget boundary and its exact relationship to the final bill
We therefore do not provide unverified commands or guarantee that setting a budget will keep your bill below a specific amount.
The same documentation explains that replacing a goal with a new one resets goal usage measurements. This does not say that it restores your remaining plan allowance. A time measurement field is also not evidence that you can guarantee a stop after two hours. Source: App Server goal management.
7. What to check when work stops
An active goal does not automatically overcome every interruption. First check the current goal and last result, then investigate in this order.
Is it complete, paused, cleared, or at its budget limit? If complete, check the evidence that the completion criteria were met.
Is it waiting for approval or required information? A pending follow-up message may need to be processed first.
Check the goal budget limit, plan usage limit, and selected model errors separately.
Are the required files, dependencies, testing tools, and connections available? For work on a local PC, also check that the PC is running.
According to the Cookbook, automatic continuation occurs when the chat is idle, the goal is active and within budget, and no other processing or user input is pending. Planning-only work does not trigger continuation, and interruptions pause the goal. If a continuation turn makes no tool calls, the next automatic continuation is suppressed to avoid unproductive repetition.
When the budget is reached, the official design stops substantive work and reports progress, blockers, and next steps. Exhausting the budget and achieving the goal are different things. Before resuming, check the remaining work and expected costs.
Starting /goal does not expand permissions or connected resources. The official documentation says it follows the existing sandbox and approval policy. It does not automatically move local work to the cloud. If you expect to lose connectivity, the guidance is to pause and resume when the environment is available. Source: Permissions and continuation conditions for long-running work.
If a model-side error such as “Selected model is at capacity” appears, investigate according to that message. Investigating and handling Codex’s at capacity error explains it separately from usage limits.
8. Evidence to check in a completion report
A reply saying “done” does not establish whether the requested result was verified. Look for evidence that matches the original goal. Even if tests passed, interaction checks that were not performed should not be treated as having passed.
| Completion criterion | Evidence to receive | Example of an inadequate report |
|---|---|---|
| The bug is fixed | Reproduction conditions, the diff, and results under the same conditions after the fix | Only suspicious code was changed, without reproducing the issue |
| Existing behavior is preserved | Commands and results for relevant regression tests | Only the new feature was checked; existing saving behavior was not |
| The interface and interactions work | Actual interactions at the specified widths, including input and reloading | Images alone were used to mark saving and button interactions as passed |
| Evidence-based research is complete | Claims matched to source text, conditions, and remaining questions | A list of links without an explanation of what was confirmed |
If evidence is missing, give a specific follow-up in the same chat: “test persistence after reloading” or “show the source text and conditions for this claim.” If you add completion criteria, update the goal too. If you want another agent to check the work, explicitly request that separately, and distinguish a review based only on the agent’s report from checks that actually rerun the work.
You can also ask for requirements, implementation, and verification in ordinary Codex development. Goal mode’s value is retaining completion criteria across multiple steps and using them to choose what to do next. For differences between products and execution modes, see Claude Code and Codex compared.
9. Before you start
- Include the outcome, constraints, and verification in the goal text
- Make the required files and testing environment available to the chat doing the work
- Manage the goal through the progress row on desktop or goal commands in the CLI
- Check the goal budget and plan usage allowance separately
- Compare the completion report against tests, diffs, actual interactions, and sources
If one change or explanation is enough, an ordinary request is suitable. Consider /goal for repeated work toward the same completion criteria where the next step depends on investigation results. Choose based on verifiable outcomes rather than how long it runs.
10. Frequently asked questions
Can I stop sending “continue” every time?
When the goal is active and the conditions for automatic continuation are met, work can proceed to the next step after a turn. It may still stop for approval, a budget limit, or a blocker. The feature does not eliminate the need for human decisions.
Is /goal cloud-only? Will it run after I close my PC?
It is not cloud-only. It is documented for desktop, Codex CLI, and the IDE extension. The environment needed for continuation depends on where the work runs. Setting a goal alone does not establish that work will continue on a local PC that is powered off.
Does a token budget guarantee a cap on charges?
It cannot be guaranteed as a strict billing cap. The goal budget is separate from plan allowances, additional credits, and API billing. In the official material we checked, we could not find an explanation of the exact relationship between goal counters and the billed amount.
Is it the same as using dot?
/goal manages a goal and continuation within the same Codex chat. dot also handles ongoing assistance, delegation to other tasks, and coordination. A direct request to Codex can be sufficient for a small development job. Choose based on your purpose and whom you want to manage progress.