Choose a GPT-6 model based on the reasoning your task requires, speed, and cost.
-
Highest intelligence
For the most demanding reasoning, coding, and professional work.
-
Balanced speed, cost, and intelligence
Near-Astra performance for complex work at a lower cost.
-
Fastest and most cost-effective
Strong performance for focused, high-volume tasks.
To get started, set model in a Responses API request. If you already use gpt-6-sol, review the migration guidance before switching to GPT-6.1 Sol.
GPT-6 Astra
GPT-6 Astra is our most intelligent model yet, with state-of-the-art performance in computer use, browsing, software engineering, science, and professional work. It can carry out multi-step workflows across code, browsers, and professional software. In several evaluations, Astra achieved stronger results using substantially fewer output tokens. Its estimated API cost per task was lower than earlier models despite its higher per-token pricing.
GPT-6 Astra is also our most aligned model yet. It excels at exercising care, respecting task boundaries, and communicating transparently. When instructions leave room for interpretation, it uses the context it has to fill in routine gaps and asks focused questions when the answer could change the outcome. It incorporates new requirements, changes course when asked, and answers side questions without losing track of the broader task.
All GPT-6 Astra users also have access to Fast mode and the new Ultrafast mode for our fastest API speeds.
GPT-6.1 Sol
Use GPT-6.1 Sol for complex coding, computer use, and professional work when you want near-Astra performance at a lower cost. Compare it with Astra on your tasks to assess the tradeoff between quality and cost.
Set reasoning.effort to low, medium (default), high, xhigh, or max.
Use the Responses API for tool calling. Chat Completions supports requests
without tools. The none and minimal reasoning efforts are not supported.
See the model page for specifications, pricing, and availability, or model selection for guidance on choosing a model.
What’s new
- Async tool calling: GPT-6 can continue reasoning, call other tools, or answer independent parts of a request while your application runs a tool. Set
async: trueon a function or custom tool and return its result when ready using the originalcall_id. Your application still executes the tool and manages pending work. See Async tool calling for basic usage and a developer-defined wait-tool pattern. - Mid-turn steering: Send additional user instructions while GPT-6 is working, such as a correction or a change in requirements. Over a WebSocket connection, the Responses API preserves completed work and includes the update in a continuation. See Mid-turn steering for the event flow and tool-result handling.
- Change reasoning mid-conversation while preserving cache: Add a
configuration_updateinput item to increase reasoning effort for difficult work or reduce it for routine follow-ups without rewriting the original prompt prefix. The updated reasoning effort applies until anotherconfiguration_updateinput item overrides it. See Change reasoning mid-conversation for examples and compatibility. - Misalignment monitoring: As part of our strengthened safeguards for GPT-6 Astra, our systems asynchronously monitor for misalignment and trigger alerts when necessary. See Misalignment monitoring for more information.
GPT-6 also supports the existing API capabilities available with GPT-5.6, including computer use, Structured Outputs, streaming, Programmatic Tool Calling, multi-agent orchestration, prompt caching, persisted reasoning, compaction, and pro mode.
Limitations
- GPT-6 Astra and GPT-6.1 Sol do not support the
nonereasoning effort; GPT-6 Sol and GPT-6 Luna do. - Fast mode is not available with EU data residency for GPT-6 Astra, GPT-6 Sol, or GPT-6 Luna. Ultrafast mode supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints. See data residency eligibility.
Prompting best practices
Use the following prompts as a starting point across the GPT-6 model family. They address behavior observed with GPT-6 Astra; evaluate them with your chosen model and workload.
GPT-6 Astra behavior
- Initiative and follow-through: The model is designed to be a more effective collaborator and is thus more likely to ask the user a question when additional input could materially change the result. This can cause it to stop when the user may expect it to make reasonable assumptions and persist.
- Instruction following: GPT-6 Astra is stronger at general instruction following than our previous models, giving you greater control over its behavior. It can be more sensitive to instructions contained in skills and other files, such as
AGENTS.md. We strongly recommend auditing skills and other files accessible to your model for instructions that could influence its behavior. - Personality and writing style: The model tends toward detailed, formatted responses and may use recurring phrases across sessions. Specify the writing style and structure your application needs.
- Subagent delegation: The model may delegate less often than desired for your workflow. Specify when and how much it should use subagents for parallel work.
- Testing and verification: For coding tasks, the model tends to be thorough in testing before considering a task complete. For smaller tasks, this can result in broader tests than the task requires.
Initiative and follow-through
GPT-6 Astra is generally better than GPT-5.6 Sol and earlier models at staying coherent during long tasks. It is also more likely to ask for clarification where earlier models would make assumptions.
To encourage more autonomous work, start with this prompt:
You should infer the user's intent and task scope from the instructions and prior conversation context. Your job is to bias towards action and carry the user's intended task to completion.
When the user expresses intent to perform new work or fix an existing issue, persist until the user's intended goal is complete. Progress autonomously towards the user's goal (e.g. creating isolated worktrees / checkouts if needed, resolving merge conflicts, read-only actions, creating draft PRs etc.) unless they are clearly destructive or irreversible.
When the user’s intent is unclear, the model is more likely to ask the user for clarification to proceed. Prompt the model to follow through if the user’s prompt implies authorization:
When the user's prompt indicates a request for action, such as "can you...", "I want to...", "help me..." and similar expressions, treat these as instructions to do the work and take action. Do not stop at acknowledging capability (e.g. "Yes…"), proposing a plan, or offering to continue. Do not settle for a partial or "helpful enough" solution that does not fully satisfy the user's task to save time, effort or tokens. If a task requires sustained work, complete all the necessary work until the intended outcome is fulfilled.
Prompt the model to ask for approval only after preparing a concrete, reviewable result. This avoids blocking the task before the model has done the work it can, and often leads to quicker task completion.
Before asking the user clarifying questions, you should complete the work that is already authorized from context and necessary to make the proposed action concrete and reviewable. The user should be approving a concrete, reviewable result. For example, before deploying a change, writing to an external application, merging a PR or publishing a site, do all the required work first so that user approval is the final step. You don't need user permission for reversible tasks, read-only actions, reviews or fixes, or anything for which authorization is provided earlier in the session or strongly implied from the task instruction.
Do not introduce unsolicited warnings, disclaimers, approval flows, or safety/compliance checklists due to hypothetical risk.
The model also likes to ask non-blocking questions as it’s working by default, so adjust these prompts to match the level of autonomy your application needs.
Instruction following
GPT-6 Astra is better able to follow longer instructions, but can also be more sensitive to information in context. For example, unclear or conflicting guidance in a skill file may cause the model to pause and block work early. Make the priority of user instructions and skills explicit.
The user's instructions take precedence over guidelines provided in a skill. If explicit user instructions conflict with a skill's instructions, prioritize the user's instructions.
Asking the model to identify the skill and instruction that caused it to pause or change direction can also be effective in providing transparency into model behavior.
If a skill causes you to ask for permission or confirmation, pause, leave requested work unfinished, or diverge from the user's intent, name and link to the exact SKILL.md file you read, quote the relevant instruction, and briefly explain how it applies. Distinguish explicit skill requirements from your interpretation of guidelines.
Use this prompt to find silent and conflicting guidance when your application loads many skills and instruction files such as AGENTS.md.
Personality and writing style
GPT-6 Astra tends to use lists, tables and Markdown to make responses scannable. If your application needs prose with less formatting, specify that preference.
Default to using clear, concise paragraphs, each developing one main idea. Use lists only when the information is genuinely parallel, sequential, or easier to compare, and avoid nested lists unless the hierarchy cannot be expressed clearly in prose. Use plain, simple language: familiar words, concrete examples, and precise verbs. Prefer active voice and direct statements.
Make sure to state the main point clearly and early, then develop it with the explanation and detail the reader needs. Let each sentence build on what came before. Develop the points that matter and provide enough support to be useful.
For technical communication, the following prompt helps strike a balance between using clear, coherent language while remaining domain appropriate:
Use plain language over jargon, and reference technical details only to the degree that it helps illustrate an idea or your work to the user. Communicate complex concepts in a clear and cohesive manner, and calibrate your writing to the level of background knowledge assumed from the user's prompt and context.
To reduce jargon and stock phrases in writing, start with this prompt:
Avoid using slop words or phrases like "Bottom Line:" in conclusions, "delve," "foster," "leverage," "it's worth noting," "importantly," "Question? Answer." or "This isn't about X. It's about Y.", "genuinely" or hyphenated compound descriptions and adjectives. Do not use concluding summary statements such as "In short:..", "The simplest mental model is:...".
State the intended action directly. Avoid adding what you won't do, what will remain unchanged, or how you'll separate or categorize results. Do not use contrastive framing such as "X, not Y" that introduces an unprompted alternative that the user didn't ask about. Avoid invented compound labels like "exact-head checks" and "editorial-row layouts", vague qualifiers, and canned transitions; use plain verbs and prepositions to state the actual relationship directly.
Subagent delegation
GPT-6 Astra is trained to be able to divide and delegate work to subagents that work in parallel. If you are implementing a multi-agent system in your harness, use the following prompt to tune how much GPT-6 Astra should delegate work:
If at any point you can parallelize work by delegating tasks to another agent (no matter if you are the root or subagent), you should do so using collaboration tools if it could save time or improve quality.
Messages between agents may contain grammar or spacing errors. Use this prompt to make inter-agent messages easier to read:
Messages that you send to other agents and your final answer may be read by a human, so ensure they are legible. Always put proper spaces between words and/or numbers.
The model tends to respond well to prompting for how and when it should delegate work to subagents, so tune this behavior to fit with your harness and multi-agent implementation.
Testing and verification
For coding tasks, calibrate how much testing and verification a change requires. This can help avoid unnecessary tests or repeated checks for small changes.
Do not write tests for reversible, low-impact changes that mirror the implementation. If you do choose to verify your work with tests, make sure that the tests are meaningful and necessary to verify implementation.
Run tests appropriate to the change and complete required checks. Once those pass, broaden or repeat testing only when new changes, failures, or unresolved concerns justify it; otherwise, continue toward completing the task.
Migration quickstart
Migrate with Codex
Codex can apply the recommended changes in this guide with the OpenAI Docs skill.
$openai-docs migrate this project to the GPT-6 model family
To use this skill in other coding agents, download it from the Codex repository.
Update API and model parameters
Set model to gpt-6-astra, gpt-6.1-sol, or gpt-6-luna, then check the following:
- Reasoning effort: Preserve your current effective reasoning effort where supported. GPT-6 Astra and GPT-6.1 Sol do not support
none; uselowinstead. GPT-6 Sol and GPT-6 Luna supportnone. If your existing request usesminimal, start withlowand compare results on representative tasks. Usereasoning.effortin Responses orreasoning_effortin Chat Completions. - Tool calling: Use the Responses API. GPT-6 Astra and GPT-6.1 Sol support Chat Completions, but tool calling requires Responses. GPT-6 Sol and GPT-6 Luna support function calling in Chat Completions only with
reasoning_effort: "none". Use Responses for reasoning with tools. - Unsupported parameters: When reasoning effort is not
none, removetemperature,top_p, andtop_logprobs. For Chat Completions, also removelogprobs. For Responses, removemessage.output_text.logprobsfrominclude. - Data residency: Fast mode is not available with EU data residency for GPT-6 Astra, GPT-6 Sol, or GPT-6 Luna. Ultrafast mode supports US data residency and global processing only. It does not support EU or other non-US regional processing endpoints. Fast mode for GPT-6 Astra does not include a latency SLA. See Fast mode compatibility.
- Changing reasoning effort: If your application changes effort between responses, use
configuration_updateitems in standard, single-agent requests. Keep request-levelreasoning.effortunchanged to preserve the prompt prefix for caching. Check the compatibility limits before adopting this feature. - Prompt caching: When migrating from GPT-5.5 or earlier, replace
prompt_cache_retentionwithprompt_cache_options.ttlset to"30m". Review the prompt caching changes, including cache boundaries and cache-write billing. - Unnecessary approval pauses: If you run into issues where the model keeps asking for approval before proceeding, use the initiative and follow-through guidance to prompt for more autonomous execution. See the rest of Prompting best practices for guidance on instruction following, writing style, subagent delegation, and testing.
Previous models
Using GPT-5.6
Learn best practices, features, and migration guidance for GPT-5.6 and the GPT-5.6 model family.
Introduction
GPT-5.6 sets a new quality and efficiency baseline for complex production workflows. GPT-5.6 is especially token-efficient and improves frontend aesthetics, including layout, visual hierarchy, and design judgment.
GPT-5.6 also introduces a new naming scheme. The gpt-5.6 alias routes requests to gpt-5.6-sol, the model for flagship capability. Use gpt-5.6-terra for strong performance at a lower price and gpt-5.6-luna for efficient, high-volume workloads.
When migrating from GPT-5.5 or GPT-5.4, start with your current GPT-5.5 or GPT-5.4 reasoning setting, then test the same setting and one level lower on representative tasks. GPT-5.6 can often maintain or improve quality with fewer tokens, but the best setting depends on your workload.
What’s new
- Programmatic Tool Calling: GPT-5.6 can write JavaScript to call eligible tools, pass results between calls, and process intermediate outputs in a hosted runtime. Use Programmatic Tool Calling for bounded, tool-heavy workflows that do not require fresh model judgment between each step. Programmatic Tool Calling is ZDR-compatible with no additional container costs.
- Multi-agent [beta]: Multi-agent lets a GPT-5.6 instance coordinate multiple subagents in parallel and synthesize their results. Similar to ultra mode in Codex, this can reduce wall-clock time and improve performance for complex tasks that divide cleanly into independent workstreams. Multi-agent is available as a beta feature in the Responses API as we iterate on developer feedback.
- Explicit prompt caching: GPT-5.6 lets you mark exactly which reusable prompt prefixes OpenAI caches. You can still use automatic caching in implicit mode. OpenAI bills cache writes at 1.25× the uncached input rate, while cache reads remain discounted. Learn how to configure prompt caching.
- Persisted reasoning: GPT-5.6 can reuse available reasoning items across turns to improve multi-turn quality and cache efficiency. Use
reasoning.contextto select the behavior. Learn how to preserve reasoning across calls. - Max reasoning effort: GPT-5.6 supports
maxreasoning effort for demanding tasks that need more exploration and verification. If you currently usexhigh, compare both settings on representative workloads. - Pro mode: GPT-5.6 can perform more model work to improve reliability on difficult tasks and return a single final answer. Enable it with
reasoning.mode: "pro"when quality matters more than latency and token usage. Learn how to use pro mode. - Token efficiency: GPT-5.6 reaches flagship-level performance with fewer output tokens.
- Frontend design: GPT-5.6 creates more polished and usable websites and applications, with stronger layout, visual hierarchy, and design judgment.
- Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work from context, so you often do not need to prescribe every step. Continue to provide domain context, hard constraints, approval boundaries, and success criteria. Tell the model when an important ambiguity should trigger a question.
- Original image detail: GPT-5.6 preserves image dimensions with
originalorautodetail, except that images larger than 65,535 pixels on either side are scaled down to fit that limit. The API rejects images that still exceed the 30,000-patch limit, rather than resizing them to fit it. Large images can use more input tokens and increase latency. Learn how to choose an image detail level.
Safeguards
When using GPT-5.6 models, users may encounter safeguards that block or refuse some requests due to real-time cyber and biology misuse classifiers that are run as model outputs are generated. Other requests may take longer because generation is paused for several seconds mid-stream while these classifiers synchronously review outputs. Safeguards may occasionally intervene on legitimate work, particularly in dual-use areas where defensive and offensive activity can initially look similar.
If your application serves individual end users, send a stable, privacy-preserving safety_identifier with each request. See Implement safety identifiers for guidance.
We are continuously evolving these safeguards so that they are robust and effective in holding up to adversarial pressure, while preserving access to legitimate work such as code review, vulnerability research, patch development, debugging, security education, and defensive testing.
Migration quickstart
Migrate with Codex
Codex can apply the recommended changes in this guide with the OpenAI Docs skill.
$openai-docs migrate this project to the GPT-5.6 model family
To use this skill in other coding agents, download it from the OpenAI skills repository.
Update API and model parameters
- Choose the target model for the workload. Use
gpt-5.6-solfor flagship capability,gpt-5.6-terrafor a balance of intelligence and cost, orgpt-5.6-lunafor efficient, high-volume workloads. Thegpt-5.6alias routes requests togpt-5.6-sol. - Use the Responses API for reasoning, tool-calling, and multi-turn workflows.
- Set
reasoning.effortintentionally. GPT-5.6 supportsnone,low,medium,high,xhigh, andmax.- If you are migrating from GPT-5.5 or GPT-5.4, preserve your current reasoning effort as the baseline, then compare one level lower.
- If you use
none, keep it as your latency baseline and also testlowwhen the workflow benefits from reasoning or tool use. - Use
mediumas a balanced starting point andlowfor latency-sensitive workloads. - Use
highorxhighwhen more reasoning produces a measured quality gain. - Reserve
maxfor the hardest quality-first workloads. Comparemaxandxhighto find the best quality, latency, and cost tradeoff for your use case.
- To use pro mode, keep your selected GPT-5.6 model and set
reasoning.modetoproin the Responses API; do not switch to a separate Pro model slug. Choosereasoning.effortindependently. If you omit it, GPT-5.6 defaults tomediumin both standard and pro modes. See reasoning mode for a request example and billing details. - Configure persisted reasoning based on how much prior reasoning is still relevant. GPT-5.6 models default to
all_turns; earlier models default tocurrent_turn.- Omit
reasoning.contextor set it toautoto useall_turns, the GPT-5.6 default. Check the response’sreasoning.contextfield to confirm the effective mode. - Set
reasoning.contexttoall_turnswhen the task’s goals, assumptions, and priorities stay stable across turns. - With
all_turns, continue withprevious_response_idto make reasoning from earlier responses available to the model. - When managing history manually, preserve and resend previous user inputs and every response output item. For
store: falseor Zero Data Retention, replay the encrypted reasoning items that the API returns by default. - Set
reasoning.contexttocurrent_turnwhen earlier reasoning is no longer relevant.
- Omit
- Review prompt caching. You do not need to change code to keep using implicit caching. Because GPT-5.6 cache writes cost 1.25× the uncached input rate, track
cached_tokensandcache_write_tokensto understand net cost. Use explicit breakpoints orprompt_cache_options.mode: "explicit"to avoid unnecessary writes, and replaceprompt_cache_retentionwithprompt_cache_options.ttl. - To use Programmatic Tool Calling, add the
programmatic_tool_callingtool and opt eligible tools in withallowed_callers. Update your application to handleprogramitems, program-issued function calls, andprogram_outputitems while preserving each call’scall_idandcallerlinkage. See the Programmatic Tool Calling guide for request and continuation examples.- Benchmark the PTC-enabled workflow on representative tasks. Compare task success, final-answer completeness, required evidence, total tokens, latency, and cost. Fewer calls, turns, or intermediate outputs are improvements only when the final answer still meets the required quality bar.
Prompting best practices
Favor leaner prompts
Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%. Results will vary by workload, so treat these ranges as directional and validate changes on representative tasks from your own application.
To simplify prompts without losing important guidance:
- Start with a prompt and tool set that already works. Remove one group of instructions, examples, or tools at a time, then rerun the same evals.
- State each instruction once.
- Expose only tools relevant to the task, and keep their descriptions concise and precise.
- Keep examples and style guidance when they encode a product requirement or correct a measured gap.
- Track context both at the start of a run and as the conversation grows. Long sessions can amplify repeated prompt and tool content.
Define autonomy and approval boundaries
GPT-5.6 can be proactive and persistent when carrying out multi-step tasks. Define what level of action each request authorizes so the model can continue safe, in-scope work without unnecessary pauses while stopping before external, destructive, costly, or scope-expanding actions.
A compact policy is usually sufficient:
For requests to answer, explain, review, diagnose, or plan, inspect the relevant
materials and report the result. Do not implement changes unless the request also
asks for them.
For requests to change, build, or fix, make the requested in-scope local changes
and run relevant non-destructive validation without asking first.
Require confirmation for external writes, destructive actions, purchases, or a
material expansion of scope.
Name safe local actions explicitly, such as reading files, inspecting logs, editing in-scope code, and running tests. Keep the policy in one place and state each rule once. Repeating instructions such as “ask first,” “do not mutate,” or “wait for approval” can cause unnecessary approval requests for safe, expected actions.
Set response length and style
GPT-5.6 tends to be more concise by default than GPT-5.5. When migrating, check whether broad brevity instructions such as “Be concise” or “Keep it short” are still useful. They may be unnecessary for some tasks and can sometimes make responses too brief. Keep them when they reliably produce the output your application needs.
For more consistent control across requests, use text.verbosity to set the default level of detail, then use the prompt for task-specific requirements.
Set a default with text.verbosity
Choose low, medium, or high as the default level of detail for a request. In the prompt, specify any task-specific length, structure, or required content. See Set up text.verbosity for an API example.
Specify what a short answer must include
When a task calls for a shorter answer, identify the information the model must preserve and the detail it can omit. For example:
Lead with the conclusion. Include the evidence needed to support it, any material
caveat, and the next action. Omit secondary detail and repetition.
Keep all required facts, decisions, caveats, and next steps. Trim introductions,
repetition, generic reassurance, and optional background first.
This gives the model a clear priority order: preserve the content needed to complete the task, then remove lower-value detail.
Define the tone
Broad labels such as “friendly” or “empathetic” can be ambiguous. Describe the writing choices that define your product’s tone, such as how directly to state the answer, when to acknowledge a problem, and whether reassurance or a sign-off is appropriate.
State the answer directly. If the user reports a problem, acknowledge the
specific issue before giving the next step. Use reassurance only when it is
relevant. Omit generic praise and unnecessary sign-offs.
Pro mode
Choose pro mode when quality matters most
Pro mode is a Responses API execution mode that applies more model work to a request before returning a single final answer. It can improve reliability on difficult tasks, but it increases latency and aggregates the tokens from that work in reported usage. Those tokens are billed at the selected model’s standard token rates.
Use pro mode when a marginal quality improvement materially affects the outcome and the task is difficult enough to benefit, such as complex optimization, high-value coding or review, or deep analysis with clear evaluation criteria. Prefer standard mode for routine, latency-sensitive, or high-volume work, and whenever your evaluations do not show a meaningful gain from pro mode.
Reasoning mode and reasoning effort are independent. Pro mode works with any GPT-5.6 model and its supported reasoning efforts. Start with the same model and effort as your standard-mode baseline, then compare configurations on representative tasks instead of assuming that the highest effort is always the best tradeoff.
Configure pro mode in the API
Enable pro mode in the API request. Keep the same outcome-focused prompt you use in standard mode: state the goal, relevant context, constraints, required evidence, success criteria, and output format. You do not need to ask the model to “use pro mode,” “think harder,” or generate several candidate answers.
For example:
Review this database migration plan for failure modes that could cause data loss
or extended downtime. For each finding, cite the relevant step, estimate impact
and likelihood, and recommend a specific mitigation. Return the five most
important risks in severity order.
Compare quality and cost
Compare standard and pro modes on the same representative tasks. Measure task success, answer completeness, required evidence, total tokens, latency, and cost. Use pro mode selectively where its quality or reliability gain justifies the extra model work.
Learn more in the reasoning mode guide.
Programmatic Tool Calling
Choose Programmatic Tool Calling by task shape
Programmatic Tool Calling (PTC) works best for bounded workflows where code can process several tool results or large intermediate outputs and return a much smaller structured result. Use it for filtering, joining, ranking, deduplication, aggregation, validation, or other predictable processing.
Multiple, parallel, or dependent calls alone do not justify Programmatic Tool Calling. Prefer direct, non-PTC tool calls when:
- One call is sufficient
- The intermediate outputs are already small
- Each result may change the model’s next decision
- An action requires approval
- The final output must preserve citations or native artifacts
Make routing instructions task-specific
Do not rely on tool availability or generic instructions such as “use Programmatic Tool Calling efficiently” to produce the right route. When both direct and programmatic calling are available, explicitly state:
- Which bounded stage should use Programmatic Tool Calling.
- Which tools it may call.
- The exact output schema and required evidence.
- Concurrency, retry, and stopping limits.
- Which work should remain direct.
Tool descriptions should document their expected return fields, types, and error behavior. If the model cannot determine the return shape before writing the program, prefer direct tool calling so it can inspect the result before deciding how to use it.
If both routes are needed, define one clear handoff and tell the model not to switch routes or repeat completed work.
For example:
<tool_orchestration>
Use Programmatic Tool Calling for [bounded stage] using only [eligible tools].
Run independent calls concurrently when safe. Use only documented tool input
and output fields.
Process and reduce the intermediate results, then emit exactly [output schema],
including the evidence needed for the final answer.
Stop when [condition] is met. Retry transient failures at most [R] times.
Do not repeat completed calls or perform side-effecting actions. If a required
result is still missing, return a clear structured failure.
Use direct tool calls for [semantic judgment, approval, or final validation].
</tool_orchestration>
Assess the final answer
The program_output item and final assistant message are separate outputs; make sure to test both. In theory, a program can return the correct records while the message omits a required field, citation, or caveat.
Compare direct and programmatic calling on the same representative tasks. Check whether the final response is correct, complete, and includes the required evidence. Then compare total tokens, latency, cost, calls, turns, and retries. Count lower resource use as an improvement only when the response still passes your existing evals.
Learn more in the Programmatic Tool Calling guide.