What changed on 22 September
OpenAI introduced Sol and Luna as reasoning models for Work and Codex and as API models. The API accepts text and images and produces text. What an employee can select depends on their plan, product surface and workspace policy. A screenshot from a different account is therefore insufficient evidence for a rollout. Record the account, model identifier and reasoning level used in your pilot. Codex desktop preserves a model that was chosen manually. Check existing tasks explicitly: installing an application update does not mean every task is already using the new model.
Compare three tasks rather than picking a universal winner
Build a small task mix: a tightly specified classification, an analysis of conflicting documents and a change to an internal tool. For classification, check whether categories are complete and reproducible. For analysis, inspect supported claims and missing information. For the tool, assess working behaviour and understandable changes. Use identical inputs and instructions for both candidates. Start with the default reasoning setting and increase it only when there is a specific quality problem. This reveals which configuration supports a particular job, rather than assuming capability from a model name or treating an impressive isolated answer as a universal result.
Example: preparing incoming quote requests
A technical service provider could classify incoming enquiries and extract required information first. A second step compares requirements with an approved service description and drafts clarification questions. This is an illustrative process, not a claimed customer deployment. Decide beforehand which missing fields are acceptable and when a person must take over. Include ambiguous deadlines, conflicting quantities and services you do not offer in the test set. The draft must not invent prices or commitments. Only use a cheaper configuration for the first step if it passes the same acceptance criteria; more demanding cases can follow a separately evaluated route.
Separate API pricing from the cost of a finished result
The checked API rates for standard requests with up to 272,000 input tokens are, per million tokens: Sol, USD 2 input, USD 0.20 cached input and USD 10 output; Luna, USD 0.10, USD 0.01 and USD 0.50 respectively. These are provider rates checked on 23 September, not ChatGPT subscription prices or a Werkverstand quotation. Longer prompts, other processing tiers and tools require separate checks. Include retries and human correction time in the pilot calculation. A short extraction has a different cost profile from a long investigation involving multiple tools. Compare cost per accepted deliverable rather than price per token alone.
Turn the experiment into a team decision
Use a fixed set of completed cases with known outcomes. Remove unnecessary personal data and keep the inputs stable throughout the comparison. Score completeness, evidence, formatting, turnaround and rework separately. Assigning information to the wrong customer or order is a critical failure that a good average must not conceal. Also record who may change model settings and which tasks pause when limits are reached. After a successful pilot, switch one workflow first. Review early production results closely and retain a manual route when access, quality or cost differs from the experiment. This makes the decision reproducible for the next person maintaining the process.
Common implementation questions
Must every task switch immediately? No. A new generation is a reason to evaluate; an existing successful workflow needs an evidence-based reason to change. Is Luna automatically best for simple work? Even a simple task can have serious error consequences, such as confusing order identifiers or personal records. Does a ChatGPT subscription include API access and usage? Treat the access and billing paths separately. Which metric decides? For a defined workflow, use the share of correctly completed cases at sustainable total cost. If you do not yet have a useful comparison case, the AI System Check helps identify a bounded task with an outcome that can be reviewed.
Keep it verifiable
Primary sources
- OpenAI: GPT-6 Sol modelSource checked:
- OpenAI: GPT-6 Luna modelSource checked:
- OpenAI: Work and Codex modelsSource checked:
- OpenAI: API changelogSource checked:



