This was the week OpenAI stopped being a company that only sells model access and started being a company that sells enterprise outcomes at consulting scale. GPT-5.1-Codex-Max became the default model across all Codex surfaces — the first model natively trained to compact its own context across multiple windows for multi-hour tasks. IBM launched a dedicated OpenAI Practice with thousands of certified consultants, and OpenAI’s enterprise revenue officially overtook consumer revenue. Meanwhile, the cybersecurity story intensified: GPT-5.6-Cyber shipped to approved researchers through the expanded Daybreak program, and the unreleased Astra model became the first OpenAI model to approach the “Critical” threshold of its Preparedness Framework.
And the deprecation clock keeps ticking. gpt-5.2-chat-latest and gpt-5.3-chat-latest were removed from the API on August 10. The Assistants API shuts down August 26 — twelve days away. o3 leaves ChatGPT the same day, and DALL·E GPT follows on August 30.
Here is our weekly breakdown across the Codex CLI, the model layer, the enterprise push, and the security landscape.
1. GPT-5.1-Codex-Max — The New Default, Built for Long-Horizon Work
On August 11, OpenAI announced that GPT-5.1-Codex-Max had replaced GPT-5.1-Codex as the default model in Codex surfaces: CLI, IDE extension, cloud, and code review.
The headline architectural innovation is compaction. GPT-5.1-Codex-Max is the first model natively trained to operate across multiple context windows. When it approaches the context limit, it automatically prunes less important history while preserving key decisions, touched files, errors, commands, test results, and implementation direction — then continues in a fresh window. This repeats until the task completes, enabling project-scale refactors, deep debugging sessions, and coherent work over millions of tokens in a single task.
| Spec | Value |
|---|---|
| Model ID | gpt-5.1-codex-max |
| Context window | 400,000 tokens |
| Max output | 128,000 tokens |
| Inputs / Outputs | Text + image / text |
| API | Responses API only |
| Pricing | $1.25/M input, $0.125/M cached, $10.00/M output |
Benchmark gains are significant:
| Benchmark | GPT-5.1-Codex (high) | GPT-5.1-Codex-Max (xhigh) |
|---|---|---|
| SWE-bench Verified (n=500) | 73.7% | 77.9% |
| SWE-Lancer IC SWE | 66.3% | 79.9% |
| Terminal-Bench 2.0 | 52.8% | 58.1% |
A new Extra High (xhigh) reasoning tier targets non-latency-sensitive tasks, and at medium effort the Max model outperforms the previous default while using 30% fewer thinking tokens. GPT-5.1-Codex-Max is also the first OpenAI model natively trained to operate in Windows environments — directly relevant to any enterprise with a mixed OS developer base.
One important caveat: this model is recommended only for agentic coding in Codex or Codex-like environments, not general-purpose chat. It was trained on real software engineering work — PR creation, code review, frontend coding, Q&A — and it is purpose-built to collaborate inside the Codex CLI specifically.
2. Codex CLI v0.148.0-alpha.5 — Export, Fork, Archive
The CLI continues its blistering cadence: ~2.2 days between releases, 14 releases in the last month, 73 tracked in 2026. On August 8, v0.148.0-alpha.5 landed with three workflow additions:
/export— writes the complete conversation as structured Markdown to clipboard or file. Audit-friendly session sharing without pasting raw transcripts.codex exec fork <SESSION_ID> [PROMPT]— creates a new thread from an existing session, letting you branch a long debugging session into parallel explorations.- Session archiving — long-running sessions can be archived rather than deleted, keeping history browsable without cluttering the active list.
These complement the v0.147.0 stable release from August 7 (portable Agent Plugins, --approve-for-me, MCP 2026-07-28 support, secrets redaction) and the July 29 v0.146.0 release (Agent Plugins 1.0, named sessions, thread forking, remote Code Mode over WebSocket). The pattern is clear: Codex is being tuned for long-horizon, multi-session agent work — compaction on the model side, session management on the client side.
Also noteworthy from the platform side: on August 11, OpenAI shipped a public preview of the ChatGPT desktop app for Linux (Ubuntu 24.04/26.04 LTS, Debian 13, Fedora 43/44; x64 and ARM64; native .deb/.rpm). It bundles ChatGPT, ChatGPT Work, and Codex in one native app — closing the last major OS gap roughly a month after Anthropic’s Claude desktop app for Linux. The import feature brings in projects, chats, skills, and plugins from other AI tools.
3. GPT-5.6-Cyber, Daybreak Blue/Red, and the Astra Threshold
GPT-5.6-Cyber (August 10). OpenAI introduced a cybersecurity-specific model built on GPT-5.6 Sol for authorized vulnerability research, exploit validation, and security testing:
- 95.0% completion rate on advanced cyber tasks — versus 1.5% for standard Sol and 57.3% for the prior GPT-5.5-Cyber
- Already found real zero-days, including CVE-2026-15903 (Chrome V8 out-of-bounds read/write, patched in Chrome 150.0.7871.128) and a mobile OS privilege escalation chain
- Rated “High” under OpenAI’s Preparedness Framework — below “Critical”
- Not on the public API — gated, application-only access through the Daybreak Red program
Daybreak expanded into two tiers. Daybreak Blue gives defensive teams GPT-5.6 Sol with tailored guardrails for vulnerability discovery, secure code review, malware analysis, incident response, and patch validation. Daybreak Red gives approved researchers GPT-5.6-Cyber for exploit validation and red teaming. Trusted partners include Accenture, Akamai, Cisco, Cloudflare, CrowdStrike, Fortinet, IBM, Palo Alto Networks, PwC, and Sophos — several of whom are now incorporating OpenAI cyber models into their security products. Governance is strict: identity verification, monitoring, approved-use restrictions, legal attestations — and hardware security keys become mandatory for all individual Daybreak accounts on September 1.
Astra approaches the “Critical” line. In a disclosure that landed August 7–10, OpenAI said it “cannot rule out” that its unreleased Astra model has reached the Critical tier of its Preparedness Framework — a first for any OpenAI model. Critical means a model can autonomously identify and develop functional zero-day exploits across hardened real-world systems, or devise and execute end-to-end novel cyberattack strategies from a high-level goal alone. OpenAI’s response: paused internal Astra activities that don’t meet strengthened security requirements, isolated test environments, enhanced weight protection and encryption, universal chain-of-thought monitoring across agentic applications, an automated response system that can interrupt high-risk activity mid-run, external testing with government agencies and the UK AI Security Institute — and the first U.S. government pre-release cybersecurity review of any model, per the June 2026 executive order. There is no release date; the language suggests a delay of months, not weeks.
4. IBM, the Enterprise Crossover, and the IPO Countdown
IBM strategic partnership (August 13). IBM is integrating OpenAI’s frontier models (GPT-5.6), Codex, and ChatGPT Work into IBM Consulting Advantage, and establishing a dedicated OpenAI Practice with thousands of consultants certified through the OpenAI Partner Network. IBM joins OpenAI’s Elite partner tier, with three pillars: transforming legacy business processes into AI-ready workflows, developing industry-specific AI solutions, and expanding cybersecurity collaboration via Daybreak plus IBM Autonomous Security. This is the deepest consulting-channel deal yet — it embeds OpenAI directly into IBM’s service delivery platform rather than a referral arrangement. It builds on Frontier Alliances (Accenture, BCG, Capgemini, McKinsey, Infosys, TCS) and the June IBM-OpenAI Daybreak Cyber Partner Program.
Enterprise revenue overtook consumer revenue. CFO Sarah Friar told shareholders on August 14 that enterprise revenue now exceeds ChatGPT consumer revenue — a crossover previously forecast for end-of-2026, hit months early. Annualized revenue run rate has climbed to ~$40 billion (July 2026), up 20% month-over-month, with business customers up 32%. For context: $13.1B official revenue in 2025, roughly $2B+/month in March 2026. Other business moves this week: a $7B tender offer completed at the $852B valuation (August 10), the NextSlide acquisition for presentation technology (August 10), the ChatGPT Business Premium seat tier (August 11), and a Yelp partnership letting ChatGPT users book reservations (August 14).
The public S-1 is imminent. Expected mid-to-late August — as of August 13 it was not yet on SEC EDGAR. Bookrunners: Goldman Sachs, Morgan Stanley, JPMorgan. Target listing September 2026 at $852B–$1T, though reports still allow a slip to 2027. The S-1 will reveal audited financials, the consumer/enterprise/API revenue split, the precise Microsoft revenue-sharing terms (Microsoft holds ~26.79% fully diluted), risk factors including the Astra pause and the Apple lawsuit, and the offering range. It will also land in the middle of the biggest potential IPO season in tech history — Anthropic’s confidential S-1 was filed June 1 with an October target at a $2T valuation, and xAI already IPO’d post-SpaceX-merger at a $1.25T valuation.
5. The August Deprecation Gauntlet
The deadline stack is now dense, and several items already bit:
gpt-5.2-chat-latestandgpt-5.3-chat-latestwere removed from the API on August 10 (announced May 8). Requests pinned to retired snapshots fail — OpenAI does not auto-route. Recommended replacement:gpt-5.6-sol($5/$30 per 1M tokens, 1,050,000-token context).- Assistants API shuts down August 26 — 12 days away — on both OpenAI and Azure. No automated migration, no grace period, 410 Gone after retirement, and conversation history does not migrate. Replacement: Responses API + Conversations API (OpenAI) or Foundry Agent Service (Microsoft).
- o3 leaves ChatGPT August 26 (API o3 snapshots survive until December 11).
- DALL·E GPT retires August 30 — users should download saved images now. Replacement: GPT Image 2.
- Legacy models (gpt-3.5-turbo, gpt-4) retire from the API October 23.
- GPT-5.4 and GPT-5.4 mini leave Codex for ChatGPT-authenticated sessions August 31 — they remain available via API key. Check workspace defaults and managed configs.
Pricing note for planning: GPT-5.6 Luna at $0.20/$1.20 per 1M tokens (80% cut, effective July 30) remains one of the cheapest frontier-tier options — viable for high-volume batch work, classification, and lightweight tasks. Claude Sonnet 5’s introductory pricing expires August 31, moving to $3.00/$15.00.
Final Thoughts
Three things engineering leaders should act on this week.
First, evaluate GPT-5.1-Codex-Max for long-horizon automation. The compaction architecture changes what a single agent run can do — project-scale refactors and multi-hour loops are now practical, and the 400K context with 128K output covers large monorepos. If you standardized on the previous Codex default, re-benchmark your workloads against Max at xhigh; the SWE-Lancer jump (66.3% → 79.9%) is the kind of number that pays for a migration.
Second, treat the August 26 Assistants API shutdown as a hard deadline. Twelve days. Inventory assistants, back up configurations, migrate to Responses/Conversations API or Foundry Agent Service, and verify with a test before the cutoff. If any production code still references gpt-5.2-chat-latest or gpt-5.3-chat-latest, it is already broken — fix it first.
Third, pay attention to the Astra disclosure as a governance signal. A model approaching the “Critical” cyber threshold changes the regulatory conversation — the U.S. government pre-release review process is now a fact of life for frontier labs. For enterprises, the practical implication is unchanged but sharper: agent deployments need the layered defense pattern — harness/compute separation, guardrails in the hot path, approval gates for irreversible actions, and credentials kept out of model-generated execution environments.
The next 90 days will be defined by the public S-1, the September listing window, IBM’s OpenAI Practice coming online, and whether Astra’s review process sets a new precedent for frontier model releases. The platform is maturing; the stakes are rising.
Follow the conversation on X.