Over the past week, the artificial intelligence sector has produced noteworthy signals across multiple dimensions—commercial competition, model safety, embodied AI, and product strategy. OpenAI’s coding agent Codex has begun to overtake Anthropic’s Claude Code in commercial growth, while DeepSeek has quietly shifted its pace on multimodality, launching an experimental visual understanding model. Meanwhile, NVIDIA has reportedly signaled price increases for full server systems, and the embodied AI field has reached what industry insiders are calling a “GPT-2 moment.”

Codex’s Commercial Growth Surges; Anthropic’s Flagship Model Hits Headwinds

Over the past month, the competitive dynamics between OpenAI and Anthropic in the coding agent space have shifted markedly. According to TrickerTrends metrics, Codex posted a 20.8% annualized revenue growth rate over the past four weeks, compared with just 5.2% for Claude Code. Enterprise data tells a similar story: Ramp corporate spending figures show OpenAI grew approximately 82% quarter-to-date in Q3, versus roughly 76% for Anthropic.

This stands in stark contrast to Q2. At that time, Anthropic reported revenue of approximately $11.6 billion, up more than 100% quarter-over-quarter. According to a Wall Street Journal report on August 18 citing people familiar with the matter, OpenAI’s Q2 revenue rose from $5.7 billion in Q1 to $6.7 billion—a mere 18% sequential increase. In just three months, the growth trajectories of the two companies have reversed.

Anthropic’s slowdown stems primarily from two factors. First, enterprise adoption of its new flagship model, Fable 5, has been underwhelming. Ramp data indicates that GPT-5.6 Sol accounts for 25% of OpenAI’s enterprise tokens and 23% of model spending, while Fable 5 represents only 6% of Anthropic’s tokens and 11.4% of spending. The model was taken offline just three days after its June 9 launch due to a U.S. government directive, and only restored roughly 18 days later. Its pricing of $10 per million input tokens and $50 per million output tokens is approximately 2.5x GPT-5.6 Sol’s current promotional rates of $4 and $20 respectively, and some enterprises also face a 30-day data retention requirement.

On the other hand, Codex itself has been catching up rapidly, with particularly successful operational strategy. It has been bundled into Plus, Pro, Business, and Enterprise subscriptions, with some users receiving free access, doubled quotas, or promotional credits. ChatGPT Business annual pricing dropped from $25 to $20 per seat, and eligible team members and students can receive $100 in credits. Codex’s weekly active users surpassed 5 million by June, a sixfold increase since February, with overall monthly active users recently exceeding 10 million.

However, truly closing the gap with Anthropic is no easy task. The latest reports put annualized revenue at approximately $65 billion for Anthropic and $40 billion for OpenAI—a difference of roughly 1.63x. Even if Anthropic stopped growing entirely, OpenAI would need to grow another 62.5% to catch up. More importantly, there is a difference in revenue quality. In Vercel’s June production data, Anthropic captured 61% of spending with 32% of tokens, and secured at least 72% of spending in tasks such as coding agents, backend agents, and application generation; OpenAI accounted for 10.3% of tokens and 16.1% of spending. Claude still controls the tasks for which enterprises are most willing to pay premium prices.

Frontier Model Safety Strategies Diverge: Withhold vs. Strict Control

Frontier labs have made superficially similar but fundamentally different choices on model safety. On August 18, OpenAI disclosed that the company had previously paused reinforcement learning training for models slated for deployment for two weeks. As of the announcement, the largest-scale frontier RL training had not yet resumed, though smaller-scale training and evaluations continued.

According to SemiAnalysis editor-in-chief Dylan Patel in an August 17 interview, Anthropic’s next-generation model Mythos 2 has completed training but will not be released. The model has not become a product, but has instead become an internal production asset—capable of generating code, tests, evaluation questions, and training trajectories to help researchers improve post-training, toolchains, and next-generation experiments.

OpenAI, by contrast, has directly tightened certain training and inference workloads because Astra may possess “critical cybersecurity capabilities.” The model has demonstrated the ability to execute code, invoke tools, access internal systems, and sometimes reach the network during RL and evaluation. OpenAI has therefore required that any model at Sol level or above that uses tools during RL or evaluation be routed into a new monitoring system. This system scans model activity token-by-token, then escalates suspicious behavior to more computationally intensive automated investigators. The highest-priority alerts simultaneously notify safety, research, and security teams. If a false positive cannot be confirmed within 30 minutes, the relevant activity should be paused.

OpenAI estimates that monitoring consumes approximately 20% of the inference compute being monitored, beginning to eat into cluster budgets and affecting training throughput, permission configurations, and experiment scheduling. The safety team has also shifted from being release reviewers to becoming part of training scheduling.

DeepSeek Quietly Launches Vision Model; Multimodality Still Positioned as a “Component”

On August 21, DeepSeek launched DeepSeek V4 Flash Vision Exp, which the company describes as a multimodal visual understanding model. Unlike V4 Flash and V4 Pro, this model comes with no technical report and is not open-sourced—the company released only a performance table and a few demonstration cases.

Notably, the benchmarks in the official release are predominantly agent benchmarks: Terminal Bench 2.1 scored 83.9, and DeepSWE scored 59.3. Even the multimodal-side benchmarks are “unconventional”: Chartography tests professionals reading specialized charts for decision-making, ApexBench is a multimodal agent benchmark, and Agents’ Last Exam is a general agent evaluation. None of them are the traditional “image question-answering” tests commonly used for vision models.

The benchmark table itself is a signal. DeepSeek founder Liang Wenfeng previously stated: “We have always been working on multimodal deployment. For products, it is important; for consumer-facing products, it is important. But for the upper limit of intelligence, it is a component, not the main line itself.” The value proposition demonstrated by DeepSeek V4 Flash Vision Exp is that multimodality has never been about letting people “chat with images”—it is about enabling agents to read webpage screenshots, interpret chart data, understand UI layouts, and then complete tasks.

The underlying technology of this model can be traced back to a technical report DeepSeek published on GitHub on April 29, titled “Thinking with Visual Primitives.” The paper proposed elevating point coordinates and bounding boxes to the status of “minimal units of thought,” embedding them directly into reasoning chains so that models can “point while reasoning.” The paper was deleted in the early hours of May 1, and related announcement posts also disappeared, but the Vision mode DeepSeek launched on its client and web platforms in June is built on this visual primitives mechanism.

Community feedback on this experimental model has been lukewarm. Some users reported that it misses key information when interpreting complex charts, handles Chinese-language scenarios relatively poorly, and even fails to recognize one of Liang Wenfeng’s most iconic photographs. But the release cadence itself signals a shift—Liang Wenfeng was previously known for “never releasing anything imperfect,” with four and a half months between V3.2 and the V4 preview, yet now the company is willing to open experimental models to the public.

NVIDIA Reportedly to Raise Server Pricing by More Than 15%

According to a Bloomberg report on August 22 citing people familiar with the matter, NVIDIA has notified some major customers that many server configurations equipped with AI chips such as Vera Rubin and Grace Blackwell will see price increases of more than 15% upon delivery in early 2027, primarily due to rising memory costs.

This development is worth watching. Over the past year, the industry has been calculating how quickly the price of per-unit intelligence drops after new model launches. But per-token price declines only reflect one marginal cost of model serving, while training and inference clusters must also pay for high-bandwidth memory, CPUs, networking, racks, power, cooling, and delivery timelines. The cheaper models become and the more densely agents are invoked, the more likely enterprises are to increase total call volume. If memory and full-system prices rise simultaneously, the budget saved on the model side will be consumed on the infrastructure side.

Embodied AI Reaches Its “GPT-2 Moment”

On August 19, Generalist released its robot foundation model GEN-1.5. The company claims that after watching 3 to 12 seconds of sensor-action demonstrations, the model can execute new tasks without updating parameters, achieving an average success rate of 59% across ten short-horizon manipulation tasks. With 1 to 5 minutes of training—approximately 10 to 50 demonstrations—the model needs to update only 1 to 10 steps, achieving an average success rate of 83% across ten steps, with weight changes of less than 0.15%.

Beyond imitation, the model can switch hands to open lids, clear obstacles, or handle stuck objects. The company also acknowledged that testing focused primarily on short-horizon tasks such as zipping, opening cans, and object retrieval, and that in-context learning is more fragile than fine-tuned models.

This release continues a year-long technical evolution. In November 2025, GEN-0 demonstrated the scaling effects of physical pretraining; in February 2026, DreamZero formally proposed the World Action Model (WAM); in April, GEN-1 used over 500,000 hours of real interaction data to achieve success rates above 99% on multiple simple tasks; and the Behavior Prompting Policy released in June found that task diversity is the primary driver of behavior prompting capability.

GEN-1.5 can be viewed as a representative system following this year’s convergence of embodied AI approaches. Generalized WAM-style temporal modeling, diverse pretraining, real interaction data with error-correction processes, and rapid test-time adaptation all appear in its documentation. Pretraining has already brought the model close to mastering many basic actions; a few seconds of demonstration mainly tells it “what to accomplish this time,” then combines existing capabilities.

However, the field has only reached the GPT-2 moment—several critical steps remain before the ChatGPT equivalent. Effective long-horizon tasks remain unsolved, and the 59% learning rate from single demonstrations only covers the short-horizon atomic tasks the company selected, with relatively narrow generalization.