SHEYU.AI舍予基业
AI Model Watch

Model price spreads hit 200x as Zhipu joins Amazon and Grok lands on Microsoft

Zhipu's GLM-5.3 arrives on Amazon Bedrock and Grok 4.7 on Microsoft Foundry, as model prices diverge by more than 200 times and capability gaps narrow.

·4 min read
Model price spreads hit 200x as Zhipu joins Amazon and Grok lands on Microsoft
AI-generated illustration, not a news photograph

Model vendors spent the past 24 hours doing two things: moving their models onto other companies' clouds, and widening the price range. Version numbers matter less than these two shifts, which will determine where businesses buy AI next year, what they pay, and whether they have a way out. Here are three things a business lead can read in ten minutes.

The sales model has changed: models now list on other clouds

Start with how models are sold. The entry point is moving from vendors' own websites to cloud providers' shelves, and revenue is increasingly shared by usage.

On October 6, Amazon Bedrock, Amazon Web Services' large-model platform, added Zhipu's GLM-5.3. Qualifying enterprise customers can call the open-weight model directly inside Bedrock, with AWS sharing revenue with Zhipu based on usage. On the day of the announcement, Zhipu (2513.HK) rose more than 7% intraday, trading at HK$702 as of 10:39, for a total market value of HK$342.3bn (source: Yicai, republished by Sina AI Hotspot Hourly on October 6, 2026).

A day later, Grok 4.7 went live on Microsoft Foundry. The company disclosed no parameters, capability metrics or pricing. What the jump from version 4 to 4.7 means will have to wait for benchmarks (source: AI Industry News Briefing, October 7, 2026).

Taken together, the message is straightforward: the path where you registered on a vendor's website, bound a card and called an API is being absorbed into cloud providers' billing systems.

The upside is that one cloud contract may cover several models. The cost is that pricing, availability and migration paths now involve a middleman. When choosing a cloud in future, it is worth asking which other models sit on the platform, and what it costs to leave.

The maths needs redoing: a 200-fold price gap, a 3% performance gap

The same job done by a different model can produce a bill an order of magnitude apart — but there is no need to buy only the expensive one.

According to the public price lists of DeepSeek, the Kimi open platform, Alibaba Cloud Bailian, Volcano Engine Doubao, OpenAI and Claude, Tongyi qwen-flash output costs Rmb1.5 per million tokens; DeepSeek's deepseek-v4-pro peak price is Rmb8.9 for input and Rmb26.6 for output; the most expensive, gpt-6-astra, costs Rmb336 per million tokens for output, a gap of more than 200 times between the two ends. Kimi's kimi-k3 output costs Rmb100 per million tokens (source: the open platform price lists above).

Capability gaps, however, are narrowing. Robert Lea, an analyst at Bloomberg Intelligence, wrote in an October 5 report that since DeepSeek released V4.1 Flash in September, the gap between China's top models and their American competitors on benchmark scores has shrunk to 3%, from about 9% in May and 15% earlier this year (source: Bloomberg, republished by Sina AI Hotspot Hourly on October 6, 2026).

Model selection therefore should not be about picking the single strongest option, but about layering by task: high-volume, fault-tolerant work goes to cheap models; anything involving external messaging, contract terms or customer-visible content goes to the expensive ones.

The hard part is not the judgement but the execution. If every task starts with someone asking in a group chat which model to use, there is no layering at all.

Where money and compute are heading: supply keeps scaling up

Over the next year, both model availability and pricing will keep moving. Do not lock your business to a single version.

配图

On October 6, Bloomberg reported that DeepSeek is close to completing a new funding round expected to raise at least Rmb80bn (about $12bn), far exceeding an earlier target of roughly Rmb50bn. Tencent Holdings and CATL have committed larger amounts, and the final figure could approach Rmb100bn (source: Bloomberg, relayed by PANews on October 6).

On the capability side, there were three developments. Google announced Guided Vision for Gemini Live, built with the blind and low-vision community, which combines shared camera footage with real-time voice description to assist travel. MIT and Meta jointly released the DREAM model, which can understand images, generate images from text, and critique and iterate on its own drafts, speeding generation by about 10% when an external reranker is used. Meta switched the TBE kernel of its recommendation model from CUDA to FBTriton, speeding up forward propagation by 1.28 times (source: AI Industry News Briefing, October 7, 2026; the kernel change came from PyTorch's official account).

The same day, OpenAI published a set of mathematical results generated by its internal frontier models, developed with the Institute for Advanced Study's Mathematics and AI advisory group on release standards. The results were published in a GitHub repository with paper revisions, citation protocols and Lean formalisations of a large number of proofs. The repository also disclosed that each result consumed compute roughly equivalent to three hours of ChatGPT Pro thinking.

These items are somewhat distant from day-to-day operations, but they point to the same thing: models are still iterating quickly in three directions — seeing in real time, revising themselves, and pushing unit costs down. Binding a business to a specific version will become painful within six months.

How to implement it: put high-frequency creation in the tool layer, not the model layer

The real problem for most small and medium-sized businesses is not which model to choose, but whether anyone is watching these changes every day. The common situation: the marketing department needs Xiaohongshu image-and-text posts and posters, operations needs copy, and every time someone has to compare models again, stitch tools together and calculate quotas.

This kind of high-frequency, highly standardised content production is better placed in an already integrated tool layer. SHEYU AGENT (舍予AI智能体) has turned this into a finished product: 16 industry advisers, each handling a professional domain, and 34 zero-threshold tools covering Xiaohongshu image-and-text posts, AI posters and copywriting, producing output in one click; log in once on desktop, mobile or web and use it everywhere. When the underlying model raises prices, changes generation or switches cloud platform, you do not have to relearn how to operate.

Version numbers will keep jumping and prices will keep moving. For businesses, the certainty lies in keeping change outside the tool layer and keeping judgement in your own hands.

AI modelscloud platformsZhipuAmazon Bedrockenterprise procurementmodel pricing

閱讀繁體中文版 →