SHEYU.AI舍予基业
AI Model Watch

Google cuts image generation costs, Mistral opens trillion-parameter weights

Google slashes image pricing to a quarter of its flagship, Mistral schedules open weights for October 27, and Zhipu's GLM-5.3 joins Amazon Bedrock on a usage-share model.

·5 min read
Google cuts image generation costs, Mistral opens trillion-parameter weights
AI-generated illustration, not a news photograph

Google has cut the price of image generation to a quarter of its own flagship model, Mistral has scheduled open weights for a trillion-parameter model on October 27, and Zhipu's GLM-5.3 has entered Amazon Web Services' model marketplace on a usage-share arrangement. For teams doing overseas business, content and compliance work, the practical impact is on budgets and workflows, not benchmark scores.

The image generation bill: unit price first, rework rate second

Google released Nano Banana 2.1, built on Gemini 3.6 Flash, on October 6. According to TMTPost's Edge AI Daily briefing on October 8, it ranks fourth, fifth and sixth on Arena's multi-image editing, text-to-image and single-image editing leaderboards respectively, with OpenAI's GPT Image series occupying the top three spots.

The rankings are not the point. Price and speed are: $0.0336 per 1K image, a quarter of Google's own Pro version, and 11 seconds per image against roughly 50 seconds for GPT Image 2.5.

More telling are two capability scores. Mask editing comes in at 1049, 122 points above the Pro version, meaning users can select a region and change only that area without redrawing the whole image. Multi-character consistency scores 1106, 128 points above the previous generation, supporting up to 14 reference images (four people plus ten objects).

Anyone producing e-commerce assets knows what those two numbers mean. Keeping a model's face consistent while changing outfits has long been the bottleneck in batch image production.

A workable approach is to split assets into two layers, person and scene. Once the person is fixed, only backgrounds and props change. Every edit then falls within a selectable region, saving retries rather than just unit cost. A home goods team, for example, may pay a significant sum for a single model shoot but need a dozen seasonal scenes. Instead of reshooting or redrawing entire images, it can now lock the person in place.

One caveat. The same source shows input pricing for this version rose threefold and reasoning output 2.5-fold. In plain terms: images are cheaper, but making the model think is more expensive. Route batch image work here, but do not move long-chain reasoning tasks over at the same time.

A non-US option for teams that want to host their own models

France's Mistral released a public preview of Mistral Large 4 on October 6: one trillion total parameters, 49 billion activated per inference, a native multimodal MoE architecture, with weights scheduled for download on October 27. Training used roughly 3,800 Nvidia Grace Blackwell GPUs in the company's own European data centre. API pricing is $1.36 per million tokens for input and $4.18 for output (ITHome via Toutiao, October 6, 2026).

What may draw a second look from security and financial compliance teams is its 82% score on vulnerability reproduction and patching tests. Closed models such as Claude Opus 5.5 and GPT-6 Astra refuse such tasks outright on safety grounds and score close to zero.

If a workflow cannot let data leave the internal network, test it on one small link first, such as internal log analysis or patch verification, rather than replacing the primary model outright. The reasoning is practical: this is only a preview, the weights arrive at the end of the month, and 49 billion activated parameters will not run on a consumer graphics card.

Procurement is shifting: models are moving through cloud providers

According to Lujiazui Finance Breakfast, Amazon Bedrock has officially added Zhipu's GLM-5.3 and shares revenue with Zhipu based on usage. Domestically, Zhipu has signed similar revenue-share agreements with leading cloud providers including Alibaba Cloud's Bailian, while Huawei Cloud has listed GLM-5.3 and reached an intent to cooperate (October 7, 2026).

The keyword is revenue sharing, not integration. This is a different path from the technology licensing and API resale of the past.

For enterprise users, the direct change is in procurement. Model bills merge into cloud bills, with settlement and invoicing running through the same process, so there is no need to open a separate account with a model company. If a team already runs on AWS, the first step is to check the console for whether the model is available in the target region. Regional availability determines whether data can stay local. For specific pricing and availability, the cloud provider's page is the reference.

配图

The edge line: worth noting for privacy-sensitive products

Google also open-sourced EmbeddingGemma 2 in the same period: 740 million parameters under the Apache 2.0 licence, roughly 270 million parameters for pure text workloads and about 191MB compressed, around 567MB for full modality on the Pixel 11 Pro. It improves 9.92 points over the previous generation on the MTEB Code retrieval benchmark and supports fully offline RAG (Tencent News, Lujiazui Finance Breakfast, October 7, 2026).

This matters more to app developers. Running cross-modal search on a phone, covering photos, recordings and documents at once, previously meant either building it yourself or going to the cloud. There is now a lightweight base to build on.

A medical or legal app team, for instance, could keep historical documents on-device for retrieval, where keeping data off the network matters far more than a benchmark score. Testing is straightforward: run a retrieval pass on an offline document set first, then decide whether to replace the existing setup.

How to act

Of these four developments, the first two are cost calculations and the latter two are procurement and technical ones. Private deployment and on-device retrieval both carry barriers that most small and mid-sized teams will not need in the short term. The first question is whether your own workflow is worth changing.

The one that can be acted on today is the first: bringing down the cost and rework rate of content production. If no one on the team handles model APIs directly, and building an image pipeline for a single campaign is not worth it, consider SHEYU AGENT (舍予AI智能体). Its 34 zero-barrier tools cover Xiaohongshu posts, AI posters and copywriting, producing finished assets in one step, with a single login across desktop, mobile and web. It does not answer which large model to choose. It answers whether this week's posts and posters will ship.

One judgment worth keeping: of these developments, falling image generation costs can go straight into the books. Open weights and on-device retrieval are worth noting down, but there is no rush to change your architecture.

AI modelsimage generationopen weightscloud procurementedge AIenterprise adoption

閱讀繁體中文版 →