OpenAI's model-picking guide and Gemini's shrinking free tier: how overseas teams should cost AI by task
OpenAI cut prices and published a selection guide, Google trimmed Gemini's free tier and Anthropic made safety checks default — a sign to budget AI by task, not by model.

Three practical moves landed in the large-model world in the past week, even without a smarter new model. OpenAI published a guide on how to choose models and priced its new release at one-fifth of Astra; Google cut Gemini's free tier down to a single option; Anthropic made a safety check the default.
For overseas teams and small-business owners, the arithmetic of model selection has shifted — from "which is strongest" to "which model fits this job".
Vendors are splitting "good enough" from "best"
On October 2, OpenAI published a usage guide for its GPT-6 series on its website. IT之家's October 3 report summarised the key points: choose models by workload rather than defaulting to the strongest; use lightweight models for simple tasks, which are usually faster and cheaper; avoid writing long step-by-step prompts and instead state the desired result, the audience, what cannot be changed, and what counts as done; and do not max out reasoning effort for everyday questions — reserve it for complex analysis, debugging and deep research.
Read alone, the guide is not news. Placed next to the September 29 DevDay pricing, it is.
According to Sohu's report on OpenAI DevDay 2026, GPT-6.1 Sol approaches GPT-6 Astra in agentic coding, computer use and professional tasks, but its standard input and output token prices are about one-fifth of Astra's — $2 per million input tokens, $0.1 for cached input, and $10 for output. At the same event, the always-on agent dots, powered by GPT-6 Astra, comes with an independent cloud computer and browser and can connect to more than 4,000 applications; Codex supports cloud execution, allowing tasks to run asynchronously.
Google is moving the other way, tightening.
A Gemini app community announcement on October 3 and an ETtoday report show that from October 9, the Pro and Flash models previously available in Gemini's free tier will no longer be accessible, leaving only Flash-Lite; the Plus tier keeps Flash and Flash-Lite, with Pro unavailable; Pro and Ultra users are unaffected.
Anthropic, meanwhile, added a hard constraint to its documentation: for accounts created on or after 00:00 UTC on August 31, 2026, Claude Sonnet 5.5's thinking block will be signed before the conversation and enforced by default across the Claude API, Amazon Bedrock and Google Cloud.
One vendor cuts prices and layers its lineup, another shrinks free quotas, a third writes verification into the default.
All three are saying the same thing: model capability is being packaged as a product at different prices, different tiers and different levels of trust.
For companies this is an accounting question, not a parameter question
The conclusion first: over the next year, the cost gap in enterprise AI will come mainly not from which model is used, but from whether tasks have been properly classified.
The most common pattern is a company cramming every AI need into one chat box with the strongest model. It writes social-media copy, reviews contracts and drafts customer-service scripts — all in the same place. That is fine during trials, but once volume rises, the bill is the first thing to spin out of control.
The line in OpenAI's guide about "what counts as done" is really drawing an acceptance line for companies. Without a clear definition of completion, there is no way to judge which work can be pushed down to cheaper tiers.
The classification does not need to be complicated. Three categories will do.
High-frequency, low-difficulty work — customer-service answers, first drafts of copy, image assets — goes to lightweight models or fixed workflows; low-frequency, high-difficulty work — code review, data migration, complex analysis — is worth higher reasoning effort; and anything that must leave a trace, such as filing materials or compliance documents, cannot rely on model strength alone — the conclusion must be traceable back to the source.
Three types of work, three types of cost. That is where the real value of tiering lies.
"Wait for the strongest model to get cheaper, then standardise" does not hold up
Some owners think: models change every six months, so tiering now is wasted effort — better to wait for the strongest one to get cheap and adopt it all at once.

That judgment fails on two counts.
First, tiering is not a temporary arrangement for a transition period; it is the vendors' own commercial structure. Pricing Sol at one-fifth of Astra, and pushing Google's free tier down to a single Flash-Lite option, are deliberate moves to steer users with different usage volumes into different tiers. Waiting will not make costs fall on their own; it will only widen the price gap for the same piece of work.
Second, tiers and safety endorsements are bound together. Anthropic's decision to make signing the default reminds us that a lightweight tier means less verification; conversely, work that must leave a trace or pass review becomes more expensive when the wrong tier is used.
How to put this into practice
In tool terms, "classifying well" means not letting one entry point do all the work.
High-frequency content production — Xiaohongshu posts, posters, copy — is essentially repetitive labour and should be handed to a fixed workflow rather than having prompts rewritten each time. This kind of work can go to SHEYU AGENT (舍予AI智能体): 16 industry advisers each covering a professional area, and 34 zero-threshold tools covering Xiaohongshu image-and-text posts, AI posters and copywriting, producing output in one click; desktop, mobile and web, one login works everywhere.
As for the material work that must be traced back to source line by line, do not force a content tool to carry it — hand it to a tool that can verify the evidence chain.
The judgment
Models will keep being replaced, but the answer to model selection will not update itself.
What really deserves a quarterly review is not the version number, but which tier you assigned to which type of work. Get that clear, and both the bill and delivery quality will hold steady at the same time.