Gemini 4 Argon: 1 Million Output Tokens and Access
Gemini 4 Argon is being rolled out in stages, starting with security partners. The announced output limit of 1 million tokens does not indicate its input capacity or mean that regular accounts can use it immediately.
The initial rollout announced on September 30, 2026, was for trusted security experts in the Fairwind Program.
Paid API customers and Google AI Ultra subscribers were named as the first groups for a later rollout.
The 1 million token output limit alone cannot establish the input context size.
The announcement gives no specific start date for general users.
Development and security results should be read as claims in Google's announcement.
Gemini 4 Argon is an AI model being rolled out to security partners first. Its announced output limit is 1 million tokens. Paid API customers and Google AI Ultra subscribers are slated for a later release. The announcement gives no specific start date for general users.
The limits and prices in this article are based on Google’s September 30, 2026 announcement.
Gemini 4 Argon rollout
Initial access and plans for a later release should be distinguished. Google’s Korean-language announcement describes a phased rollout. There is no basis for treating the announcement date as the date every account can start using the model.
Group | Status at announcement | Timing
Security experts in the Fairwind Program | Initial rollout | Selected participants get access first
Paid API customers and Google AI Ultra subscribers | First group for a later release | No specific date given
Other developers, businesses, and general users | Subsequent phased expansion | No specific date given
Comparing output tokens and input context
Gemini 4 Argon has an output limit of 1 million tokens. Google says this is an increase from the previous 64,000 tokens. The simple ratio is 15.625 times. This compares maximum output amounts, not accuracy or speed.
Tokens are the units a model uses to process content. They do not always correspond one-to-one with characters or words. So 1 million tokens should not be read as 1 million Korean characters. These distinctions follow the Gemini API token documentation.
Category | Meaning | Can it be determined from the output limit alone?
Input token limit | How much content a request can contain | No
Output token limit | Maximum number of tokens the model can generate | The announced figure can be confirmed
Context window | Total capacity covering both input and output | Separate specifications must be checked
Actual usage | Number of tokens processed in an individual request | Usage information in the response must be checked
The output capacity can be used for long code or multistep tasks. But the maximum is not the amount generated every time. Argon’s input limit must be checked in the separate model specifications. Nor should you assume that input and output can both reach their respective maximums at the same time.
Gemini 4 Argon API pricing and calculation examples
Announced API pricing is separate from account-specific access. Google’s English-language announcement gives introductory prices. A footnote also gives prices after the introductory period. (Confirmed figures: US$4 per 1 million input tokens and US$20 per 1 million output tokens after the introductory period · Source: blog.google · checked 2026-09-30) The article does not specify when the introductory period ends.
Period | Per 1 million input tokens | Per 1 million output tokens
Introductory period | US$2 | US$10
After the introductory period | US$4 | US$20
Assuming 1 million output tokens are billed, the calculation is as follows. This is an arithmetic example based on the announced rates. It does not guarantee the actual bill or that a request can be executed.
· Output cost during the introductory period: 1 million ÷ 1 million × US$10 = US$10
· Output cost after the introductory period: 1 million ÷ 1 million × US$20 = US$20
· Total request cost: Other separately billed items, including input, must also be checked
You do not have to use the full output limit. Cost calculations must distinguish between the permitted maximum and the amount actually billed. How reasoning and tool use are billed should be checked in the applicable API pricing documentation.
Interpreting results for development, business work, and security
Google identified long development tasks and business work as major uses. Financial research and legal document drafting are also included. For security, it emphasized finding, verifying, and fixing vulnerabilities. These descriptions are based on Google’s performance announcement.
Announced example | How to interpret it
Converting C/C++ code to Rust | Work requiring auditing, testing, and review before deployment
Optimizing data center memory | An example reported in Google’s internal environment
Wiz’s discovery of a medical software vulnerability | An early use case from a security partner
Evaluations of financial and legal work | Performance reported under the conditions of those evaluations
Internal results cannot simply be applied to general accounts. There is no basis for assuming they have the same work materials and tool permissions. The long output limit alone cannot be used to calculate a task’s success rate, either. Deciding whether to adopt the model requires separate testing on the work in question.
Security partner access and restrictions on resale
Security partners cannot pass their access rights on to general users. Google DeepMind’s Fairwind Program information describes controlled access. It also specifies conditions for defensive and research use.
The Fairwind Program’s access management rules include this statement:
“we do not allow partner organizations to share, redistribute, or sell access to our frontier models.”
This prohibits partners from sharing, redistributing, or selling access. It therefore cannot be interpreted as a public release through partner accounts. Participation in the program must also be distinguished from access to individual models.
· Priority groups: Governments, major critical infrastructure operators, and core technology platforms
· Teams with access within an organization: Security, incident response, and penetration testing teams
· Management requirements: User authentication, access controls, and records of employee use
· Selection process: Review of applicant organizations’ security history and operational records
These conditions provide a basis for interpreting security-only access. Some partners’ use of Argon does not mean every participant automatically has access. It is also not the same as a general consumer’s subscription access.
Common mistakes
An announcement, access rights, and performance are different kinds of information. Using one to draw conclusions about the others leads to misunderstandings. The following distinctions are useful when translating an announcement into actual conditions of use.
Incorrect interpretation | What to check
It has been announced, so it can be used immediately | Whether the model is available to the account
Subscribers are named, so it has already been rolled out | An announcement that access has actually begun
A large output limit means more material can be submitted | The input limit and overall context
A long answer is more accurate | The accuracy and completeness of the result
Security partners use it, so resale is allowed | The program’s access management rules
Frequently asked questions
Can I use it immediately if I subscribe to Google AI Ultra?
There is no basis for guaranteeing immediate access through a subscription alone. Being named as a group for a later release is different from actually having access. Check your account’s model selection screen and later announcements.
Is Google AI Pro permanently excluded?
The lack of a specific date in the announcement does not establish permanent exclusion. Naming a group for access is different from explicitly excluding another. Until later access conditions are confirmed, this remains undecided.
Are 1 million tokens a free allowance?
No. The figure is the maximum amount the model can generate. Billing units and usage limits must be considered separately.
How many pages of Korean can 1 million tokens produce?
There is no fixed conversion to pages. Token counts vary with sentence structure. Page counts also depend on the font and layout.
Can the time a task will take be determined from the output limit?
The output limit alone cannot be used to calculate completion time. It does not measure processing speed. The time required must be measured by running the task.