Google Launches Gemini 4 Argon, a Frontier Model Built for Professional Workflows
Google introduced Gemini 4 Argon, a new frontier model aimed at complex professional workflows such as software engineering, enterprise knowledge work, and cybersecurity, headlined by a massive jump in output length and strong results on coding and security benchmarks.
Details
- Output limit: supports up to 1 million output tokens per response, up from 64K tokens in prior models, letting it generate far longer code changes, reports, or documents in a single pass
- Benchmark results: scored 77.9% on DeepSWE v1.1 (real-world software engineering tasks), led the Vals Index across finance, coding, legal, and tax domains, ranked #1 on AutomationBench with 51.3%, hit 91.7% on LVBench (long video understanding), and tied for first on CWE-bench v1 at 68% (vulnerability remediation)
- Internal use cases: Google says the model is already powering internal workflows for thousands of employees, including quantum optimization work that “beat the published baseline by 40%,” freeing over 300 TiB of memory across data centers, scaling a C/C++-to-Rust codebase migration from tens of thousands to more than 800,000 lines, and a video decoder rewrite that runs 2.7x faster with added memory safety
- Pricing: introductory pricing of $2 per million input tokens and $10 per million output tokens, rising to a standard rate of $4 and $20 per million tokens respectively; cached input tokens get a 95% discount
- Rollout: launching first through the “Fairwind Program” to trusted cyber defenders, with broader availability to developers, enterprises, and consumers planned afterward
What happened next
Gemini 4 Argon lands the same week OpenAI shipped GPT-6.1 Sol and its “dots” agents and Anthropic shipped Claude Sonnet 5.5, underscoring how compressed the frontier-model release cycle has become. The model’s emphasis on cybersecurity — including a phased rollout that starts with vetted defenders rather than the general public — also signals Google treating security-sensitive deployment as a differentiator, not just a safety afterthought, as rival labs race to ship ever-larger coding and agentic capabilities.