Google has a new flagship AI model. You just can't use it.
Gemini 4 Argon is Google's answer to GPT-6 Astra, Fable 5.1, and Opus 5.5, and the company says it leads on coding, knowledge work, and cybersecurity benchmarks, including a 77.9 percent score on the DeepSWE v1.1 software engineering test. Google isn't shipping it to outside developers yet. For now, Argon is strictly an internal tool. Engineers have reportedly used it to migrate C/C++ codebases to Rust, including the re2 and libgav1 libraries and more than 800,000 lines in Fuchsia's Zircon kernel, while fleet-wide telemetry analysis helped save 300 TiB of memory across Google's data centers.
Announcing benchmark scores before anyone outside the building can verify them is a familiar move in this race, but Google's internal use cases here are more concrete than most. A model doing real migration work on safety-critical kernel code is a more convincing demo than another chatbot leaderboard screenshot. It also signals where Google thinks the near-term payoff is: not chat, but enterprise-scale code maintenance and infrastructure efficiency.
This was supposed to be Gemini 3.5 Pro's summer. Instead Google spent it shipping smaller Flash models and quietly building something bigger, which says as much about its release calendar as it does about Argon's benchmarks.