Series : Tech news · Part 2 of 2
Tech news: Gemini 4 Argon, when the model works longer than you do
Published
Argon can chain analysis, code, and tests in one thread — up to 1 million output tokens, far beyond usual limits.

Google announced Gemini 4 Argon on September 30, 2026. In one line: this is not a slightly smarter chatbot. It is a model built to stick with a long job — like a colleague who stays on the same problem for hours, not someone who answers and forgets.
For now you cannot freely play with it: Google is opening it first to trusted cyber defenders, then widening access.
“Tokens”? Think of a very long conversation
Models measure work in tokens (chunks of text). Before, Argon (like many others) was mostly capped on how much it could write back. Google now says up to 1 million output tokens — up from 64,000.
Concrete translation: the model has enough “breath” to chain analysis → plan → code → tests → fixes in one thread, without restarting from scratch every message. That is the difference between a text and a real work file.
Inside Google, it is already running — not vague marketing
Google shares internal cases, not only leaderboard scores:
- helping optimize quantum-computing routines (in one example: 40% better than a published baseline, in minutes);
- agents (several model instances working together) that helped free over 300 terabytes of memory across data centers — with potentially much more;
- large migrations from older C/C++ to Rust (a memory-safer language), including critical libraries and part of the Fuchsia Zircon kernel.
On an open-source video decoder (libgav1), the cited result: 2.7× faster than a prior Rust port, same video output. In plain terms: not just “suggest an idea,” but rewriting a big chunk of code under human review and tests.
Cybersecurity: full power… for a closed circle
Argon is also trained for defense: find bugs, validate them, propose patches. Google ships a version without cyber guardrails to a program called Fairwind (trusted defenders) and to its own teams.
Example with Wiz: a critical flaw in healthcare software used by hospitals — personal data at risk — that other frontier models had missed.
Everyone else waits. Google talks about a phased rollout, safeguards against misuse, and monitors so the model does not “overdo” what you asked. That is today’s paradox: the stronger the model, the more carefully it is released.
What it will cost (when it opens up)
Intro pricing announced: about $2 per million input tokens, $10 per million output; later $4 / $20. Broader access expected first via paid API and Google AI Ultra.
The benchmarks look strong — DeepSWE, finance/legal, long video, and more. Useful as a signal. Not as a single proof: Google highlights what it measures.
In short
- Argon = marathon, not a chat sprint: enough output to hold a real workstream.
- Google already uses it on big internal jobs (code, memory, quantum).
- Very strong cyber for a small circle first; the public arrives later, with more brakes.
One article at a time.