Google has introduced Gemini 4 Argon, a new large language model (LLM) designed to handle complex tasks in fields like software engineering, law, finance, and cybersecurity. The model was announced on September 30, 2026, and is currently accessible only to a select group of cybersecurity experts through Google's Fairwind program. The company is also participating in a voluntary U.S. government initiative that allows testing of models before their public release. Google plans to gradually expand access to developers, businesses, and eventually the general public after refining its security measures based on feedback from initial users. According to internal tests, Gemini 4 Argon outperforms several competing models on multiple benchmarks. On DeepSWE v1.1, which evaluates the ability to perform long software engineering tasks, Argon achieved a score of 77.9%, compared to 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. It also achieved a score of 68% on CWE-bench v1, tying with GPT-6 Astra and Grok 4.7 in detecting and correcting security vulnerabilities. On AutomationBench, a test of end-to-end execution of business tasks, Argon scored 51.3%, compared to 41.4% for GPT-6 Astra, 40% for Claude Opus 5.5, and 31.4% for Claude Fable 5.1. On LVBench, which measures long video understanding, Argon achieved a score of 91.7%, compared to 87.5% for GPT-6 Astra. The model's output token limit has been significantly increased from 64,000 to 1 million, allowing it to generate more complex responses in a single session. This feature is particularly useful for tasks such as large-scale code migrations, legal document analysis, and financial research. Google states that thousands of its employees are already using Argon internally for tasks such as optimizing quantum computing algorithms, improving memory efficiency in data centers, and rewriting large code bases in Rust. The initial price for using Gemini 4 Argon via the API is $2 per million input tokens and $10 per million output tokens. Cached input tokens benefit from a 95% discount. Google warns that this introductory price may increase after an unspecified period. The model is expected to be available to paying API customers and Google AI Ultra subscribers after the initial cybersecurity phase, though no specific date has been announced. Google emphasizes that it is implementing enhanced safeguards to prevent malicious use of the model, including measures to detect and prevent prompt injection attacks and to monitor the model's reasoning and actions. For trusted cybersecurity experts and internal teams, a version of Argon without cybersecurity safeguards is available to fully exploit its capabilities in detecting and correcting vulnerabilities. Despite its strong performance on internal benchmarks, independent evaluations have shown mixed results. On the Intelligence Index from Artificial Analysis, which aggregates ten tests, Argon scored 52.6 points at its "High" setting, compared to 57.6 points for Claude Opus 5.5 and 52.7 points for GPT-6 Astra. On GDPval-AA, which replicates professional tasks from 44 professions, Argon achieved an Elo score of 1,611, compared to 1,542 for GPT-6 Astra and 1,846 for Claude Opus 5.5. Google has not yet made Gemini 4 Argon available to all developers, businesses, and consumers, as the company is still testing the model's security measures and gathering feedback from initial users. The model's performance on real-world tasks, particularly in front-end coding, has been questioned by some employees, who suggest that it may perform less well in practical applications despite its strong benchmark results. The model has already demonstrated its capabilities in real-world applications, such as detecting a critical vulnerability in a medical software used by hospitals worldwide, which had previously gone unnoticed by existing models. Google also states that Argon has been used to optimize memory usage in its data centers, freeing up more than 300 terabytes of memory.