NAIROBI, Kenya – Google has launched Gemini 3.8 Flash, a new artificial intelligence model focused on coding, reasoning and agentic workloads as the technology giant intensifies competition with OpenAI and Anthropic.
The model was unveiled on September 2, 2026, alongside Gemini 3.8 Flash Cyber, a specialised version designed for cybersecurity applications.
Google is positioning Gemini 3.8 Flash as a fast and cost-efficient model for developers and businesses seeking advanced AI capabilities without the higher operating costs associated with some larger frontier models.
The launch highlights a growing shift in the AI industry, where competition is increasingly centred not only on which model is the most capable, but also on how much useful work it can perform for every dollar spent.

Google makes coding a central focus
Software development is one of the main areas Google is targeting with Gemini 3.8 Flash.
The company describes the model as its most capable Flash model for coding and says it has been designed for demanding software-engineering and agentic workloads.
That means moving beyond simple code generation.
AI coding agents can analyse existing projects, identify problems, modify multiple files, run tests and continue refining their work based on the results.
As developers increasingly use AI throughout the software-development process, the ability to perform these multi-step tasks is becoming an important measure of a model’s usefulness.
Gemini 3.8 Flash is therefore aimed at workloads where speed, reasoning capability and operating cost all matter.
Google claims frontier-level performance
Google is also making an aggressive performance case for the new model.
The company says Gemini 3.8 Flash can deliver results comparable with advanced rival models, including Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.6 Sol, on important workloads.
Those comparisons should be treated as company-reported benchmark claims, rather than evidence that Gemini 3.8 Flash is universally superior.
AI benchmark results can vary depending on the test, prompting approach, reasoning configuration and the type of task being evaluated.
A model may perform particularly well on coding, mathematics or agentic tasks while producing different results on other workloads.
Google’s decision to compare a Flash model with larger frontier systems nevertheless illustrates how aggressively the company is pursuing efficiency.
OpenAI has similarly emphasised efficiency with GPT-5.6 Sol, positioning it as a model capable of handling coding, knowledge work, cybersecurity and scientific tasks while using fewer tokens and lowering estimated costs on some workloads.

Gemini’s price is a major selling point
The strongest argument for Gemini 3.8 Flash may ultimately be its combination of capability and price.
Google is maintaining pricing of $0.75 per million input tokens and $3.75 per million output tokens, according to reporting following the launch.
At those rates, developers can access an advanced model without paying the higher per-token prices associated with some larger AI systems.
The distinction becomes particularly important for businesses deploying AI agents at scale.
A small price difference on an individual request may have little impact on a developer. But organisations processing thousands or millions of AI operations must account for token consumption across every interaction.
Coding agents can be particularly demanding because they may repeatedly inspect files, generate code, run tests, analyse errors and make corrections.

Lower token prices do not always mean lower costs
There is, however, an important distinction between token price and cost per completed task.
A model can be cheaper per token but consume more tokens while solving a complicated problem.
The Verge reported early indications of roughly a 40 per cent increase in cost per task in some situations, linked to greater output-token usage and additional agentic evaluation steps.
That means businesses evaluating AI models may need to measure more than the advertised price of one million tokens.
The more useful question is not simply how much a million tokens costs, but how much it costs the model to complete a particular job.
For developers deploying autonomous coding agents, task-level economics could ultimately matter more than headline API pricing.
Gemini 3.8 Flash Cyber targets security teams
Alongside the general-purpose model, Google announced Gemini 3.8 Flash Cyber, a specialised model aimed at cybersecurity.
Unlike the standard Flash model, the Cyber version is being offered through Google’s Fairwind Program to a limited group of trusted cybersecurity defenders rather than being released broadly.
Google says the model is designed to help security teams identify and fix vulnerabilities.
The company also says Gemini 3.8 Flash Cyber delivers frontier-level performance on the CyberGym benchmark for autonomous vulnerability discovery, outperforming its previous cybersecurity model and significantly larger frontier systems.
Those claims remain subject to independent evaluation, but they demonstrate Google’s broader strategy of applying increasingly capable AI agents to specialised technical work.

What Gemini 3.8 Flash could be used for
A faster and relatively inexpensive model capable of handling complex programming tasks could expand the range of activities developers are willing to automate.
Potential applications include:
- Generating and modifying software
- Debugging code
- Reviewing existing projects
- Analysing large codebases
- Running and evaluating software tests
- Automating repetitive development tasks
- Building AI-powered agents
- Supporting enterprise workflows
- Identifying and addressing cybersecurity vulnerabilities
The economics of these applications will depend on more than token prices. Reliability, latency, context handling, tool use and the number of attempts required to complete a task can all affect the final cost.




