Update, 19 September 2026: Google now offers stable Gemini 3.8 Flash. This article preserves the July comparison between 3.6 and 3.5. If you are choosing a model now, start with our Gemini 3.8 Flash guide and test it against your existing workflow before switching.
Google launched Gemini 3.6 Flash on 21 July 2026, only nine weeks after 3.5 Flash and barely seven months after Gemini 3 Flash arrived in December. Three Flash generations in eight months is a pace worth noticing in itself. The headline specification has barely moved: both models in this comparison accept roughly one million input tokens, both can return up to 65,536 tokens, and both work with text, images, video, audio and PDFs. The practical changes are easier to miss and more important to a business. Output tokens are cheaper. Google says the model completes agentic work with fewer reasoning steps and tool calls. Its selected coding, computer-use and knowledge-work benchmarks have improved. Independent testing, meanwhile, found the clearest gains in speed and efficiency rather than in the aggregate intelligence score. That makes this less a dramatic new generation than a careful attempt to remove waste from the previous one, and for a business that may be the more valuable kind of improvement. The answer in brief: Gemini 3.6 Flash became generally available on 21 July 2026 and has since been followed by 3.7 and 3.8 Flash. Standard paid API pricing is $1.50 per million input tokens and $7.50 per million output tokens, against $1.50 and $9 for 3.5 Flash. In pre-release high-thinking tests published by Artificial Analysis on launch day, both models scored 50 on its composite Intelligence Index, while 3.6 completed the Index tasks in half the time and at lower measured cost. For a new build, include 3.8 Flash in your evaluation. Existing workflows should move only after regression testing, and Google’s current deprecation schedule should determine migration deadlines.What is Gemini 3.6 Flash?
Gemini 3.6 Flash is a proprietary, natively multimodal model from Google. It sits between the cheapest high-volume models and the most expensive frontier models: capable enough for coding, document work and multi-step agents, but fast and economical enough to use repeatedly. The stable API model ID is gemini-3.6-flash. It accepts text, images, video, audio and PDFs and returns text. The official model page lists Search and Maps grounding, URL context, file search, function calling, code execution, structured outputs and thinking as supported features, with Computer Use in preview. Native image generation, audio generation and the Live API are not supported by this model. Google describes 3.6 Flash as a model for rapid agentic loops, complex coding cycles and spatial reasoning. It is generally available through the Gemini API and Google AI Studio, and Google’s launch announcement also lists the Gemini app, Android Studio, Antigravity and its enterprise products. Availability does not mean identical limits for every user: plan, region, product and usage caps still apply.Gemini 3.6 Flash at a glance
If the table extends beyond the screen, scroll sideways to view all columns.
| Specification | Gemini 3.6 Flash | Why it matters |
|---|---|---|
| Release status | Gemini 3.6 FlashGenerally available | Why it mattersSuitable for production evaluation, not merely a preview |
| Release date | Gemini 3.6 Flash21 July 2026 | Why it mattersVery new at the time of this review |
| Stable model ID | Gemini 3.6 Flashgemini-3.6-flash | Why it mattersThe identifier used in API requests |
| Inputs | Gemini 3.6 FlashText, image, video, audio and PDF | Why it mattersOne workflow can interpret several media types |
| Output | Gemini 3.6 FlashText | Why it mattersUse another model for native image or audio generation |
| Maximum input | Gemini 3.6 Flash1,048,576 tokens | Why it mattersUseful for substantial document sets and long working context |
| Maximum output | Gemini 3.6 Flash65,536 tokens | Why it mattersSupports long reports and code output |
| Default thinking level | Gemini 3.6 FlashMedium | Why it mattersThe same default as 3.5 Flash |
| Knowledge cut-off | Gemini 3.6 FlashMarch 2026 | Why it mattersNewer facts still need supplied sources or grounding |
| Standard paid API price | Gemini 3.6 Flash$1.50 input / $7.50 output per million tokens | Why it mattersOutput is 16.7% cheaper than 3.5 Flash |
| Batch and Flex price | Gemini 3.6 Flash$0.75 input / $3.75 output per million tokens | Why it mattersLower-cost routes for suitable workloads |
| Built-in tools | Gemini 3.6 FlashSearch, Maps, file search, URL context, code execution, function calling, Computer Use (preview) | Why it mattersThe model can take part in workflows that act as well as answer |
Gemini 3.6 Flash vs Gemini 3.5 Flash
This is an iteration, not a clean-sheet model. The input formats and headline token limits are unchanged. The difference lies in the route to the answer. Google says 3.6 reaches its results with fewer reasoning steps and tool calls, reports better instruction-following, and highlights stronger work on charts and multi-element visual layouts. These improvements matter if they survive contact with a real workflow: an unnecessary tool call costs money, and an unnecessary file edit costs review time.If the table extends beyond the screen, scroll sideways to view all columns.
| Area | Gemini 3.5 Flash | Gemini 3.6 Flash | Practical reading |
|---|---|---|---|
| GA release | Gemini 3.5 Flash19 May 2026 | Gemini 3.6 Flash21 July 2026 | Practical reading3.6 is the newer stable model |
| Input context | Gemini 3.5 Flash1,048,576 tokens | Gemini 3.6 Flash1,048,576 tokens | Practical readingNo increase |
| Maximum output | Gemini 3.5 Flash65,536 tokens | Gemini 3.6 Flash65,536 tokens | Practical readingNo increase |
| Standard input price | Gemini 3.5 Flash$1.50 per million | Gemini 3.6 Flash$1.50 per million | Practical readingNo change |
| Standard output price | Gemini 3.5 Flash$9 per million | Gemini 3.6 Flash$7.50 per million | Practical reading3.6 is 16.7% cheaper on output |
| Standard cache-read price | Gemini 3.5 Flash$0.15 per million | Gemini 3.6 Flash$0.15 per million | Practical readingNo change |
| Cache storage | Gemini 3.5 Flash$1.00 per million tokens per hour | Gemini 3.6 Flash$1.00 per million tokens per hour | Practical readingRetained context has a running cost |
| Default thinking | Gemini 3.5 FlashMedium | Gemini 3.6 FlashMedium | Practical readingA familiar starting point, not identical behaviour |
| Knowledge cut-off | Gemini 3.5 FlashJanuary 2025 | Gemini 3.6 FlashMarch 2026 | Practical readingThe newer model has a later training cut-off |
| Model lifecycle | Gemini 3.5 FlashNo shutdown date announced | Gemini 3.6 FlashNo shutdown date announced | Practical readingThere is no forced emergency migration |
What does Gemini 3.6 Flash cost?
The standard paid input price has not changed. Output has fallen from $9 to $7.50 per million tokens. Google also offers lower Batch and Flex rates, while Priority costs more in exchange for prioritised inference.If the table extends beyond the screen, scroll sideways to view all columns.
| Gemini Developer API route | Input per million tokens | Output per million tokens | Typical fit |
|---|---|---|---|
| Standard | Input per million tokens$1.50 | Output per million tokens$7.50 | Typical fitInteractive and ordinary production requests |
| Batch | Input per million tokens$0.75 | Output per million tokens$3.75 | Typical fitLarge asynchronous jobs |
| Flex | Input per million tokens$0.75 | Output per million tokens$3.75 | Typical fitFlexible work where lower cost matters more than guaranteed immediacy |
| Priority | Input per million tokens$2.70 | Output per million tokens$13.50 | Typical fitWorkloads paying for prioritised inference |

Worked monthly examples
If the table extends beyond the screen, scroll sideways to view all columns.
| Illustrative workload | 3.5 Flash | 3.6 Flash | Saving at identical token use |
|---|---|---|---|
| 10m input + 2m output tokens | $33 | $30 | $3 (9.1%) |
| 100m input + 20m output tokens | $330 | $300 | $30 (9.1%) |
| 20m input + 20m output tokens | $210 | $180 | $30 (14.3%) |
Google’s benchmarks and the independent evidence
Launch benchmarks answer a narrow question: how did the models perform in the publisher’s chosen test set-up, with its chosen settings? They can be useful without being the final word. In the four comparisons highlighted in Google’s launch post, 3.6 scored higher:If the table extends beyond the screen, scroll sideways to view all columns.
| Evaluation | Gemini 3.5 Flash | Gemini 3.6 Flash | What it covers |
|---|---|---|---|
| DeepSWE | Gemini 3.5 Flash37% | Gemini 3.6 Flash49% | What it coversLong-horizon software engineering |
| MLE-Bench | Gemini 3.5 Flash49.7% | Gemini 3.6 Flash63.9% | What it coversMachine-learning engineering and research |
| OSWorld-Verified | Gemini 3.5 Flash78.4% | Gemini 3.6 Flash83.0% | What it coversComputer-use tasks in desktop environments |
| GDPval-AA v2 | Gemini 3.5 Flash1,349 Elo | Gemini 3.6 Flash1,421 Elo | What it coversReal-world knowledge work |

What Artificial Analysis found
Artificial Analysis tested the high-thinking variants ahead of release, publishing on launch day. Its aggregate Intelligence Index score was 50 for both models. That tie does not establish equal capability on every task; it says the gains and losses across that particular composite produced the same rounded total. The efficiency differences were clearer:If the table extends beyond the screen, scroll sideways to view all columns.
| Artificial Analysis launch snapshot | 3.5 Flash (high) | 3.6 Flash (high) |
|---|---|---|
| Intelligence Index | 50 | 50 |
| Average time per Index task | 2.7 minutes | 1.3 minutes |
| Output speed | 159 tokens/ |
304 tokens/ |
| Weighted cost per Index task | $0.59 | $0.50 |
| Output tokens across the Index | 75m | 59m |
| Humanity’s Last Exam | 41% | 38% |
What else shipped alongside it?
Three related announcements matter for planning. Gemini 3.5 Flash-Lite arrived the same day at $0.30 input and $2.50 output per million tokens, and Google says it is rolling into Google Search, which makes it one of the models reading and summarising web content inside AI-powered results. Gemini 3.5 Flash Cyber, a security-focused variant built on 3.5 Flash, remains in a limited pilot for governments and trusted partners. And Gemini 3.5 Pro, promised for June when the 3.5 family launched, was still in partner testing at the time of writing, while Google says pre-training for Gemini 4 has begun. The Flash tier is carrying the roadmap while the larger models finish.Where the newer model makes business sense
The strongest use cases sit between a cheap classifier and a costly frontier model: complicated enough to need reasoning and tools, frequent enough for efficiency to matter.Document and report analysis
A 1,048,576-token input limit can accommodate substantial document sets, depending on file size, extraction quality and tokenisation. File search, code execution and improved chart reasoning make 3.6 a plausible first-pass engine for evidence extraction, report comparison and management summaries. The model should not become the source of record. Keep the original files, require source references in the output and validate important figures before they enter a board paper, client report or public article.Marketing analysis and content operations
There is a sensible role between the dashboard and the strategist. A controlled workflow can combine campaign exports, search-query reports, creative results and landing-page data, then draft a weekly narrative or flag patterns for review. Structured outputs make the result easier to validate before another system acts on it. Do not give the model an undefined mandate: spending changes, regulated claims, publication and customer-data decisions need explicit controls and named owners.Coding and website maintenance
Google has concentrated much of the update on code and agentic execution. Fewer unwanted edits and shorter debugging loops would be valuable in a real repository, where a technically valid change can still damage an unrelated page. There is a useful caveat in Google’s latest-model guidance: human evaluators preferred earlier models for some visual layout and styling tasks. Functional code and good art direction are not the same skill. Give the model a design system, reference screens and acceptance criteria, then test the result.Customer-service assistance
Long context, file search and grounding suit an assistant that finds the relevant policy and prepares a response for an adviser. Latency still needs testing in the actual channel: a fast output rate is little comfort if a customer waits through a long reasoning pause before the first useful sentence. Access control matters just as much. The assistant should retrieve only the records needed for the current case, with a clear route to a person when confidence is low or the consequence is high.Multi-step agents
This is the job Google has designed around: a workflow that plans, calls tools, checks the result and continues. Competitor monitoring, catalogue quality checks, code maintenance and research packs are plausible examples. The unglamorous parts make the agent usable: permissions, logs, step and cost limits, schema validation, source checks and human approval for material actions.If the table extends beyond the screen, scroll sideways to view all columns.
| Workload | Fit | Main reason | Main caveat |
|---|---|---|---|
| Multi-step research agent | FitStrong | Main reasonLong context, grounding and improved tool use | Main caveatVerify sources and recent facts |
| Document and chart analysis | FitStrong | Main reasonMultimodal input and stronger spatial reasoning | Main caveatTest complex layouts and extraction quality |
| Coding and migrations | FitStrong | Main reasonBetter instruction-following and fewer reported edit loops | Main caveatRun tests and review every diff |
| High-volume extraction | FitMixed | Main reasonCapable, but possibly more model than the job needs | Main caveatCompare with 3.5 Flash-Lite |
| Live customer chat | FitMixed | Main reasonFast generation once output begins | Main caveatMeasure time to first useful answer |
| Brand-led visual design | FitMixed | Main reasonCan produce functional interfaces | Main caveatSupply clear art direction |
| Native image or audio generation | FitPoor | Main reasonThese output modes are unsupported | Main caveatUse a suitable image, audio or Live model |
| Unsupervised high-stakes decisions | FitPoor | Main reasonHallucinations remain possible | Main caveatRequire qualified human judgement |
When should a business keep Gemini 3.5 Flash?
Keep it temporarily when a production integration is stable and the cost of behavioural change is higher than the likely saving. Prompts, parsers, safety rules and tool contracts can depend on habits that never appear on a model card. Deep-reasoning workloads deserve particular attention, because the independent Humanity’s Last Exam result moved backwards in the launch snapshot. That is not proof that your task will regress. It is a reason to include difficult, multi-step cases in the test set instead of relying on an average. At the original July check, there was no announced deadline. Google’s deprecation schedule listed gemini-3.5-flash with no shutdown date announced when checked on 24 July 2026, and gemini-3.6-flash did not appear on the schedule. That is a dated observation, not a promise of continued support. Before starting or migrating a project, check current availability and compare 3.8 Flash with the model you already use. The July price comparison below remains useful for understanding the change from 3.5 to 3.6.A practical migration checklist
Changing the model ID is the easy line. The surrounding behaviour is the work.- Build a fixed evaluation set from real, permission-cleared production tasks.
- Run both models at equivalent thinking levels and through the same tools.
- Record task success, latency, token use, tool calls, retries and human correction time.
- Inspect failure types, not only averages.
- Validate every structured output and external tool result.
- Review permission boundaries, logs, safety behaviour and cost caps.
- Route a small share of live traffic to 3.6 while keeping a rollback path.
- Increase traffic only when the accepted-result rate and end-to-end cost are better.
Limits, privacy and controls
The DeepMind model card explicitly notes hallucinations, occasional slowness and timeouts. The March 2026 knowledge cut-off is another ordinary constraint. Search grounding can improve freshness, but a grounded answer can still omit context or overstate what a source supports. Five controls cover most serious deployments. Source control: supply approved material or grounding routes, preserve citations and check each consequential claim. Action control: give an agent the minimum permissions needed for the job; reading a report and changing a campaign budget are different risk classes. Cost control: cap steps, tool calls, tokens and retries, and alert when a workflow loops. Quality control: validate schemas automatically and keep qualified people in the review chain. Data control: decide what may be sent to the service, under which terms, with what retention and regional requirements. Google’s pricing page says free-tier content may be used to improve its products, while paid-tier content is not. That table row is not a complete privacy assessment. Organisations handling confidential, personal or regulated data should review the applicable API or enterprise terms and configure the workflow around their obligations.What does this change for SEO and AEO?
The model matters to search marketing in two ways. It can be used as a production tool for research, clustering, reporting and content operations. It also illustrates the wider shift towards answer-led interfaces and agents that retrieve information from websites on a user’s behalf. It does not create a secret ranking recipe. Google’s guide to optimising for generative AI features, published in May 2026, says ordinary SEO still underpins visibility. Pages need to be indexed and eligible for snippets, and a site must be included in Search generative AI features in Search Console. There is no special AI schema, no required llms.txt file and no reason to split useful prose into tiny AI-sized chunks; Google states plainly that Search does not use such files.Structured data still earns its keep
Not required is not the same as not useful. Structured data remains one of the clearest ways to tell every machine, not only Google, who published a page, who wrote it, when it was last reviewed and what entity stands behind the site, and other search and answer systems read it too. The value now lies less in individual snippets than in the connected graph: Organization, WebSite, WebPage, BlogPosting, Person and image nodes linked together by ID references, so that a crawler can assemble one coherent picture of the site instead of a drawer of loose facts. This is what some SEO tools describe as schema aggregation, and the good news for WordPress publishers is that Yoast and Rank Math already build the graph automatically. The real work is feeding it properly: a correct organisation profile with logo and sameAs links to your public profiles, genuine author entities rather than “admin”, honest published and modified dates, and markup that always matches the visible page. Spend nothing on retired display types; FAQ rich results ended in May 2026, so FAQPage markup earns no Google feature, even though the questions and answers themselves remain exactly what answer engines like to quote.Where llms.txt fits
llms.txt deserves a candid paragraph, because the tooling has run ahead of the evidence. It is a community proposal, a markdown index at the root of a site pointing machines at your most important pages. Google is unambiguous that Search and its generative features do not use it, and no major assistant has documented reading it for answers at the time of writing. At the same time, Yoast now generates one from Site features settings, Rank Math and several dedicated plugins do the same, and Yoast’s own framing is carefully modest: a way to offer language models a preview of your most important, current content, with no promises about who consumes it. Our advice sits in the same place. If your SEO plugin can maintain the file automatically, switch it on; it costs nothing, keeps a tidy machine-readable index of the pages you care about should assistants adopt the convention, and does no harm meanwhile. Treat it as preparedness rather than performance, and be sceptical of anyone selling llms.txt optimisation as an AEO strategy. What verifiably does govern AI visibility today is crawler access, and the user agents are not interchangeable. For ChatGPT, OAI-SearchBot governs appearance in search answers while GPTBot concerns model training, and OpenAI notes that a site opted out of OAI-SearchBot can still surface as a bare navigational link. Anthropic separates ClaudeBot, its training crawler, from Claude-SearchBot and Claude-User, which handle search quality and user-requested visits. Perplexity draws the same line, with one wrinkle worth knowing: PerplexityBot respects robots.txt, but Perplexity-User generally ignores it because a human asked for the page. And Google-Extended controls certain Gemini training and grounding uses without affecting inclusion or ranking in Google Search, which follow ordinary indexing and snippet controls. Review robots.txt and firewall rules service by service rather than blocking by reflex; each door you close removes a place you can be cited.The same principles on custom-built sites
Not every site has a plugin to lean on, and the principles port directly to custom builds. The first check on a React or otherwise JavaScript-heavy front end is rendering: Google executes JavaScript when it indexes, but assistant crawlers have been consistently observed reading only the raw HTML response, so content that exists solely after client-side rendering is invisible to the very systems this section is about. Server-side rendering or static generation through a framework such as Next.js closes that gap, and Google’s own guide points the same way when it asks for important facts to be visible in HTML rather than only in scripts or images. Structured data follows the same rule: emit the JSON-LD graph server-side, whether through Google’s schema-dts types in a TypeScript build, a package such as Spatie’s schema-org for Laravel, or hand-authored templates in plain PHP, and keep the organisation and author nodes identical on every page so the graph stays coherent. An llms.txt file on a custom stack is simply a build artefact, best generated from your routes or CMS at deploy time rather than maintained by hand. And treat heading discipline as a template concern, not an author’s memory test: one H1 per page, sections at H2, questions at H3, whatever the framework renders. The work is familiar, but the standard is higher. Answer the main question early, then earn the detail. Publish evidence or a useful judgement rather than another launch summary. Keep prices, specifications and lifecycle information current, and make important facts visible in HTML rather than only in an image. Use descriptive links to primary sources, add tables where readers genuinely compare the same fields, name the editorial team, show review dates, connect the subject to the wider site with internal links, and keep structured data in agreement with the visible page. Google’s guide also describes query fan-out, where the system may run related searches across subtopics before composing an answer. That supports one thorough page covering status, price, limits, uses and migration. It does not justify a thin page for every rearrangement of the keyword. Google now points site owners to the Generative AI performance report in Search Console rather than to third-party claims about secret AI metrics. For the wider strategy, read AIWIZ’s analysis of how AI is changing Search, SEO and advertising. Using generative AI in the writing process is not itself an SEO violation. Google’s concern is scaled, unoriginal content created to manipulate search. The safer test is plain: does the page contain accurate, useful work that a reader would choose over the launch post it summarises?The verdict
Gemini 3.6 Flash looks like a worthwhile efficiency update. The output price is lower, Google reports fewer tokens and tool calls, and both the vendor benchmarks and the independent launch snapshot point towards a cheaper route through demanding work. It is not a reason to change production traffic on trust. Gemini 3.6 Flash launched on 21 July 2026, its pre-release throughput figures are not yet reflected on live measurement pages, and a single aggregate score cannot describe every workload. Test the jobs that consume staff time now, include the awkward failures, and compare cost per accepted result. For a new agentic, coding or multimodal build, 3.6 is the first Flash model to evaluate. For an existing 3.5 integration, run both in parallel and migrate when the evidence from your own tasks is better. There is no announced shutdown forcing the schedule. The search lesson is similar. As agents become better at reading, comparing and acting on web content, brands benefit from being easy to retrieve and safe to cite. Clear answers, named editorial responsibility, current facts and sound technical SEO have not become old-fashioned. They have become the admission price. AIWIZ helps organisations plan practical AI workflows and improve visibility across search and answer engines. To discuss an automation opportunity or an AI-search visibility review, contact the AIWIZ Digital Marketing team.Frequently Asked Questions
Is Gemini 3.6 Flash officially available?
Yes. Google made it generally available on 21 July 2026 through the Gemini API and Google AI Studio, and announced access through the Gemini app, Android Studio, Antigravity and enterprise products. Product, plan, regional and usage limits may differ.
What is the Gemini 3.6 Flash model ID?
The stable API model ID is gemini-3.6-flash.
Is Gemini 3.6 Flash better than Gemini 3.5 Flash?
It is the better first model to test for most new Flash workloads. It has a lower standard output price, a later knowledge cut-off and stronger Google-reported results in selected coding, computer-use and knowledge-work tests. Artificial Analysis gave both high-thinking variants an aggregate Index score of 50, so workload-specific testing remains essential.
How much does Gemini 3.6 Flash cost?
On the standard paid Gemini Developer API tier, $1.50 per million input tokens and $7.50 per million output tokens. Batch and Flex rates are $0.75 for input and $3.75 for output, and Priority is $2.70 and $13.50. Caching, storage, grounding, tool, tax and enterprise charges can change the total.
What context window does it have?
It accepts up to 1,048,576 input tokens and can produce up to 65,536 output tokens. The practical amount of material that fits depends on file processing, media resolution, tokenisation and the space needed for instructions and tool results.
Should existing 3.5 Flash users migrate now?
They should evaluate the move now, not switch blindly. Google has announced no shutdown date for 3.5 Flash. Compare both models on real tasks, including failure handling, latency, tool calls, human correction and total cost, and note the API changes: sampling parameters are deprecated and prefilled model turns are no longer supported.
Does AEO require special schema or an llms.txt file?
No. Google says there is no special AI schema, no llms.txt requirement and no extra markup route into generative answers, and it retired FAQ rich results in May 2026. Use accurate structured data that matches the visible page, keep the site technically eligible for Search, and focus on useful, non-commodity content. A plugin-generated llms.txt file is preparedness, not a requirement, and a connected schema graph with real organisation and author entities is worth more than any single markup type.