


We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.
The Gemini 3.6 Flash Family is a new lineup of AI models from Google, including Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. These models are designed to deliver the efficiency, latency, and reliability needed to build AI agents at scale. Building on the success of Gemini 3.5 Flash, the family focuses on higher token efficiency, lower latency, and more reliable performance for production-grade agentic workflows.
Gemini 3.6 Flash delivers better coding, knowledge work, and multimodal performance while reducing output token usage by 17% compared to 3.5 Flash. It achieves this with fewer reasoning steps and tool calls, making multi-step workflows more cost-effective at $1.50/1M input tokens and $7.50/1M output tokens.
This model delivers 350 output tokens per second according to the Artificial Analysis Index, making it the fastest 3.5-class model. It significantly outperforms prior Flash-Lite generations in agentic workflows while maintaining low costs.
A specialized cyber-focused model paired with the CodeMender code security agent, designed for successful cybersecurity applications that require careful orchestration of a model alongside agent infrastructure. It delivers competitive performance at the frontier of code security.
"3.6 Flash reduces output token usage by 17% compared to 3.5 Flash, all at a lower cost per output token."
This efficiency gain is remarkable because it comes alongside performance improvements across multiple benchmarks. 3.6 Flash achieves 49% vs. 37% on DeepSWE for code edits, 63.9% vs. 49.7% on MLE Bench for ML research, and 83.0% vs. 78.4% on OSWorld-Verified for computer use. The combination of lower cost, fewer tokens, and better results makes this family uniquely suited for scaling agentic workflows in production.
You're building production AI agents and need a model family that balances efficiency, latency, and reliability. The Gemini 3.6 Flash Family is especially worth exploring if you're working on coding tasks, knowledge work, multimodal applications, or cybersecurity workflows where token efficiency and cost control are critical.
Other tools you might consider
The world can't build compute fast enough to keep up with AI demand. So we took a different path. ZeroGPU is AI infrastructure powered by small language models running on a hybrid edge network reusing compute that already exists. Not every task needs a frontier model. Our purpose-built, edge-optimized models run 10x faster, 50% cheaper and offload 70–80% of production tasks to small models with frontier-level accuracy.
Run state-of-the-art open-source models (GLM 5.1, Kimi K2.7 Code, MiniMax M2.7, and more) in Claude Code at up to 4× the speed (up to 200 tok/s) for a flat $29/month. Set up in minutes, no code changes.
BaseRT is the fastest LLM runtime on Apple Silicon. Install it with one command and run local models on your own device.
GitHits gives coding agents access to the open-source code your app depends on. Get real implementation examples, dependency source navigation, package inspection and documentation. Agents can grep and read your codebase. They can't grep and read the open-source code your app depends on. That's where they start guessing, retrying, and looping. GitHits builds a version-aware index on demand. Agents can search, navigate, and inspect the code behind their dependencies. CLI: npx githits@latest init
Loading comments…
Maker
sandbyte
Project Info
Product Keywords