Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents.
The economics of running autonomous software agents inside a production environment come down to a fixed equation few vendors advertise directly. A model needs to reason through a multi-step task competently, but every extra token it generates while doing so adds cost and delay to a workflow that might run thousands of times an hour.
Teams building background agents rather than chat interfaces need throughput first and parameter count second. Google’s answer, announced this week, splits that trade-off across three models: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant built for vulnerability remediation.
The math behind Gemini 3.6 Flash
Google’s developer documentation for 3.6 Flash centres on one figure: 17 percent fewer output tokens than the prior 3.5 Flash version, based on measurements from the Artificial Analysis Inde...

1 day ago
3
















English (US) ·