# Google launches Gemini 3 Flash API and open‑source Gemma 4 12B model

**Published:** 2026-06-29T15:30:24.206Z  
**Topic:** Google Ai  
**Sentiment:** neutral  
**Publisher:** TrendWatcher — https://www.trendwatcher.in/article/987d6cf6-5aab-4311-9cd5-cf2546a2de7c

Google rolls out Gemini 3 Flash API at $0.50/1M input tokens and unveils Gemma 4 12B, an 11.95B‑parameter model that runs on a 16 GB laptop.

Google made Gemini 3 Flash the default model in its Gemini app and opened the model to developers via a preview API, pricing it at $0.50 per 1 million input tokens and $3.00 per 1 million output tokens [1]. The move follows the launch of Gemma 4 12B, an 11.95‑billion‑parameter open‑weights model that can run locally on a typical 16 GB enterprise laptop [4].

| At a glance | |
|---|---|
| Model | Gemini 3 Flash (API) |
| Price | $0.50 / 1 M input tokens; $3.00 / 1 M output tokens |
| Benchmark | 33.7% on Humanity’s Last Exam (no tools) |
| Open model | Gemma 4 12B (11.95 B params) |

## Gemini 3 Flash: performance and rollout  
Gemini 3 Flash improves on its predecessor Gemini 2.5 Flash, scoring 33.7% on the Humanity’s Last Exam benchmark—up from 11% for Gemini 2.5 Flash and just 0.8 points shy of GPT‑5.2’s 34.5% [1]. On the multimodal MMMU‑Pro benchmark it posted an 81.2% score, outpacing all listed rivals [1]. Google positions the model as a “workhorse” for bulk tasks, noting it uses 30% fewer tokens than Gemini 2.5 Pro on thinking‑type workloads [1]. The model is already integrated into products like the Gemini app, Vertex AI, and the Antigravity coding tool, and is being used by companies such as JetBrains and Figma [1].

## Gemma 4 12B: architecture and edge potential  
Gemma 4 12B is released under an Apache 2.0 license, enabling free download and local deployment on consumer hardware via Hugging Face or Kaggle [4]. Its “Unified” architecture removes separate audio and vision encoders, replacing them with lightweight linear layers that cut VRAM needs to 16 GB and lower inference latency [4]. The model supports a 256 K token context window and native tool‑use, bridging the gap between edge devices and data‑center‑scale models [4]. Google also introduced Multi‑Token Prediction (MTP) drafter models for Gemma, using speculative decoding to accelerate token generation up to three‑fold compared with standard autoregressive sampling [2].

## Competitive landscape  
Gemini 3 Flash’s benchmark scores place it within striking distance of frontier models like Gemini 3 Pro and GPT‑5.2, while its pricing remains higher than Gemini 2.5 Flash but justified by speed gains of three‑times faster inference [1]. Gemma 4 12B’s ability to run on a laptop contrasts with larger, cloud‑only offerings from OpenAI and Anthropic, offering enterprises a privacy‑preserving alternative for multimodal tasks. The introduction of MTP drafter models further narrows the performance gap between edge and data‑center hardware [2].

## What to watch  
- **API rollout milestones** – Google will track adoption of the Gemini 3 Flash preview API and may adjust pricing based on usage patterns.  
- **Edge adoption of Gemma 4** – Monitor downloads and integration of Gemma 4 12B on platforms like Hugging Face and its performance on consumer GPUs.  
- **Competitive responses** – Expect OpenAI and other AI firms to announce faster or cheaper models to counter Google’s speed‑focused releases.

Google’s dual push—commercializing a faster, cheaper Gemini 3 Flash for bulk cloud workloads while opening a highly efficient Gemma 4 model for edge deployment—highlights a strategy to dominate both enterprise AI services and the emerging local‑model market. The next test will be whether developers and enterprises adopt the new pricing and hardware‑light architecture at scale.

## Sources
1. TechCrunch — [Google launches Gemini 3 Flash, makes it the default model in the Gemini app](https://techcrunch.com/2025/12/17/google-launches-gemini-3-flash-makes-it-the-default-model-in-the-gemini-app/)
2. Ars Technica — [Google’s Gemma 4 AI models get 3x speed boost by predicting future tokens](https://arstechnica.com/ai/2026/05/googles-gemma-4-open-ai-models-use-speculative-decoding-to-get-up-to-3x-faster/)
3. MSN — [Gemma 4’s 31B model ranks third among all open AI models on the Arena AI leaderboard](https://www.msn.com/en-us/news/technology/gemma-4-s-31b-model-ranks-third-among-all-open-ai-models-on-the-arena-ai-leaderboard/ar-AA22izrM?ocid=BingNewsVerp)
4. VentureBeat — [Google's new open source Gemma 4 12B analyzes audio, video — and runs entirely locally on a typical 16GB enterprise laptop](https://venturebeat.com/technology/googles-new-open-source-gemma-4-12b-analyzes-audio-video-and-runs-entirely-locally-on-a-typical-16gb-enterprise-laptop)

---
Cite as: TrendWatcher, "Google launches Gemini 3 Flash API and open‑source Gemma 4 12B model", https://www.trendwatcher.in/article/987d6cf6-5aab-4311-9cd5-cf2546a2de7c
