Gemini 3.1 Pro: Google’s New Benchmark Leader Is Here

Gemini 3.1 Pro: Google's New Benchmark Leader Is Here

Google’s Gemini 3.1 Pro went into production on March 9, 2026, replacing Gemini 3 Pro Preview across all Google platforms. It leads 13 of 16 major AI benchmarks and costs the same as the model it replaced.

What Has Changed

Gemini 3.1 Pro is not a new architecture. Google describes it as a targeted improvement to the reasoning system already present in Gemini 3 Pro, with the same pricing and context window. What has changed is the quality of reasoning within that window.

The model accepts five input types in a single API call: text, images, audio, video, and code. Earlier Gemini versions required separate preprocessing steps for different modalities. Gemini 3.1 Pro handles all five natively.

Context Window and Output

The input context window supports up to one million tokens, equivalent to roughly 1,500 pages of text or 30,000 lines of code in a single session. Maximum output is 64,000 tokens per response. Google’s documentation confirms the model achieves near-perfect retrieval accuracy above 99% across the full context window.

Pricing

Via the API, Gemini 3.1 Pro is priced at $2.00 per million input tokens and $12.00 per million output tokens for contexts up to 200,000 tokens. Prompts exceeding 200,000 tokens are charged at $4.00 input and $18.00 output per million tokens. Google confirmed this is identical to Gemini 3 Pro Preview pricing, making it a free capability upgrade for existing users.

Where to Access It

The model is available now via the Gemini API in Google AI Studio (model ID: gemini-3.1-pro-preview), Vertex AI, the Gemini CLI, and NotebookLM. Consumer access is rolling out in the Gemini app for Google AI Pro and Ultra subscribers. The Gemini 3 Pro Preview model was shut down on March 9 and is no longer available.

Benchmarks

Gemini 3.1 Pro leads 13 of 16 major benchmark categories tracked by Artificial Analysis, tying GPT-5.4 Pro on the overall Intelligence Index at approximately one third of the API cost. On GPQA Diamond, a graduate-level science reasoning test, the model scores 94.3%.

Discover more from fastai.news

Subscribe now to keep reading and get access to the full archive.

Continue reading