On May 19, Sundar Pichai stood on stage and said, "Give us until next month." The crowd groaned. Sixty-two days later, Gemini 3.5 Pro still hasn't shipped.

30-second summary
I/O promise: "next month" June slips July 17 slips too full base-model retrain 4 senior researchers gone + $225B wiped out so what do you actually do now

Here's what Google promised on stage

At Google I/O on May 19, Google unveiled Gemini 3.5 Pro: a 2-million-token context window, a "Deep Think" reasoning mode for hard problems, and frontier-level multimodal understanding — a configuration built to absorb everything Google used to route to its top-tier "Ultra" model.

Pichai told the crowd to wait until next month, and developers groaned right there in the room. That reaction alone tells you expectations were already worn thin — Google had already slipped Gemini Ultra 1.5 by three months earlier this year.

The lighter sibling model, Gemini 3.5 Flash, actually shipped that same day — with real, published benchmarks. Per Google's own launch post, it scored 76.2% on Terminal-Bench 2.1 (versus 70.3% for the prior 3.1 Pro) and 83.6% on MCP Atlas (versus 78.2%). Flash was real. Pro was the problem.

And that promise? It's now missed three times

June came and went with nothing. Google restated July 17 — but even that wasn't an official announcement, just leaked reporting. As of July 13, there was still no model card, no pricing page, no gemini-3.5-pro listing anywhere in the public API docs. Then mid-July arrived, and July 17 slipped too. According to Geeky Gadgets, the rebuilt model is still plagued by hallucinations and inconsistency that keep it from clearing basic reliability bars.

This wasn't a routine tuning delay. HackerNoon and Geeky Gadgets, citing unnamed internal sources, reported that Google DeepMind scrapped its near-complete base model entirely and restarted pre-training from scratch on a native Gemini 3 foundation. Pre-training is the most expensive phase of building a frontier model, and it sets the capability ceiling that fine-tuning can't raise. Choosing to redo it means Google's own engineers concluded the gap wasn't cosmetic — it was structural.

Specifically: the scrapped model couldn't hold structural consistency in complex, multi-layered SVG scenes, and it broke down in recursive tool-calling chains — the multi-step sequences where an agent calls one tool, feeds the result into another, and so on. Recursive tool-call stability is the defining requirement for an agentic coding model, which is exactly the use case Google staked the entire 3.5 generation on. Flash already proved the category was achievable; a Pro model that regressed on the same tasks would have been an embarrassment, not a flagship.

The schedule wasn't the only thing that broke. Between June 18 and 24, four senior DeepMind researchers left within the same ten days — Gemini co-lead Noam Shazeer went to OpenAI, Nobel laureate John Jumper went to Anthropic, and Jonas Adler and Alexander Pritzel also joined Anthropic that same week. That same week, the market erased roughly $225 billion from Alphabet's market cap — its steepest single-day drop in over a year.

Gemini 3.5 Pro (roadmap)What you can actually use now
StatusMissed June, then July 17Already shipped, benchmarked
Official docsNo model card, no pricing pagePublic API docs and pricing
Performance claimsKeynote demos + anonymous leaksVerifiable by independent testers
CompetitionGPT-5.6 Sol, Grok 4.5 already public since July 9You can test and compare today

While Google restarted pre-training, rival models shipped into the vacuum. OpenAI's GPT-5.6 Sol and xAI's Grok 4.5 both went public on July 9. The pricing rumor floating around is roughly 10x Flash's rate — about $15 per million input tokens and $60 per million output tokens, with Deep Think reportedly gated behind the $250/month Ultra tier. All of it still carries the "rumor" tag.

This isn't just Google's problem

If you're a founder, marketer, or developer building workflows around AI, this story matters beyond the headlines. Anyone who watched the May keynote and thought "I'll build my long-document workflow once the 2-million-token context ships" has built exactly nothing in the 62 days since.

The core lesson: benchmark numbers from a keynote demo and a shipped, independently-tested product are not the same thing. No model card and no pricing page means nobody has officially stood behind those numbers yet. And this isn't Google's first time — Gemini Ultra 1.5 slipped the same way, by three months. Frontier model delays aren't rare on their own, but a full base-model retrain colliding with a simultaneous exodus of senior researchers is what makes this one different.

How to filter AI roadmap promises from now on

  1. Separate "confirmed" from "rumored"
    Only numbers on an official blog, model card, or pricing page count as confirmed. Keynote demos and "sources say" reporting are rumors until proven otherwise.
  2. Check the slip history with one search
    Find out how many times this model has already been delayed. Two or more slips means you should assume another one by default.
  3. Trace the benchmark source
    Is the number self-reported by the company, or verified by an independent evaluator or third-party tester? Self-reported-only numbers may not hold up in practice.
  4. Price out the cost of waiting
    Can you handle the work with what's already available, or do you genuinely need a specific spec the new model promises (like a 2M context window)? Most of the time, it's the former.
  5. Start with what's shipped today
    Design your workflow so swapping in the new model later is just an API endpoint change — don't freeze your roadmap waiting on someone else's.

Go deeper

The DeepMind talent exodus, timelined A day-by-day breakdown of the ten days that shook Google's AI team the-agent-report.com

Google's "worst fortnight," in full A comprehensive look at the $225B market-cap hit and the researcher departures tech-insider.org

Why Google restarted pre-training The specific SVG-generation and recursive tool-calling failures behind the rebuild hackernoon.com

Confirmed specs vs. rumored specs A breakdown of the 2-million-token context window and pricing estimates blog.getbind.co

Why mid-July got so crowded with AI launches How GPT-5.6, Grok 4.5, and others ended up shipping into the same week swisherpost.com

What the rebuild is actually targeting Math, SVG generation, image quality, and the side models riding along finance.biggo.com