Gemini 4 Launches to Fanfare but Internal Staff Question Its Real-World Performance

Deep News
Yesterday

Google officially released Gemini 4 Argon to external users on Wednesday, positioning it as the company's long-developed flagship artificial intelligence model.

However, despite the model achieving leading scores across multiple benchmark tests, some Google insiders with direct access to the model have expressed reservations about its actual performance, particularly in coding capabilities, according to Bloomberg.

Google has firmly denied these concerns, stating that claims about Gemini 4 underperforming in coding and other areas are "inaccurate," and cited comments from Google DeepMind head Koray Kavukcuoglu as supporting evidence. Speaking at a conference hosted by The Information last week, Kavukcuoglu said: "I have full confidence in the team. In my view, it is a foregone conclusion that we remain at the frontier."

Facing competitive pressure from OpenAI and Anthropic advancing on both model and product fronts, whether Gemini 4 can deliver on its promises in real-world scenarios has direct implications for Google's search, advertising, and cloud businesses. Following the news, Alphabet (NASDAQ: GOOGL) saw its after-hours gains narrow significantly.

Internal divisions: Benchmark success masks real-world shortcomings

According to Bloomberg, citing people familiar with the matter, Gemini 4 performs well on industry-standard benchmarks but shows a noticeable drop-off in actual employee usage, particularly exposing clear weaknesses in coding tasks.

The report cited one of the people as saying that the model lacks capability in front-end design, which directly determines the visual presentation and interactive experience of applications and websites.

Some employees believe that Anthropic's Fable and OpenAI's Astra have surpassed Gemini in the pace of capability iteration, and that even at its best, Gemini 4 will lag behind in several areas.

Another group of employees believes the new model has already caught up with the industry's top standard. A Google employee familiar with model development told Bloomberg that there is "broad consensus" internally that Gemini 4 is at the frontier level, and denied that it has difficulties with real coding tasks.

Additionally, the report cited people familiar with the matter as saying that Gemini 4 is massive in scale. Large-scale models typically mean higher inference costs, which will place additional pressure on Alphabet's profit margins amid the company's continued increase in capital expenditure.

"Over-optimization" or genuine capability?

Industry insiders attribute this phenomenon to an industry malaise known as "Benchmaxxing" — where engineers focus more on achieving high scores on standardized tests rather than building genuinely useful products.

According to Bloomberg, two people familiar with Gemini 4 said the model appears to be affected by this problem. Edwin Chen, founder of AI startup Surge AI, was blunt about it: "To use an analogy, it's like saying 'my kid got a high SAT score' — but SAT scores don't translate into real-world performance. This is an extremely harmful problem."

AI labs widely face this dilemma because customers often use benchmark scores as the primary basis for evaluating model quality, which objectively pushes labs to tilt their efforts toward "leaderboard climbing."

Notably, Gemini 4 is not without advantages. According to one person familiar with the matter, the model performs outstandingly on multimodal tasks, such as extracting metadata from videos, and is also competitive in security and cybersecurity domains.

Google claims it surpassed OpenAI's Astra model on a benchmark measuring safety skills. Additionally, the model can generate up to 1 million tokens per single pass, equivalent to approximately 750,000 English words.

A costly detour: The 3.5 Pro setback and talent drain

Behind the release of Gemini 4 lies a costly research and development detour.

Google originally planned to launch Gemini 3.5 Pro in June this year, following the I/O conference in May, but that version was ultimately abandoned after failing to meet internal targets.

Bloomberg Intelligence analyst Mandeep Singh estimates that the cost of a single training run for a frontier model can reach up to $400 million, and when combined with the labor costs of highly paid AI researchers, actual spending would be even more substantial.

Meanwhile, Google has also suffered continuous talent losses. Multiple star AI researchers have departed successively, including legendary engineer Jeff Dean, Nobel laureate John Jumper, and Noam Shazeer, who made important contributions to the foundational technology of the current AI wave.

In August this year, Demis Hassabis, who had long led Google's AI research, transitioned to chairman, handing over day-to-day operations of DeepMind to Kavukcuoglu.

In November last year, Google released Gemini 3, which was widely praised and generally regarded as an important turning point in the company's AI competitiveness. Since then, a series of stalled and repeated developments have raised external questions about its execution capability.

Narrowing competitive window: Rivals closing in

The external pressure facing Gemini 4 should not be underestimated either.

Google's most direct risk is that the Gemini series models serve as the technological foundation for core products including Search, Maps, Gmail, and Chrome, each with a user base exceeding 1 billion.

Once the model's competitiveness falls behind, competitors will gain more time to persuade developers, enterprises, and consumers to choose their platforms as the foundation for future search and software.

At the same time, OpenAI and Anthropic are accelerating their transformation from model providers to product companies, including advancing AI coding agents and other applications. According to reports, Meta released its AI agent platform Muse earlier this month, which quickly topped app download charts after launch.

Google stated that despite its previous Pro version being released in February this year, the company's AI products have continued to grow, with users of enterprise Gemini, consumer chatbot applications, and Google Search AI mode all expanding, with the latter two surpassing 1 billion users combined.

Google emphasized that its massive user distribution system is a significant advantage over some competitors.

However, whether this advantage can translate into sustained leadership in the AI era still depends on whether the Gemini series can convince users in real-world application scenarios.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10