Google Rolls Out Gemini, the Most Powerful AI Model
- The Guardian
For now, Google is trying to create most of the underlying technology behind the current AI boom, having called itself an "AI-first" organization for nearly a decade, and was so surprised by how good ChatGPT was and how quickly OpenAI's technology took over the industry, Google is finally ready to fight back.
"We've done a very deep analysis of these two systems together, and done benchmark testing," Hassabis said.
Google now runs 32 existing benchmarks to compare the two models, ranging from overall tests such as Multi-task Language Understanding to tests that compare the ability of the two models to generate Python code.
"I think we're far ahead in 30 of those 32 benchmarks," Hassabis said.
"Some of them are very narrow. Some of them are bigger," he added.
In those benchmarks (most of which are very close), Gemini's biggest advantage comes from its ability to understand and interact with video and audio.
This is very deliberate, multimodality has been part of Gemini's plan from the start. Google didn't train separate models for image and sound, like OpenAI did for DALL-E and Whisper, they built one multisensor model from scratch.
"We've always been interested in very general systems," says Hassabis.
He was particularly interested in how to combine all those modes to collect as much data as possible from a number of inputs and senses, and then provide a response with the same variety.
Currently, the basic Gemini model is text input and text output, but more advanced models like Gemini Ultra can work with images, video and audio.
