🟢 T5Gemma-2 is impressive return of encoder-decoder architecture.
Like the first T5Gemma, which was released in the summer, T5Gemma-2 was trained on the basis of the regular Gemma, which is a decoder-only, with a long context of up to 128K tokens and multimodality.
In the summer, Google demonstrated the main idea of adaptation: we initialize the encoder-decoder with the weights of the decoder-only and continue pretraining through UL2. Now, this approach has been transferred to multimodal and long-context, plus some architectural optimizations have been added.
As a result, this time it's not just an experiment with the architecture, but a really useful model. It has the brains of Gemma, but it handles long contexts better (+cheaper) because it has an encoder. So if you have a task like summarization or working with large documents - feel free to use it.
There are variants for 270M-270M, 1B-1B, 4B-4B (estimated total parameters ~370M / ~1.7B / ~7B). Unfortunately, Google doesn't publish the instruct, but the pretraining checkpoints are available here.
🟡 FunctionGemma is a tiny tool-caller for agents, basis for an autonomous local agent.
Generator of structured function calls with only 270M parameters. It even has a different tokenizer from the regular Gemma. It can generate text, but its main role is to call the necessary tools to perform the task. In short, something like Siri specifically for performing offline tasks on the device.
Google emphasizes that the model is designed for retraining (not prompting) on specific tasks. For example, in the case from the blog post, it was retrained quite cheaply on Mobile Actions, and the accuracy increased from 58% to 85%,
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
