Four names, four jobs. The exam asks which one suits a described situation, so the useful thing is what separates them rather than what they have in common.
Gemini
Google's flagship family, multimodal from the ground up rather than a text model with image handling added. It reads and reasons over text, images, audio and video, and it is what sits behind the Gemini app, Gemini for Google Workspace and Gemini Enterprise.
It comes in several sizes, which is the part that matters commercially. A larger model reasons better over hard, open ended work. A smaller one costs and weighs far less and is frequently just as good on a narrow task. A question that mentions high volume, tight latency or cost pressure is usually asking you to move down the range rather than to a different product.
Gemma
A family of open models. Google publishes the weights, so you can download them, run them on your own infrastructure or another cloud, inspect them and adapt them.
The reasons to choose Gemma are sovereignty, cost at scale, and the ability to run somewhere a managed API cannot reach, such as inside a restricted network or on a device. The cost is that hosting, scaling, updating and securing the model become yours.
This open against managed distinction is the one the exam returns to. A question about a regulated environment where data cannot leave, or about running on hardware you control, is pointing at Gemma.
Imagen
Image generation and editing. Producing visuals from a text description, editing an existing image, and generating variations.
The business cases are marketing and creative production, product imagery, and anywhere a team is currently commissioning or licensing stock images.
Veo
Video generation. Producing video from a text description or from an image, at higher fidelity and longer duration than earlier generations of the technology.
The cases are short form marketing, product demonstration and prototyping a concept before commissioning a production shoot.
Reading a question
The exam describes a business situation and expects one name.
Reasoning, summarising, extracting, answering, or anything mixing modalities is Gemini. Needing the weights, self hosting, or a restricted environment is Gemma. Still images are Imagen. Moving images are Veo.
Where two seem to fit, the deciding factor is usually stated in the question and is rarely capability. It is cost, latency, where inference may happen, or whether the organisation needs to hold the model itself.