Every model costs a fixed number of credits per image, per second of video, per job, or per 1,000 characters of speech, and every model draws from the same balance. Picking well is about matching the model to the job and not paying for more than you need.

Start with the tool's default

On a tool page the default model is already chosen for that job, and many tools narrow the model list to ones that suit it. Change it only when the result is not what you want. In the studio, the model list shows only models that support the mode you picked.

Photos of people

For headshots, photoshoots and edits of your own photo, you want a model that keeps the face the same. Nano Banana is the default on most photo tools.

Cheap drafts from text

When you are exploring ideas, use a low-cost text to image model and move to a stronger one for the final version.

Bringing a photo to life

For image to video from a photo of a person, Kling 3.0 Standard is the default on kiss, hug and old photo tools.

Video from a prompt, with sound

For text to video with sound, Veo 3.1 Fast is the default on the text to video and UGC ad tools. Video is priced per second, so the length you pick sets most of the cost.

Voice and music

For voiceovers and narration, use a text to speech model; for instrumental tracks, a music model.

Compare before you spend

  • The Generate button shows the exact cost for your settings before anything is spent.
  • Each model page shows its price, inputs, aspect ratios and lengths.
  • The pricing page shows roughly how many results each plan buys on common models.

If a model fails or its provider's filter rejects the request, see Why was my generation blocked. Failed generations are refunded automatically.