Images & video

What is Text-to-image?

Creating a picture from a written description. The core feature of ChatGPT images, Gemini, Midjourney and Flux.

You describe the image and the model generates it from scratch. The more specific you are about subject, clothing, setting, lighting, camera angle and style, the more control you have.

Text-to-image models are excellent at mood, style and composition, and keep improving at hands, text in images and consistent faces — the three classic weak spots.

Each tool has its own taste. Some lean photorealistic, some lean artistic, and some follow long prompts more literally than others. When a result is close but not right, change one detail at a time — the light, the lens, the outfit — instead of rewriting the whole prompt, so you learn which words actually move the image.

Prompt skeleton

Subject + outfit/details + setting + lighting + camera/lens + style + aspect ratio. Example: “Bride in red Banarasi saree, temple courtyard, soft morning light, 50mm portrait, cinematic colour grade, 4:5.”

Best portrait prompts

Related terms

In the news