content is an array of parts rather
than a string. That is OpenAI’s shape, and it travels here unchanged, so a
request your client already builds works.
Which models take one
Every catalog entry carries avision word in its capabilities block,
answering yes, no or unknown. See
Model capabilities.
Read unknown as “nobody has established this”, not as “no”. A yes is a
recorded answer rather than a guess, which is why it is worth choosing on.
Nothing here checks the image
The request path forwards your body and does not inspect it. Consequences, all of them the same consequence seen from different sides:- An image sent to a text-only model is not refused here. It reaches the provider and the provider refuses it, in its own words and on its own status.
- Format, size and dimension limits are the provider’s, and so are the errors about them. We publish no limits of our own on an image, because we impose none.
- A URL you send is fetched by the provider, not by us. Nothing on our side downloads it, so a URL that is unreachable from the provider fails at the provider even if it resolves for you. If that is a risk, send a data URI.
- We do not resize, re-encode, or strip metadata. The bytes you put in the body are the bytes that go on.
What an image costs
Images are charged as prompt tokens, at the model’s input rate, and the count is the provider’s own. Different providers count an image differently, and the same picture can be a different number of tokens on two models. Nothing here re-counts or estimates: the figure you read back came from the provider. See Usage and metering. Two habits are worth having: A data URI is prompt tokens and body bytes both. Base64 makes a large image noticeably larger on the wire before it is a single token of anything. An image stays in the conversation. Every turn of a multi-turn chat resends the wholemessages array, so an image in turn one is paid for again in turn
five. Drop it from the history once the model has described it, if the
description is what the rest of the conversation needs.
Audio, video, and generated images
The catalog carries anaudio word alongside vision, answering the same three
ways, for models that take audio input.
Beyond that, this site documents no other modality, and the reason is the same
one that governs every page here. The operations this site documents are
POST /v1/chat/completions, GET /v1/models and GET /v1/models/{id}, so
there is no image generation endpoint, no speech endpoint, and no video endpoint
described anywhere on it. Nothing here says what a path this site does not
document answers. See
OpenAI compatibility.
Where a provider exposes a modality through the chat completions body itself,
that body travels like any other, and the model’s own capability words are what
tell you whether it will be understood.