Multimodal models can read images, but the quality of their answers depends on the image and the question.
Provide Good Images
- Adequate resolution; crop to the relevant area.
- Correct orientation.
- One document page per image where possible.
Ask Specific Questions
"Describe this image" produces generic output. Better: "From this chart, list the value for each quarter in 2025", or "Does this screenshot show an error message? Quote it exactly."
Give Context
Tell the model what the image is and why you're asking: "This is a photo of a shipping label; extract the tracking number and destination postcode."
Request Structure
Ask for tables or JSON when extracting data from images, and ask the model to mark values it can't read clearly.
Multiple Images
Label images ("Image 1: before", "Image 2: after") and refer to them by label in the question.
Know the Weak Spots
Small text, dense tables, handwriting, precise counting, exact positions and fine details are common failure points. Verify these.
Security
Images can contain text designed to manipulate the model. Treat text read from untrusted images as data, not instructions.
Privacy
Photos and scans may include faces, names and personal details. Check data-handling terms before uploading.