AI Image Description Generator
Generate product descriptions, captions, or accessibility text from photos. Export as CSV.
Description length
Drop files here
or click to browse · paste from clipboard
Accepts .JPG, .JPEG, .PNG, .WEBP, .BMP, .TIFF, .HEIC · Up to 1,000 files
Images stay on your device. Verify in DevTools → Network.
Frequently asked questions
No. The Florence-2 model runs locally in your browser via transformers.js. Your images are never sent to a server — all processing happens on your device.
Alt text is short (50–100 characters) and optimised for screen readers and SEO. Image descriptions are longer (1–3 sentences), human-readable, and suited for product listings, social captions, and accessibility documentation.
Use Detailed length. It generates 2–3 sentences describing the product, colour, texture, and context — a solid starting point to edit before publishing.
Very dark images, images dominated by text (infographics, screenshots), abstract art, and images with multiple equal-prominence subjects tend to produce generic or inaccurate descriptions. The model describes visual content, not meaning — it cannot infer brand context.
Up to 200 images recommended. Florence-2 is memory-efficient (~260 MB). Processing is sequential — one image at a time — at roughly 3–10 seconds per image on a modern laptop.
Yes. Click any description to edit it inline. The CSV download uses your edited text, not the original AI output.