Send uploaded images to a vision-capable model, extract the details you need as structured data, and fill your forms automatically.
Users upload a photo of a receipt. Have AI read it and fill in the form
How it works
- Store the upload: Vovy asks Claude Code to save uploads to a private Supabase Storage bucket with a policy so users only see their own images.
- Resize before sending: Claude Code shrinks images to a sensible size before sending. Huge phone photos cost more tokens without better results.
- Ask the model to read it: A server function sends the image and a prompt like "extract merchant, date, total" to a vision model, asking for JSON back.
- Fill the form for review: Claude Code fills your form fields with the extracted values and highlights them, so users check before saving.
- Test on messy photos: Vovy tries blurry, rotated and crumpled examples and shows a table of which fields came out right, plus cost per image.
What you provide
- Sample images your users will upload
- An AI API key
- A Supabase project
What you get
- Private image uploads
- AI extraction into your form
- An accuracy and cost table
FAQ
Is this better than OCR?
For messy real-world photos, usually yes, since the model understands layout and context, not just letters.
Is it accurate enough to trust blindly?
Not for money or legal data. Let users confirm the fields, which this setup does.
What does each image cost?
Typically a fraction of a cent to a cent or two, depending on image size and model.
Related tasks
All tasks