Multimodal Prompt Design
IntermediacreativeContexto mínimo: 32K
Writes prompts that combine images, video, documents, and text effectively. Covers ordering images relative to instructions, resolution and tiling tradeoffs that affect token cost and detail, grounding questions to specific regions or pages, requesting structured output from visual input, and reducing visual hallucination on charts, tables, and scanned documents.
Casos de uso
- Extracting structured data from scanned invoices or forms
- Asking precise questions about charts and diagrams
- Ordering images and instructions for best accuracy
- Cutting token cost on high-resolution image input
Prompt de ejemplo
I need to extract line items from scanned invoices of varying quality into JSON. Design the multimodal prompt: how to order the image and instructions, what resolution to send and the cost tradeoff, how to define the output schema, how the model should signal low confidence on unreadable fields instead of guessing, and how to handle multi-page invoices. Include the prompt and a validation step for the returned JSON.
Modelos recomendados
Herramientas compatibles
claude-codekirogemini-cliany
Modalidades
Entrada: text, image, video
→Salida: text, code
Skills relacionadas
Autor
OpenModels Community