Multimodal Prompt Design

IntermediacreativeContexto mínimo: 32K

Writes prompts that combine images, video, documents, and text effectively. Covers ordering images relative to instructions, resolution and tiling tradeoffs that affect token cost and detail, grounding questions to specific regions or pages, requesting structured output from visual input, and reducing visual hallucination on charts, tables, and scanned documents.

Casos de uso

  • Extracting structured data from scanned invoices or forms
  • Asking precise questions about charts and diagrams
  • Ordering images and instructions for best accuracy
  • Cutting token cost on high-resolution image input

Prompt de ejemplo

I need to extract line items from scanned invoices of varying quality into JSON.

Design the multimodal prompt: how to order the image and instructions, what resolution to send
and the cost tradeoff, how to define the output schema, how the model should signal low
confidence on unreadable fields instead of guessing, and how to handle multi-page invoices.
Include the prompt and a validation step for the returned JSON.

Modelos recomendados

Herramientas compatibles

claude-codekirogemini-cliany

Modalidades

Entrada: text, image, video
→
Salida: text, code

Skills relacionadas

Autor

OpenModels Community

@openmodelsrun