Open-Weight Model Deployment
AvanzadaopsContexto mínimo: 32K
Designs a production deployment for an open-weight language or multimodal model from its model card, checkpoint layout, target workload, and available hardware. Verifies licensing and runtime support, estimates weights and KV-cache memory, selects a serving framework and parallelism plan, defines quantization and long-context settings, and creates quality, latency, security, and rollback gates before traffic is moved to the new model.
Casos de uso
- Turning a newly released checkpoint and model card into a deployment plan
- Choosing between vLLM, SGLang, and another model-supported serving runtime
- Estimating weight, KV-cache, and multimodal encoder memory before provisioning GPUs
- Defining canary, quality, latency, safety, and rollback gates for a model upgrade
- Separating verified model capabilities from deployment assumptions that require testing
Prompt de ejemplo
Design a production deployment for this open-weight model. Model card or repository: [URL] Hardware: [GPU type and count] Traffic: [requests per second, input/output token distribution, image or video usage] SLOs: [time to first token, throughput, availability] Verify the license, supported runtimes, native and extended context limits, checkpoint precision, and multimodal requirements. Then propose the serving framework, tensor/pipeline parallelism, quantization, memory budget, batching, autoscaling, observability, canary evaluation, and rollback plan. Clearly label any value that must be benchmarked instead of treating it as known.
Modelos recomendados
Herramientas compatibles
claude-codecursoropencodekiroany
Modalidades
Entrada: text, code, file
→Salida: text, code
Skills relacionadas
Autor
OpenModels Community