GPU Capacity Planner
FortgeschrittenopsMindestens 32K Kontext
Sizes GPU capacity for self-hosted LLM inference from a target traffic profile. Estimates KV cache and weight memory, derives concurrency limits from context length and batch size, models throughput against latency targets, and compares instance types and autoscaling policies so capacity matches demand without paying for idle accelerators.
Anwendungsfälle
- Estimating how many GPUs a target QPS requires
- Computing KV cache memory for a given context length and concurrency
- Choosing batch size to balance throughput against p95 latency
- Comparing instance types and autoscaling policies on cost
Beispiel-Prompt
Plan GPU capacity for a self-hosted inference service. Model: 32B parameters, FP8. Traffic: 40 requests/second peak, average 8K input and 1K output tokens. Target: p95 time-to-first-token under 800ms. Walk through the memory math (weights plus KV cache), the concurrency each GPU can sustain, how many GPUs I need at peak, and the batching and autoscaling configuration. Show your assumptions so I can adjust them.
Empfohlene Modelle
Kompatible Werkzeuge
claude-codecursorkiroany
Modalitäten
Eingabe: text, code
→Ausgabe: text, code
Ähnliche Skills
Autor
OpenModels Community