GPU Capacity Planner

FortgeschrittenopsMindestens 32K Kontext

Sizes GPU capacity for self-hosted LLM inference from a target traffic profile. Estimates KV cache and weight memory, derives concurrency limits from context length and batch size, models throughput against latency targets, and compares instance types and autoscaling policies so capacity matches demand without paying for idle accelerators.

Anwendungsfälle

  • Estimating how many GPUs a target QPS requires
  • Computing KV cache memory for a given context length and concurrency
  • Choosing batch size to balance throughput against p95 latency
  • Comparing instance types and autoscaling policies on cost

Beispiel-Prompt

Plan GPU capacity for a self-hosted inference service.

Model: 32B parameters, FP8. Traffic: 40 requests/second peak, average 8K input and 1K output
tokens. Target: p95 time-to-first-token under 800ms.

Walk through the memory math (weights plus KV cache), the concurrency each GPU can sustain, how
many GPUs I need at peak, and the batching and autoscaling configuration. Show your assumptions
so I can adjust them.

Empfohlene Modelle

Kompatible Werkzeuge

claude-codecursorkiroany

Modalitäten

Eingabe: text, code
→
Ausgabe: text, code

Ähnliche Skills

Autor

OpenModels Community

@openmodelsrun