Model Quantization Advisor
AvanzadaopsContexto mínimo: 32K
Advises on quantizing open-weight models for cheaper, faster inference while keeping quality acceptable. Compares formats and precisions (GGUF, AWQ, GPTQ, FP8, INT8, 4-bit), estimates memory footprint and throughput on target hardware, flags which workloads degrade most under quantization, and defines the evaluation needed to confirm the tradeoff is safe.
Casos de uso
- Choosing a quantization format for a target GPU
- Estimating VRAM footprint and throughput before deploying
- Deciding how much quality loss is acceptable per workload
- Designing an eval to verify a quantized model is still good enough
Prompt de ejemplo
I want to self-host a 70B open-weight model for an internal coding assistant on 2x A100 80GB. Recommend a quantization approach: which format and precision, expected VRAM use and tokens per second, and where quality is most likely to drop (long context, code correctness, tool calling). Define the evaluation I should run before and after to confirm the tradeoff is acceptable.
Modelos recomendados
Herramientas compatibles
claude-codecursorkiroany
Modalidades
Entrada: text, code
→Salida: text, code
Skills relacionadas
Autor
OpenModels Community