Workbench / Online
BUILD YOUR COMPUTE STACK
Start with the model. MuseBoard maps the memory and hardware. Settings are mirrored to the URL, so any configuration can be shared.
Compute workbench
B / Estimated VRAM
Inference227.5GB
Serving: weights + KV cache + runtime overhead. Estimate — not a guarantee.
- Model weights202.9 GB
- KV cache4.2 GB
- Runtime overhead20.3 GB
Memory headroom
Free after estimated load
Utilization
Target ≤ 90% of 320.0 GB
Decode ceiling
Theoretical, batch 1, bandwidth-bound
VRAM usage
227.5 GB / 320.0 GB
Pinned configuration
8 × NVIDIA A100 40GB
- Total VRAM
- 320.0 GB
- Topology
- Single node · TP
- Interconnect
- NVLink 3 · 600 GB/s
- Board power
- 3,200 W
Compatibility / 8 × A100 40GB
- Ready
Inference
227.5 GB · 71% of 8 × A100 40GB
- Limited
Fine-tuning
305.8 GB · Needs 16 × A100 40GB
- Cluster
Training
7,191 GB · Needs 200 × A100 40GB (25 nodes)
Alternative configurations
- 01SERVE / READY
Inference
Estimate memory requirements for serving models.
weights + KV cache + overhead
Open in workbench - 02ADAPT / LORA
Fine-tuning
Explore memory requirements for adapting existing models.
frozen base + adapters + activations
Open in workbench - 03TRAIN / ADAM
Training
Understand large-scale compute requirements.
weights + grads + optimizer + activations
Open in workbench
MuseBoard gives first-order estimates for planning conversations. It does not replace profiling, capacity testing or detailed ML infrastructure planning.