Skip to content
MUSEBOARD

AI Compute Workbench / 01

MAP THE COMPUTEBEHIND THE MODEL.

Plan model memory, compare GPU configurations, and understand the infrastructure required to run modern AI workloads.

Model presets
24
GPU profiles
15
Estimates
Live
NODE_01MUSEBOARD / REV.A
MuseBoard illustration: a mechanical keyboard on an exposed circuit board with the M, U, S and E keys highlighted in blue.
Model input70.6B · INT4Memory map41.5 GBGPU match1 × RTX 6000 AdaCompute ready86% util.
SYS / COMPUTEX 0.00 · Y 0.00

02Compute workbench

BUILD YOUR COMPUTE STACK

Start with the model. MuseBoard maps the memory and hardware.

Full-screen workbench
Workbench / Online
Model / Llama 70BUnits / GB (10⁹ B)

A / Model configuration

Meta
B

Editing the count switches to a custom model with an estimated architecture.

Precision0.5 B / param
Context length8,192 tokens
Batch sizeConcurrent sequences
Workload
Best fit

Architecture / published config

Layers
80
Hidden
8,192
KV heads
8 × 128
KV / token
320 KB

B / Estimated VRAM

Inference

41.5GB

Serving: weights + KV cache + runtime overhead. Estimate — not a guarantee.

  • Model weights35.3 GB
  • KV cache2.7 GB
  • Runtime overhead3.5 GB

Memory headroom

6.5 GB

Free after estimated load

Utilization

86%

Target ≤ 90% of 48.0 GB

Decode ceiling

~27 tok/s

Theoretical, batch 1, bandwidth-bound

VRAM usage

41.5 GB / 48.0 GB

0 GB48 GB

Recommended configuration

1 × NVIDIA RTX 6000 Ada 48GB

Total VRAM
48.0 GB
Topology
Single GPU
Interconnect
PCIe 4.0
Board power
300 W

Compatibility / 1 × RTX 6000 Ada

  • Inference

    41.5 GB · 86% of 1 × RTX 6000 Ada

    Ready
  • Fine-tuning

    64.0 GB · Needs 2 × RTX 6000 Ada

    Limited
  • Training

    1,262 GB · Needs 32 GPUs — exceeds one workstation node

    Cluster

Alternative configurations

03Model memory explorer

MEMORY BEFORE HARDWARE.

Slide from 1B to 405B parameters and watch the footprint cross each GPU capacity tier.

Memory / Explorer
MEM / 140GB
70B
1Blog scale405B
Precision

Weights only

140.0GB

70B × 2 B = 140.0 GB

+10% runtime → 154.0 GB

Memory footprint

1 segment = 4 GB

0 GBSingle-GPU scale (GB)256 GB

GPU fit

3/15 fit on one GPU · ≥10% headroom

  • L424 GB
  • A1024 GB
  • RTX 309024 GB
  • RTX 409024 GB
  • RTX 509032 GB
  • A100 40GB40 GB
  • L40S48 GB
  • RTX 6000 Ada48 GB
  • A100 80GB80 GB
  • H100 SXM80 GB
  • H100 NVL94 GB
  • H200141 GB
  • MI300X192 GB1× FIT
  • B200192 GB1× FIT
  • MI325X256 GB1× FIT

04GPU explorer

KNOW YOUR HARDWARE.

Capacity, bandwidth and class for the accelerators most AI teams actually plan around.

Class
Sort

Showing 15 GPUs

  • NVIDIAGPU_04

    L4

    Ada Lovelace

    24GBGDDR6

    Memory
    24GB
    Bandwidth
    300 GB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_05

    A10

    Ampere

    24GBGDDR6

    Memory
    24GB
    Bandwidth
    600 GB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_01

    RTX 3090

    Ampere

    24GBGDDR6X

    Memory
    24GB
    Bandwidth
    936 GB/s
    Class
    Consumer
    View configuration
  • NVIDIAGPU_02

    RTX 4090

    Ada Lovelace

    24GBGDDR6X

    Memory
    24GB
    Bandwidth
    1.01 TB/s
    Class
    Consumer
    View configuration
  • NVIDIAGPU_03

    RTX 5090

    Blackwell

    32GBGDDR7

    Memory
    32GB
    Bandwidth
    1.79 TB/s
    Class
    Consumer
    View configuration
  • NVIDIAGPU_08

    A100 40GB

    Ampere

    40GBHBM2

    Memory
    40GB
    Bandwidth
    1.55 TB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_07

    L40S

    Ada Lovelace

    48GBGDDR6 ECC

    Memory
    48GB
    Bandwidth
    864 GB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_06

    RTX 6000 Ada

    Ada Lovelace

    48GBGDDR6 ECC

    Memory
    48GB
    Bandwidth
    960 GB/s
    Class
    Workstation
    View configuration
  • NVIDIAGPU_09

    A100 80GB

    Ampere

    80GBHBM2e

    Memory
    80GB
    Bandwidth
    2.04 TB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_10

    H100 SXM

    Hopper

    80GBHBM3

    Memory
    80GB
    Bandwidth
    3.35 TB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_11

    H100 NVL

    Hopper

    94GBHBM3

    Memory
    94GB
    Bandwidth
    3.9 TB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_12

    H200

    Hopper

    141GBHBM3e

    Memory
    141GB
    Bandwidth
    4.8 TB/s
    Class
    Datacenter
    View configuration
  • AMDGPU_14

    MI300X

    CDNA 3

    192GBHBM3

    Memory
    192GB
    Bandwidth
    5.3 TB/s
    Class
    Datacenter
    View configuration
  • NVIDIAGPU_13

    B200

    Blackwell

    192GBHBM3e

    Memory
    192GB
    Bandwidth
    8 TB/s
    Class
    Datacenter
    View configuration
  • AMDGPU_15

    MI325X

    CDNA 3

    256GBHBM3e

    Memory
    256GB
    Bandwidth
    6 TB/s
    Class
    Datacenter
    View configuration

05GPU compare

COMPARE COMPUTE.

Put up to three GPUs side by side and check how a real model fits on each.

Compare / Up to 3 GPUs
3 selected
Precision
Workload

Model fit uses an estimated 158.0 GB at 8K context, batch 1. Blue marks the strongest value in a row — suitability depends on your workload, so there is no universal “best GPU”.

NVIDIA

RTX 4090
VRAM
24 GB
Memory type
GDDR6X
Memory bandwidth
1.01 TB/s
GPU class
Consumer
Architecture
Ada Lovelace
Workload type
InferenceFine-tuningTraining
Model fit · Llama 70B FP16
8 × RTX 409034.0 GB headroom
Power
450 W
Interconnect
PCIe 4.0
Typical use
Developer workstation inference and small-model fine-tuning

NVIDIA

H100 SXM
VRAM
80 GB
Memory type
HBM3
Memory bandwidth
3.35 TB/s
GPU class
Datacenter
Architecture
Hopper
Workload type
InferenceFine-tuningTraining
Model fit · Llama 70B FP16
4 × H100 SXM162.0 GB headroom
Power
700 W
Interconnect
NVLink 4 · 900 GB/s
Typical use
Large-scale training and high-throughput serving

NVIDIA

H200
VRAM
141 GB
Memory type
HBM3e
Memory bandwidth
4.8 TB/s
GPU class
Datacenter
Architecture
Hopper
Workload type
InferenceFine-tuningTraining
Model fit · Llama 70B FP16
2 × H200124.0 GB headroom
Power
700 W
Interconnect
NVLink 4 · 900 GB/s
Typical use
Memory-heavy inference and long-context workloads

06Compute matrix

MODEL × HARDWARE

Reference footprints across precisions, with the smallest configuration that keeps 10% headroom.

Matrix / Model × Hardware
CTX / 8K · BATCH / 1

Estimated inference VRAM (weights + KV cache + overhead) at 8K context, batch 1.

Suggest for
  • Llama 3.1 8B

    8.03B
    FP16
    18.7 GB
    INT8
    9.9 GB
    INT4
    5.5 GB

    GPU / 1 × RTX 4090

    Workbench →
  • Llama 3.1 70B

    70.6B
    FP16
    158.0 GB
    INT8
    80.3 GB
    INT4
    41.5 GB

    GPU / 1 × B200

    Workbench →
  • Qwen2.5 14B

    14.7B
    FP16
    34.0 GB
    INT8
    17.8 GB
    INT4
    9.7 GB

    GPU / 1 × A100 40GB

    Workbench →
  • Qwen2.5 32B

    32.5B
    FP16
    73.6 GB
    INT8
    37.9 GB
    INT4
    20.0 GB

    GPU / 1 × H100 NVL

    Workbench →
  • Qwen2.5 72B

    72.7B
    FP16
    162.6 GB
    INT8
    82.7 GB
    INT4
    42.7 GB

    GPU / 1 × B200

    Workbench →
  • Mistral 7B

    7.25B
    FP16
    17.0 GB
    INT8
    9.0 GB
    INT4
    5.1 GB

    GPU / 1 × RTX 4090

    Workbench →
  • Mixtral 8x7B

    46.7B
    FP16
    103.8 GB
    INT8
    52.4 GB
    INT4
    26.8 GB

    GPU / 1 × H200

    Workbench →
  • DeepSeek V3 671B

    671B
    FP16
    1,477 GB
    INT8
    738.7 GB
    INT4
    369.6 GB

    GPU / 8 × MI325X

    Workbench →
  • Custom

    FP16
    45.8 GB
    INT8
    23.8 GB
    INT4
    12.8 GB

    GPU / 1 × H100 SXM

    Workbench →

08 — Origin

FROM KEYSTROKE
TO COMPUTE.

Every model begins as an idea.

Every idea eventually needs hardware.

MUSEBOARD connects the two.

  1. 01InputKEY / 01
  2. 02ModelPARAMS / B
  3. 03MemoryMEM / GB
  4. 04GPUHW / MATCH
  5. 05ComputeSYS / READY

SYS / READY

DON'T GUESS
YOUR COMPUTE.

MAP IT.

Launch Workbench

Interactive compute planning for AI builders.