{
  "schema_version": 1,
  "as_of": "2026-10-07",
  "description": "GPU specifications from each vendor's own datasheet, product page or architecture whitepaper. Compute figures are dense tensor-core TFLOPS (no structured sparsity). null means the vendor does not publish the figure (or the GPU lacks that data type); it is never filled from third-party sources.",
  "fields": {
    "vram_gb": "Memory as marketed. The calculators treat it as GiB; the capacity the driver reports can be slightly lower (for example with ECC enabled).",
    "memory_bandwidth_gbs": "Peak memory bandwidth in GB/s.",
    "fp16_tflops_dense": "FP16/BF16 tensor throughput without sparsity. GeForce cards: the FP32-accumulate rate, which is what BF16/FP16 inference kernels use; the FP16-accumulate rate is twice that and is in the note.",
    "fp8_tflops_dense": "FP8 tensor throughput without sparsity; null on GPUs without FP8 tensor cores (Ampere).",
    "fp4_tflops_dense": "FP4 tensor throughput without sparsity where the vendor lists it.",
    "multi_gpu_link": "nvlink (NVLink/NVSwitch baseboard), nvlink-bridge (pairs or small groups of PCIe cards), infinity-fabric (AMD), or pcie (no GPU-to-GPU link)."
  },
  "gpus": [
    {
      "id": "rtx-4090",
      "name": "NVIDIA GeForce RTX 4090",
      "vendor": "NVIDIA",
      "architecture": "Ada Lovelace",
      "vram_gb": 24,
      "memory_type": "GDDR6X",
      "memory_bandwidth_gbs": 1008,
      "fp16_tflops_dense": 165.2,
      "fp8_tflops_dense": 330.3,
      "fp4_tflops_dense": null,
      "tdp_w": 450,
      "form_factor": "PCIe Gen4 card, 3-slot (Founders Edition)",
      "interconnect": "PCIe Gen4, no NVLink",
      "multi_gpu_link": "pcie",
      "category": "consumer",
      "source_url": "https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/",
      "source_url_2": "https://images.nvidia.com/aem-dam/Solutions/Data-Center/l4/nvidia-ada-gpu-architecture-whitepaper-V2.02.pdf",
      "as_of": "2026-10-07",
      "note": "Product page lists only '1321 AI TOPS', 24 GB GDDR6X, 'Total Graphics Power (W) 450', 'NVLink (SLI-Ready) No'. Ada whitepaper V2.02 Appendix A: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 165.2/330.4' (used) vs 'with FP16 Accumulate 330.3/660.6'; FP8 'with FP32 Accumulate 330.3/660.6' (used, matches FP32-accumulate choice) vs 'with FP16 Accumulate 660.6/1321.2'; second figure = sparsity. Bandwidth '1008 GB/sec' from whitepaper."
    },
    {
      "id": "rtx-5090",
      "name": "NVIDIA GeForce RTX 5090",
      "vendor": "NVIDIA",
      "architecture": "Blackwell",
      "vram_gb": 32,
      "memory_type": "GDDR7",
      "memory_bandwidth_gbs": 1792,
      "fp16_tflops_dense": 209.5,
      "fp8_tflops_dense": 419,
      "fp4_tflops_dense": 1676,
      "tdp_w": 575,
      "form_factor": "PCIe Gen5 card, 2-slot (Founders Edition)",
      "interconnect": "PCIe Gen5, no NVLink",
      "multi_gpu_link": "pcie",
      "category": "consumer",
      "source_url": "https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/",
      "source_url_2": "https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf",
      "as_of": "2026-10-07",
      "note": "Product page lists only '3352 AI TOPS', '32 GB GDDR7', 'Total Graphics Power (W) 575', 'NVLink (SLI-Ready) No'. RTX Blackwell whitepaper v1.1 Table 3: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 209.5/419' (used) vs 'with FP16 Accumulate 419/838'; FP8 'with FP32 Accumulate 419/838' (used) vs 'with FP16 Accumulate 838/1676'; 'Peak FP4 Tensor TFLOPS with FP32 Accumulate (FP4 AI TOPS) 1676/3352'; second figure = sparsity. Bandwidth '1792 GB/sec' from whitepaper."
    },
    {
      "id": "rtx-a6000",
      "name": "NVIDIA RTX A6000",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 48,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 768,
      "fp16_tflops_dense": 154.8,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 300,
      "form_factor": "PCIe dual-slot, full height (4.4\" H x 10.5\" L)",
      "interconnect": "NVLink bridge (2 GPUs) 112.5 GB/s bidirectional; PCIe 4.0 x16",
      "multi_gpu_link": "nvlink-bridge",
      "category": "workstation",
      "source_url": "https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/proviz-print-nvidia-rtx-a6000-datasheet-us-nvidia-1454980-r9-web%20(1).pdf",
      "source_url_2": "https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/pdf/NVIDIA-RTX-Blackwell-PRO-GPU-Architecture-v1_1.pdf",
      "as_of": "2026-10-07",
      "note": "Datasheet: 'Tensor performance 309.7 TFLOPS' with footnote 'Effective teraFLOPS (TFLOPS) using the new sparsity feature'. RTX PRO Blackwell whitepaper v1.1 Table 4 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 154.8/309.6' (same as FP16 accumulate for this pro card). FP8 'N/A' (Ampere has no FP8 tensor cores)."
    },
    {
      "id": "rtx-6000-ada",
      "name": "NVIDIA RTX 6000 Ada Generation",
      "vendor": "NVIDIA",
      "architecture": "Ada Lovelace",
      "vram_gb": 48,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 960,
      "fp16_tflops_dense": 364,
      "fp8_tflops_dense": 728.5,
      "fp4_tflops_dense": null,
      "tdp_w": 300,
      "form_factor": "PCIe dual-slot, full height (4.4\" H x 10.5\" L)",
      "interconnect": "PCIe 4.0 x16, no NVLink",
      "multi_gpu_link": "pcie",
      "category": "workstation",
      "source_url": "https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/rtx-6000/proviz-print-rtx6000-datasheet-web-2504660.pdf",
      "source_url_2": "https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/pdf/NVIDIA-RTX-Blackwell-PRO-GPU-Architecture-v1_1.pdf",
      "as_of": "2026-10-07",
      "note": "Datasheet: 'Tensor performance 1457.0 TFLOPS' = 'Effective FP8 teraFLOPS (TFLOPS) using the new sparsity feature'; 'NVIDIA NVLink No'. RTX PRO Blackwell whitepaper v1.1 Table 4: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 364/728' and 'Peak FP8 Tensor TFLOPS with FP32 Accumulate 728.5/1457' (FP16-accumulate rates identical on this pro card)."
    },
    {
      "id": "rtx-pro-6000-blackwell",
      "name": "NVIDIA RTX PRO 6000 Blackwell Server Edition",
      "vendor": "NVIDIA",
      "architecture": "Blackwell",
      "vram_gb": 96,
      "memory_type": "GDDR7",
      "memory_bandwidth_gbs": 1597,
      "fp16_tflops_dense": null,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 600,
      "form_factor": "PCIe Gen5 dual-slot FHFL (air) / single-slot FHXL (liquid)",
      "interconnect": "PCIe Gen5 x16, no NVLink",
      "multi_gpu_link": "pcie",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/",
      "source_url_2": "https://dam-cdn.nvd.orangelogic.com/AssetLink/3km2720jiy76r06ctf8mxg21743p1058.pdf",
      "as_of": "2026-10-07",
      "note": "Compute left null: product page lists 'FP4 Tensor Core 4 PFLOPS', 'FP8 Tensor Core 2 PFLOPS', 'FP16 | BF16 Tensor Core 1 PFLOP' with NO sparsity marker, and the datasheet only gives 'Peak FP4 AI PFLOPS 4 PFLOPS'. NVIDIA's RTX PRO Blackwell whitepaper v1.1 lists the same-GPU Workstation Edition (GB202, 752 Tensor Cores, 126 TFLOPS FP32 vs 120 for Server Edition) at FP16 503.8/1007.6, FP8 1007.6/2015.2, FP4 2015.2/4030.4 TFLOPS (dense/sparse), so the page figures appear to be sparse; halving would give about 500/1000/2000 TFLOPS, but NVIDIA publishes no dense figure for the Server Edition. Product brief: 'Memory type GDDR7', 'Peak memory bandwidth 1,597 GB/s', 600 W, 'NVIDIA NVLink Not supported'."
    },
    {
      "id": "l4",
      "name": "NVIDIA L4",
      "vendor": "NVIDIA",
      "architecture": "Ada Lovelace",
      "vram_gb": 24,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 300,
      "fp16_tflops_dense": 121,
      "fp8_tflops_dense": 242,
      "fp4_tflops_dense": null,
      "tdp_w": 72,
      "form_factor": "PCIe 1-slot low-profile (HHHL)",
      "interconnect": "PCIe Gen4 x16 64 GB/s, no NVLink listed",
      "multi_gpu_link": "pcie",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/l4/",
      "source_url_2": "https://images.nvidia.com/aem-dam/Solutions/Data-Center/l4/nvidia-ada-gpu-architecture-whitepaper-V2.02.pdf",
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core 242 teraFLOPS*', 'FP8 Tensor Core 485 teraFLOPs*', '* Shown with sparsity. Specifications are one-half lower without sparsity.' Ada whitepaper V2.02 Table 5 lists dense explicitly: 'FP16 Tensor Core Performance 121 | 242 TFLOPS', 'FP8 ... 242 | 485 TFLOPS', '24GB GDDR6 w/ ECC'. Neither source mentions NVLink; interconnect listed as PCIe only."
    },
    {
      "id": "a10",
      "name": "NVIDIA A10",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 24,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 600,
      "fp16_tflops_dense": 125,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 150,
      "form_factor": "PCIe single-slot FHFL",
      "interconnect": "PCIe Gen4 64 GB/s, no NVLink listed",
      "multi_gpu_link": "pcie",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/products/a10-gpu/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core 125 teraFLOPS | 250 teraFLOPS*', '*With Sparsity' (dense explicit). No FP8 tensor cores (Ampere). Interconnect listed only as 'PCIe Gen4 64GB/s'."
    },
    {
      "id": "a40",
      "name": "NVIDIA A40",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 48,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 696,
      "fp16_tflops_dense": 149.7,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 300,
      "form_factor": "PCIe dual-slot (4.4\" H x 10.5\" L), passive",
      "interconnect": "NVLink bridge (2-way) 112.5 GB/s bidirectional; PCIe Gen4",
      "multi_gpu_link": "nvlink-bridge",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a40/proviz-print-nvidia-a40-datasheet-us-nvidia-1469711-r8-web.pdf",
      "source_url_2": "https://www.nvidia.com/en-us/data-center/a40/",
      "as_of": "2026-10-07",
      "note": "Datasheet: 'Peak FP16 Tensor TFLOPS with FP16 Accumulate 149.7 | 299.4*' and 'Peak BF16 Tensor TFLOPS with FP32 Accumulate 149.7 | 299.4*', '* Structural sparsity enabled'. No FP8 tensor cores (Ampere). Datasheet lists PCIe Gen4 as 31.5 GB/s, product page as 64GB/s."
    },
    {
      "id": "l40s",
      "name": "NVIDIA L40S",
      "vendor": "NVIDIA",
      "architecture": "Ada Lovelace",
      "vram_gb": 48,
      "memory_type": "GDDR6",
      "memory_bandwidth_gbs": 864,
      "fp16_tflops_dense": 362.05,
      "fp8_tflops_dense": 733,
      "fp4_tflops_dense": null,
      "tdp_w": 350,
      "form_factor": "PCIe dual-slot (4.4\" H x 10.5\" L), passive",
      "interconnect": "PCIe Gen4 x16 64 GB/s, no NVLink",
      "multi_gpu_link": "pcie",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/l40s/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page full spec table: 'FP16 Tensor Core 362.05 I 733*', 'FP8 Tensor Core 733 I 1,466*', '*With Sparsity' (dense listed explicitly; note 362.05 is not exactly half of 733). '48GB GDDR6 with ECC', 'Max Power Consumption 350W', 'NVIDIA NVLink Support: No'."
    },
    {
      "id": "a100-sxm-40gb",
      "name": "NVIDIA A100 40GB SXM",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 40,
      "memory_type": "HBM2",
      "memory_bandwidth_gbs": 1555,
      "fp16_tflops_dense": 312,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 400,
      "form_factor": "SXM",
      "interconnect": "NVLink 600 GB/s; PCIe Gen4 64 GB/s",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "A100 datasheet (r4, 2021): 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity'; A100 40GB SXM column: '40GB HBM2', '1,555GB/s', TDP 400W. No FP8 tensor cores (Ampere). The current A100 product page lists only the 80GB variants."
    },
    {
      "id": "a100-pcie-80gb",
      "name": "NVIDIA A100 80GB PCIe",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 80,
      "memory_type": "HBM2e",
      "memory_bandwidth_gbs": 1935,
      "fp16_tflops_dense": 312,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 300,
      "form_factor": "PCIe dual-slot air-cooled or single-slot liquid-cooled",
      "interconnect": "NVLink bridge for 2 GPUs 600 GB/s; PCIe Gen4 64 GB/s",
      "multi_gpu_link": "nvlink-bridge",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/a100/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense explicit). No FP8 tensor cores (Ampere). NVLink only via bridge pairing two cards."
    },
    {
      "id": "a100-sxm-80gb",
      "name": "NVIDIA A100 80GB SXM",
      "vendor": "NVIDIA",
      "architecture": "Ampere",
      "vram_gb": 80,
      "memory_type": "HBM2e",
      "memory_bandwidth_gbs": 2039,
      "fp16_tflops_dense": 312,
      "fp8_tflops_dense": null,
      "fp4_tflops_dense": null,
      "tdp_w": 400,
      "form_factor": "SXM",
      "interconnect": "NVLink 600 GB/s; PCIe Gen4 64 GB/s",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/a100/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense listed explicitly). No FP8 tensor cores (Ampere). '400W TDP for standard configuration. HGX A100-80GB custom thermal solution (CTS) SKU can support TDPs up to 500W'."
    },
    {
      "id": "h100-pcie",
      "name": "NVIDIA H100 PCIe",
      "vendor": "NVIDIA",
      "architecture": "Hopper",
      "vram_gb": 80,
      "memory_type": "HBM2e",
      "memory_bandwidth_gbs": 2000,
      "fp16_tflops_dense": 756,
      "fp8_tflops_dense": 1513,
      "fp4_tflops_dense": null,
      "tdp_w": 350,
      "form_factor": "PCIe Gen5 dual-slot FHFL",
      "interconnect": "NVLink bridge (2 GPUs) 600 GB/s; PCIe Gen5 x16",
      "multi_gpu_link": "nvlink-bridge",
      "category": "datacenter",
      "source_url": "https://dam-cdn.nvd.orangelogic.com/AssetLink/705n6ur546g0uk43w0117r17n8042d73.pdf",
      "source_url_2": "https://www.nvidia.com/content/dam/en-zz/Solutions/gtcs22/data-center/h100/PB-11133-001_v01.pdf",
      "as_of": "2026-10-07",
      "note": "Hopper whitepaper v1.04 ('Includes final GPU / memory clocks and final TFLOPS performance specs') Table 3, H100 PCIe: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 756/1513' and 'Peak FP8 Tensor TFLOPS 1513/3026' (second figure = sparsity). Product brief PB-11133 v02: 'Memory type HBM2e', 'Peak memory bandwidth 2,000 GB/s', 350 W max, 'Total maximum NVLink bandwidth 600 Gbytes per second' (its overview text says 900 GB/s; table value used). Whitepaper lists bandwidth as 2039 GB/sec. NVIDIA's current H100 product page no longer lists H100 PCIe; an older 2022 H100 datasheet listed preliminary rounded figures (1,600 TFLOPS* FP16, 2TB/s), not used."
    },
    {
      "id": "h100-nvl",
      "name": "NVIDIA H100 NVL",
      "vendor": "NVIDIA",
      "architecture": "Hopper",
      "vram_gb": 94,
      "memory_type": "HBM3",
      "memory_bandwidth_gbs": 3900,
      "fp16_tflops_dense": 835.5,
      "fp8_tflops_dense": 1670.5,
      "fp4_tflops_dense": null,
      "tdp_w": 400,
      "form_factor": "PCIe dual-slot air-cooled",
      "interconnect": "NVLink bridge 600 GB/s; PCIe Gen5 128 GB/s",
      "multi_gpu_link": "nvlink-bridge",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/h100/",
      "source_url_2": "https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/h100/PB-11773-001_v01.pdf",
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core* 1,671 teraFLOPS', 'FP8 Tensor Core* 3,341 teraFLOPS', '* With sparsity'; dense values are half. TDP '350-400W (configurable)' (400 W recorded). Product brief PB-11773: 'Memory type HBM3', 'Memory size 94 GB', 'Peak memory bandwidth 3,938 GB/s' (page says 3.9TB/s)."
    },
    {
      "id": "h100-sxm",
      "name": "NVIDIA H100 SXM",
      "vendor": "NVIDIA",
      "architecture": "Hopper",
      "vram_gb": 80,
      "memory_type": "HBM3",
      "memory_bandwidth_gbs": 3350,
      "fp16_tflops_dense": 989.4,
      "fp8_tflops_dense": 1978.9,
      "fp4_tflops_dense": null,
      "tdp_w": 700,
      "form_factor": "SXM (SXM5)",
      "interconnect": "NVLink 900 GB/s; PCIe Gen5 128 GB/s",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/h100/",
      "source_url_2": "https://dam-cdn.nvd.orangelogic.com/AssetLink/705n6ur546g0uk43w0117r17n8042d73.pdf",
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core* 1,979 teraFLOPS', 'FP8 Tensor Core* 3,958 teraFLOPS', '* With sparsity'; Hopper whitepaper v1.04 Table 3 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 989.4/1978.9', FP8 1978.9/3957.8 (sparse after slash), and '80 GB HBM3'. TDP listed as 'Up to 700W (configurable)'."
    },
    {
      "id": "h200-nvl",
      "name": "NVIDIA H200 NVL",
      "vendor": "NVIDIA",
      "architecture": "Hopper",
      "vram_gb": 141,
      "memory_type": "HBM3e",
      "memory_bandwidth_gbs": 4800,
      "fp16_tflops_dense": 835.5,
      "fp8_tflops_dense": 1670.5,
      "fp4_tflops_dense": null,
      "tdp_w": 600,
      "form_factor": "PCIe dual-slot air-cooled",
      "interconnect": "2- or 4-way NVLink bridge 900 GB/s per GPU; PCIe Gen5 128 GB/s",
      "multi_gpu_link": "nvlink-bridge",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/h200/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core² 1,671 TFLOPS', 'FP8 Tensor Core² 3,341 TFLOPS', '² With sparsity'; dense = half. Marked '¹ Preliminary specifications'. TDP 'Up to 600W (configurable)'. HBM3e from page text about H200 memory."
    },
    {
      "id": "h200-sxm",
      "name": "NVIDIA H200 SXM",
      "vendor": "NVIDIA",
      "architecture": "Hopper",
      "vram_gb": 141,
      "memory_type": "HBM3e",
      "memory_bandwidth_gbs": 4800,
      "fp16_tflops_dense": 989.5,
      "fp8_tflops_dense": 1979,
      "fp4_tflops_dense": null,
      "tdp_w": 700,
      "form_factor": "SXM",
      "interconnect": "NVLink 900 GB/s; PCIe Gen5 128 GB/s",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/h200/",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'FP16 Tensor Core² 1,979 TFLOPS', 'FP8 Tensor Core² 3,958 TFLOPS', '² With sparsity'; dense = half. Page footnote '¹ Preliminary specifications. May be subject to change.' Page text: '141 gigabytes (GB) of HBM3e memory at 4.8 terabytes per second'. TDP 'Up to 700W (configurable)'."
    },
    {
      "id": "b200",
      "name": "NVIDIA B200 (HGX B200)",
      "vendor": "NVIDIA",
      "architecture": "Blackwell",
      "vram_gb": 180,
      "memory_type": "HBM3E",
      "memory_bandwidth_gbs": 8000,
      "fp16_tflops_dense": 2250,
      "fp8_tflops_dense": 4500,
      "fp4_tflops_dense": 9000,
      "tdp_w": 1000,
      "form_factor": "SXM (SXM6, 8-GPU HGX baseboard)",
      "interconnect": "NVLink 5 1.8 TB/s GPU-to-GPU (NVLink Switch)",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://www.nvidia.com/en-us/data-center/hgx/",
      "source_url_2": "https://images.nvidia.com/aem-dam/Solutions/documents/HGX-B200-PCF-Summary.pdf",
      "as_of": "2026-10-07",
      "note": "No per-GPU B200 datasheet found; compute derived by dividing NVIDIA's 8-GPU HGX B200 totals by 8. HGX page: 'FP4 Tensor Core 144 PFLOPS | 72 PFLOPS' ('Sparse | Dense'), 'FP8/FP6 Tensor Core 72 PFLOPS', 'FP16/BF16 Tensor Core 36 PFLOPS' ('Specification in Sparse. Dense is 1/2 sparse spec shown.'). PCF summary: 'eight NVIDIA Blackwell B200 GPUs, each with 180 GB of HBM3E', 'Per individual GPU: Configurable up to 1000 W', total bandwidth 'Up to 62 TB/s'. Per-GPU bandwidth 'Up to 8TB/s' from https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/components.html (62/8 = 7.75 TB/s would follow from the PCF total)."
    },
    {
      "id": "b300",
      "name": "NVIDIA B300 (HGX B300, Blackwell Ultra)",
      "vendor": "NVIDIA",
      "architecture": "Blackwell Ultra",
      "vram_gb": 270,
      "memory_type": "HBM3E",
      "memory_bandwidth_gbs": 7700,
      "fp16_tflops_dense": 2250,
      "fp8_tflops_dense": 4500,
      "fp4_tflops_dense": 14000,
      "tdp_w": 1100,
      "form_factor": "SXM (8-GPU HGX baseboard)",
      "interconnect": "NVLink 5 1.8 TB/s; PCIe Gen6 256 GB/s",
      "multi_gpu_link": "nvlink",
      "category": "datacenter",
      "source_url": "https://dam-cdn.nvd.orangelogic.com/AssetLink/1k0p832eq8r5ca0u5383ie5o4tp3bst1.pdf",
      "source_url_2": "https://www.nvidia.com/en-us/data-center/hgx/",
      "as_of": "2026-10-07",
      "note": "NVIDIA Blackwell Ultra Datasheet, 'Individual Blackwell Ultra GPU Specifications', HGX B300 column: 'FP4 Tensor Core 18 PFLOPS | 14 PFLOPS' (Sparse | Dense), 'FP8/FP6 9 PFLOPS', 'FP16/BF16 4.5 PFLOPS' ('Specification in sparse. Dense is 1/2 sparse spec shown'), '270 GB HBM3E | 7.7 TB/s', 'Configurable up to 1,100 W' (GB300 NVL72 variant: 279 GB, 8 TB/s, 1,400 W). Conflicts: HGX page 8-GPU FP4 dense 108 PFLOPS (=13.5/GPU) and docs.nvidia.com reference architecture says B300 SXM 288GB, up to 8TB/s; datasheet per-GPU values used."
    },
    {
      "id": "mi300x",
      "name": "AMD Instinct MI300X",
      "vendor": "AMD",
      "architecture": "CDNA 3",
      "vram_gb": 192,
      "memory_type": "HBM3",
      "memory_bandwidth_gbs": 5300,
      "fp16_tflops_dense": 1307.4,
      "fp8_tflops_dense": 2614.9,
      "fp4_tflops_dense": null,
      "tdp_w": 750,
      "form_factor": "OAM",
      "interconnect": "AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16",
      "multi_gpu_link": "infinity-fabric",
      "category": "datacenter",
      "source_url": "https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: 'Peak Half Precision (FP16) Performance 1.3 PFLOPs' and 'with Structured Sparsity 2.61 PFLOPs'; FP8 '2.61 PFLOPs' / sparsity '5.22 PFLOPs'. Page footnote gives precise dense values: '1307.4 TFLOPS peak theoretical half precision (FP16)', '2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '750W Peak'. No FP4 listed."
    },
    {
      "id": "mi325x",
      "name": "AMD Instinct MI325X",
      "vendor": "AMD",
      "architecture": "CDNA 3",
      "vram_gb": 256,
      "memory_type": "HBM3E",
      "memory_bandwidth_gbs": 6000,
      "fp16_tflops_dense": 1307.4,
      "fp8_tflops_dense": 2614.9,
      "fp4_tflops_dense": null,
      "tdp_w": 1000,
      "form_factor": "OAM",
      "interconnect": "AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16",
      "multi_gpu_link": "infinity-fabric",
      "category": "datacenter",
      "source_url": "https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html",
      "source_url_2": null,
      "as_of": "2026-10-07",
      "note": "Product page: FP16 '1.3 PFLOPs' (sparsity '2.61 PFLOPs'), FP8 '2.61 PFLOPs' (sparsity '5.22 PFLOPs'); footnote MI325-002: '1307.4 TFLOPS peak theoretical half precision (FP16)... 2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '1000W Peak'. No FP4 listed."
    },
    {
      "id": "mi355x",
      "name": "AMD Instinct MI355X",
      "vendor": "AMD",
      "architecture": "CDNA 4",
      "vram_gb": 288,
      "memory_type": "HBM3E",
      "memory_bandwidth_gbs": 8000,
      "fp16_tflops_dense": 2516.6,
      "fp8_tflops_dense": 5033.2,
      "fp4_tflops_dense": 10066.3,
      "tdp_w": 1400,
      "form_factor": "OAM",
      "interconnect": "AMD Infinity Fabric: 7 x 153.6 GB/s scale-up links; PCIe Gen5 x16 128 GB/s",
      "multi_gpu_link": "infinity-fabric",
      "category": "datacenter",
      "source_url": "https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html",
      "source_url_2": "https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/product-briefs/amd-instinct-mi355x-gpu-brochure.pdf",
      "as_of": "2026-10-07",
      "note": "Product page: 'Peak Half Precision Matrix (FP16) Performance 2.5 PFLOPs' (sparsity '5 PFLOPs'), OCP-FP8 '5 PFLOPs' (sparsity '10.1 PFLOPs'), 'MXFP4 Performance 10.1 PFLOPs' (no sparsity figure), TBP '1400W'. Brochure gives precise values: FP16 matrix 2.5166 PFLOPS (5.0332 w/ sparsity), OCP-FP8 5.0332 (10.0664 w/ sparsity), MXFP4 10.0663 (sparsity N/A). FP4 is MXFP4 (microscaling)."
    }
  ]
}
