Skip to content
‹ All benchmarks

2026-09-09 · Linux · PyTorch 2.13.0+cu130

N = 65,536 · Scale 250

ScalingThroughput
Log Y

Workloads and operations

PT × CT matrix multiplication

64 × 64 · Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)012.52537.550124816Batch sizeNVIDIA H100 NVL · 400 W · 39.9NVIDIA H100 NVL · 400 W · 38.3NVIDIA H100 NVL · 400 W · 38.5NVIDIA H100 NVL · 400 W · 38.8NVIDIA H100 NVL · 400 W · 39
HoistingOn

CT × CT matrix multiplication

64 × 64 · Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)06.2512.518.825124816Batch sizeNVIDIA H100 NVL · 400 W · 21.3NVIDIA H100 NVL · 400 W · 21.2NVIDIA H100 NVL · 400 W · 21.4NVIDIA H100 NVL · 400 W · 21.6NVIDIA H100 NVL · 400 W · 21.7
HoistingOn

Nonlinear polynomial evaluation

Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)07.51522.530124816Batch sizeNVIDIA H100 NVL · 400 W · 23.1NVIDIA H100 NVL · 400 W · 21.1NVIDIA H100 NVL · 400 W · 21.5NVIDIA H100 NVL · 400 W · 21.8NVIDIA H100 NVL · 400 W · 21.9

NTT

Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)03,7507,50011,30015,000124816Batch sizeNVIDIA H100 NVL · 400 W · 7,930NVIDIA H100 NVL · 400 W · 8,830NVIDIA H100 NVL · 400 W · 9,180NVIDIA H100 NVL · 400 W · 9,550NVIDIA H100 NVL · 400 W · 9,740

Key switching

Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)0175350525700124816Batch sizeNVIDIA H100 NVL · 400 W · 555NVIDIA H100 NVL · 400 W · 579NVIDIA H100 NVL · 400 W · 573NVIDIA H100 NVL · 400 W · 576NVIDIA H100 NVL · 400 W · 531

RNS arithmetic

Entry depth 0
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)020,00040,00060,00080,000124816Batch sizeNVIDIA H100 NVL · 400 W · 46,200NVIDIA H100 NVL · 400 W · 52,900NVIDIA H100 NVL · 400 W · 61,600NVIDIA H100 NVL · 400 W · 67,600NVIDIA H100 NVL · 400 W · 69,300

Matrix operations

8 operations

Rotate (1 offset)

Entry depth 0 · Shift +1
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)0163325488650124816Batch sizeNVIDIA H100 NVL · 400 W · 513NVIDIA H100 NVL · 400 W · 533NVIDIA H100 NVL · 400 W · 531NVIDIA H100 NVL · 400 W · 536NVIDIA H100 NVL · 400 W · 507

Rotate-many (7 offsets)

HoistingOn
Entry depth 0 · Shifts +1…+7 · one input
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)050100150200124816Batch sizeNVIDIA H100 NVL · 400 W · 151NVIDIA H100 NVL · 400 W · 151NVIDIA H100 NVL · 400 W · 144NVIDIA H100 NVL · 400 W · 147NVIDIA H100 NVL · 400 W · 147

Multiply (PT × CT)

Entry depth 0 · No rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)07,50015,00022,50030,000124816Batch sizeNVIDIA H100 NVL · 400 W · 25,700NVIDIA H100 NVL · 400 W · 18,600NVIDIA H100 NVL · 400 W · 20,200NVIDIA H100 NVL · 400 W · 20,800NVIDIA H100 NVL · 400 W · 21,000

Sum products (8 PT × CT terms)

Entry depth 0 · Sum of eight products · no rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)02,0004,0006,0008,000124816Batch sizeNVIDIA H100 NVL · 400 W · 6,120NVIDIA H100 NVL · 400 W · 6,580NVIDIA H100 NVL · 400 W · 6,690NVIDIA H100 NVL · 400 W · 6,670NVIDIA H100 NVL · 400 W · 6,130

Rescale (2 components)

Entry depth 0 · One Q depth group
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)01,0002,0003,0004,000124816Batch sizeNVIDIA H100 NVL · 400 W · 2,900NVIDIA H100 NVL · 400 W · 3,110NVIDIA H100 NVL · 400 W · 3,250NVIDIA H100 NVL · 400 W · 3,170NVIDIA H100 NVL · 400 W · 3,090

Multiply (CT × CT)

Entry depth 0 · Distinct inputs · no relinearization or rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)07,50015,00022,50030,000124816Batch sizeNVIDIA H100 NVL · 400 W · 20,000NVIDIA H100 NVL · 400 W · 23,700NVIDIA H100 NVL · 400 W · 24,400NVIDIA H100 NVL · 400 W · 24,600NVIDIA H100 NVL · 400 W · 24,700

Add (CT + CT)

Entry depth 0 · Same depth and scale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)011,30022,50033,80045,000124816Batch sizeNVIDIA H100 NVL · 400 W · 24,800NVIDIA H100 NVL · 400 W · 30,600NVIDIA H100 NVL · 400 W · 34,700NVIDIA H100 NVL · 400 W · 37,200NVIDIA H100 NVL · 400 W · 38,200

Relinearize

Entry depth 0 · 3 → 2 components · no rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)0150300450600124816Batch sizeNVIDIA H100 NVL · 400 W · 519NVIDIA H100 NVL · 400 W · 528NVIDIA H100 NVL · 400 W · 514NVIDIA H100 NVL · 400 W · 516NVIDIA H100 NVL · 400 W · 498

Polynomial operations

Additional operations

Square (CT)

Entry depth 0 · Same input · no relinearization or rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)08,75017,50026,30035,000124816Batch sizeNVIDIA H100 NVL · 400 W · 20,400NVIDIA H100 NVL · 400 W · 24,700NVIDIA H100 NVL · 400 W · 26,800NVIDIA H100 NVL · 400 W · 27,900NVIDIA H100 NVL · 400 W · 27,900

Multiply scalar & rescale

Entry depth 0 · Scalar −0.3 · one rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)08751,7502,6303,500124816Batch sizeNVIDIA H100 NVL · 400 W · 2,720NVIDIA H100 NVL · 400 W · 2,890NVIDIA H100 NVL · 400 W · 3,020NVIDIA H100 NVL · 400 W · 2,970NVIDIA H100 NVL · 400 W · 2,850

Advance depth

Entry depth 0 · Multiply by one & rescale
NVIDIA H100 NVL · 400 W
Throughput (tasks/s)08751,7502,6303,500124816Batch sizeNVIDIA H100 NVL · 400 W · 2,730NVIDIA H100 NVL · 400 W · 2,890NVIDIA H100 NVL · 400 W · 3,020NVIDIA H100 NVL · 400 W · 2,950NVIDIA H100 NVL · 400 W · 2,860
Measurement conditions
HardwareNVIDIA H100 NVL
Host CPUAMD EPYC 9115 16-Core Processor
Intra-op / inter-op threads16 / 16
PyTorch / CUDA2.13.0+cu130 / 13.0
GPU power limit400 W
Host allocationSlurm allocation; selected GPU idle before measurement
ExecutionEager CKKS and Backend RNS/NTT · prepared inputs and resources
Additional measurements · 2026-09-10245 cases · 20260910T003931Z-6b6de71f