Skip to content
‹ All benchmarks

2026-09-08 · Linux · PyTorch 2.13.0+cu130

N = 65,536 · Scale 250

ScalingThroughput
Log Y

Workloads and operations

PT × CT matrix multiplication

64 × 64 · Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)05101520124816Batch sizeNVIDIA RTX A6000 · 300 W · 13.2NVIDIA RTX A6000 · 300 W · 13.5NVIDIA RTX A6000 · 300 W · 13.6NVIDIA RTX A6000 · 300 W · 13.7NVIDIA RTX A6000 · 300 W · 13.7
HoistingOn

CT × CT matrix multiplication

64 × 64 · Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)02.254.56.759124816Batch sizeNVIDIA RTX A6000 · 300 W · 7.33NVIDIA RTX A6000 · 300 W · 7.45NVIDIA RTX A6000 · 300 W · 7.53NVIDIA RTX A6000 · 300 W · 7.58NVIDIA RTX A6000 · 300 W · 7.6
HoistingOn

Nonlinear polynomial evaluation

Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)02.254.56.759124816Batch sizeNVIDIA RTX A6000 · 300 W · 7.53NVIDIA RTX A6000 · 300 W · 7.69NVIDIA RTX A6000 · 300 W · 7.78NVIDIA RTX A6000 · 300 W · 7.82NVIDIA RTX A6000 · 300 W · 7.84

NTT

Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)08751,7502,6303,500124816Batch sizeNVIDIA RTX A6000 · 300 W · 2,640NVIDIA RTX A6000 · 300 W · 2,730NVIDIA RTX A6000 · 300 W · 2,790NVIDIA RTX A6000 · 300 W · 2,820NVIDIA RTX A6000 · 300 W · 2,840

Key switching

Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)062.5125188250124816Batch sizeNVIDIA RTX A6000 · 300 W · 192NVIDIA RTX A6000 · 300 W · 195NVIDIA RTX A6000 · 300 W · 198NVIDIA RTX A6000 · 300 W · 199NVIDIA RTX A6000 · 300 W · 200

RNS arithmetic

Entry depth 0
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)05,00010,00015,00020,000124816Batch sizeNVIDIA RTX A6000 · 300 W · 13,100NVIDIA RTX A6000 · 300 W · 14,200NVIDIA RTX A6000 · 300 W · 14,800NVIDIA RTX A6000 · 300 W · 15,100NVIDIA RTX A6000 · 300 W · 15,300

Matrix operations

8 operations

Rotate (1 offset)

Entry depth 0 · Shift +1
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)062.5125188250124816Batch sizeNVIDIA RTX A6000 · 300 W · 173NVIDIA RTX A6000 · 300 W · 176NVIDIA RTX A6000 · 300 W · 179NVIDIA RTX A6000 · 300 W · 180NVIDIA RTX A6000 · 300 W · 180

Rotate-many (7 offsets)

HoistingOn
Entry depth 0 · Shifts +1…+7 · one input
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)013.827.541.355124816Batch sizeNVIDIA RTX A6000 · 300 W · 45.7NVIDIA RTX A6000 · 300 W · 46.1NVIDIA RTX A6000 · 300 W · 46.5NVIDIA RTX A6000 · 300 W · 46.8NVIDIA RTX A6000 · 300 W · 46.9

Multiply (PT × CT)

Entry depth 0 · No rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)02,0004,0006,0008,000124816Batch sizeNVIDIA RTX A6000 · 300 W · 6,960NVIDIA RTX A6000 · 300 W · 4,390NVIDIA RTX A6000 · 300 W · 4,500NVIDIA RTX A6000 · 300 W · 4,560NVIDIA RTX A6000 · 300 W · 4,610

Sum products (8 PT × CT terms)

Entry depth 0 · Sum of eight products · no rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)07501,5002,2503,000124816Batch sizeNVIDIA RTX A6000 · 300 W · 2,270NVIDIA RTX A6000 · 300 W · 2,320NVIDIA RTX A6000 · 300 W · 2,350NVIDIA RTX A6000 · 300 W · 2,370NVIDIA RTX A6000 · 300 W · 2,380

Rescale (2 components)

Entry depth 0 · One Q depth group
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)03757501,1301,500124816Batch sizeNVIDIA RTX A6000 · 300 W · 1,070NVIDIA RTX A6000 · 300 W · 1,100NVIDIA RTX A6000 · 300 W · 1,110NVIDIA RTX A6000 · 300 W · 1,110NVIDIA RTX A6000 · 300 W · 1,120

Multiply (CT × CT)

Entry depth 0 · Distinct inputs · no relinearization or rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)01,8803,7505,6307,500124816Batch sizeNVIDIA RTX A6000 · 300 W · 5,940NVIDIA RTX A6000 · 300 W · 6,250NVIDIA RTX A6000 · 300 W · 6,410NVIDIA RTX A6000 · 300 W · 6,490NVIDIA RTX A6000 · 300 W · 6,530

Add (CT + CT)

Entry depth 0 · Same depth and scale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)02,2504,5006,7509,000124816Batch sizeNVIDIA RTX A6000 · 300 W · 6,890NVIDIA RTX A6000 · 300 W · 7,280NVIDIA RTX A6000 · 300 W · 7,510NVIDIA RTX A6000 · 300 W · 7,630NVIDIA RTX A6000 · 300 W · 7,670

Relinearize

Entry depth 0 · 3 → 2 components · no rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)062.5125188250124816Batch sizeNVIDIA RTX A6000 · 300 W · 174NVIDIA RTX A6000 · 300 W · 178NVIDIA RTX A6000 · 300 W · 180NVIDIA RTX A6000 · 300 W · 181NVIDIA RTX A6000 · 300 W · 182

Polynomial operations

Additional operations

Square (CT)

Entry depth 0 · Same input · no relinearization or rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)03,7507,50011,30015,000124816Batch sizeNVIDIA RTX A6000 · 300 W · 8,060NVIDIA RTX A6000 · 300 W · 8,580NVIDIA RTX A6000 · 300 W · 8,890NVIDIA RTX A6000 · 300 W · 9,090NVIDIA RTX A6000 · 300 W · 9,180

Multiply scalar & rescale

Entry depth 0 · Scalar −0.3 · one rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)03757501,1301,500124816Batch sizeNVIDIA RTX A6000 · 300 W · 942NVIDIA RTX A6000 · 300 W · 966NVIDIA RTX A6000 · 300 W · 973NVIDIA RTX A6000 · 300 W · 975NVIDIA RTX A6000 · 300 W · 979

Advance depth

Entry depth 0 · Multiply by one & rescale
NVIDIA RTX A6000 · 300 W
Throughput (tasks/s)03757501,1301,500124816Batch sizeNVIDIA RTX A6000 · 300 W · 943NVIDIA RTX A6000 · 300 W · 968NVIDIA RTX A6000 · 300 W · 973NVIDIA RTX A6000 · 300 W · 976NVIDIA RTX A6000 · 300 W · 979
Measurement conditions
HardwareNVIDIA RTX A6000
Host CPUAMD Ryzen Threadripper PRO 7975WX 32-Cores
Intra-op / inter-op threads32 / 32
PyTorch / CUDA2.13.0+cu130 / 13.0
GPU power limit300 W
Host allocationShared workstation; selected GPU idle before measurement
ExecutionEager CKKS and Backend RNS/NTT · prepared inputs and resources
Additional measurements · 2026-09-09245 cases · 20260909T144553Z-b787d848