FHElium gives both usage models access to the same CKKS stack—from public values and mixed-level IR to Backend linking, RNS and NTT, native CPU/CUDA arithmetic, and rank-local execution.
python -m pip install --index-url https://download.pytorch.org/whl/cu130 "torch==2.13.0+cu130"
python -m pip install --only-binary=fhelium --extra-index-url https://download.fhelium.550w.host/torch213-cu130/simple/ "fhelium==0.30.0"FHElium is under active development. APIs may change significantly between releases.
Build through the full FHE stack while choosing the level of control that fits each workload. Eager execution and Compile Programs use the same Backend, arithmetic resources, and native CPU/CUDA implementations.
Choose a Program workflow, immediate execution with runtime mechanisms, or a rank-local distributed program.
Eager
Create Tensor-backed values, apply visible CKKS state transitions, and dispatch from operand placement on CPU or CUDA.
Open tutorialimport torch
import fhelium as fh
from fhelium.eager import Engine
torch.set_default_device("cuda:0")
engine = Engine(
fh.Preset.slots8192_scale40_depth7_int64
)
ciphertext = engine.encrypt_message(message)
rotated = engine.rotate_with_key(
ciphertext, engine.rotation_key(1)
)
triplet = engine.multiply(
engine.coefficient_domain_to_ntt_domain(ciphertext),
engine.coefficient_domain_to_ntt_domain(rotated),
)
result = engine.rescale_to_next_depth(
engine.relinearize(triplet)
)The compiler captures a call, fuses connected operations into generated kernels, and prepares the program so repeated evaluations run with a fraction of the launch overhead. One BSGS matrix-vector call then finishes in about half the time.
Loading measured timelines…
See JIT compilation for how the compiler does this.
The same conventional cyclic-diagonal BSGS formulation measures both packed plaintext-matrix × ciphertext-vector (PT×CT) and ciphertext-matrix × ciphertext-vector (CT×CT) evaluation. View source