Native operator workflow
Add or modify a native operator from the public mathematical specification outward. One operator uses a backend-neutral schema and device-specific registrations; its Python semantics, dispatcher mutation rules, row mapping, CPU implementation, CUDA implementation, generated wrappers, and tests must agree wherever those backends are supported.
1. Define the operation requirements first
Write down:
mathematical operation
input/output tensor shapes and axes
dtypes and devices
depth and active prime-row mapping
Q or QP basis
coefficient/NTT domain
standard/Montgomery representation
standard `[0, q)` or lazy residue range
functional or mutating behavior
supported singleton/partial-layout cases2
3
4
5
6
7
8
9
10
Decide which layer reconstructs public output metadata and which invalid inputs must fail before native launch.
2. Add a deterministic public reproducer
Before editing CUDA, create a minimal test/harness that:
- builds a fixed preset and indexed device;
- reaches the target operation through public state transitions;
- compares against a cleartext or trusted reference;
- checks output state as well as tensor values;
- identifies the first failing operation/depth;
- synchronizes narrowly enough to locate asynchronous errors.
For a chained bug, add checkpoints after every legal materialization step.
3. Define or update the dispatcher schema
The C++ registration layer owns:
- operator namespace and name;
- argument and return schema;
- device implementation registration;
- mutation and alias annotations;
- pre-launch shape/dtype/device checks.
Use a trailing underscore and correct alias schema for mutating operators. Do not make a functional name mutate storage silently.
Names should describe mathematical/state transitions; backend execution policy such as grouping or shared-memory strategy belongs at the backend/operator variant layer rather than public CKKS semantics.
4. Implement the selected device paths
Pass operand, table, and parameter tensors. Avoid hidden device-global context whose state cannot be represented in the dispatcher schema.
For CPU support, register the schema under the CPU dispatch key and use ATen tensor accessors, integral dtype dispatch, and at::parallel_for where the work size justifies intra-op parallelism. Compile against the parallel backend selected by the installed Torch package; do not introduce an independent FHElium thread pool or link a second OpenMP runtime.
For CUDA support, register the same schema under the CUDA dispatch key. The C++ adapter validates the tensors before launch; the CUDA implementation uses the operand device and PyTorch's current CUDA stream. Do not add a hidden host copy or a device fallback to make a schema appear portable.
Audit:
- configured prime row for every compact input row;
- depth-specific table/parameter offsets;
- Q/QP row order;
- key-digit index versus active local digit index;
- tensor strides and contiguous assumptions;
- current CUDA stream behavior;
- temporary ownership and lifetime;
- lazy/standard residue-range preconditions and outputs;
- integer overflow and modular reduction bounds.
If an operation is intentionally supported by only one backend, document that support in its Python owner and tests. Missing CPU or CUDA registration must fail through normal PyTorch dispatch or the engine's backend validation rather than execute a different algorithm silently.
5. Build from a clean enough state
Rebuild the editable native extension through the selected environment frontend. The uv-managed environment uses:
uv --preview-features extra-build-dependencies \
sync --locked --reinstall-package fhelium2
The pip-managed environment uses:
python -m pip install \
--editable . --verbose --no-build-isolation --no-cache-dir2
Set CMAKE_ARGS=-DFHELIUM_NATIVE_BACKENDS=CPU or CPU+CUDA in the current shell before either command when backend coverage matters. Set CMAKE_BUILD_PARALLEL_LEVEL to control build parallelism. Contributors who install just may use the corresponding optional shortcuts:
just NATIVE_BACKENDS=CPU build-uv
just NATIVE_BACKENDS=CPU+CUDA build-pip2
Use the project environment and the selected Python/Torch/CUDA toolchain. Confirm which shared library Python actually loads and which native backends its ABI manifest records.
6. Regenerate and check wrappers
Generated wrapper files live under fhelium/native/wrapper/. After compiled schemas are available, generate or verify them through the generator rather than editing files manually:
python scripts/generate_native_wrappers.py \
--path fhelium/native/torchops
python scripts/generate_native_wrappers.py \
--path fhelium/native/torchops \
--check2
3
4
5
6
The direct script is the supported invocation; it is deliberately outside the runtime package so generating wrappers never imports a partially initialized fhelium.native.wrapper package. Callable wrappers retain require_native() guards, resolved lazily at call time. The generated FakeTensor registration module omits that import and is loaded by fhelium.native only after _ops has registered its schemas.
For a configured CMake tree, prefer its fresh-target check:
cmake --build <build-directory> --target native_wrappers_checkThat target depends on _ops, so the schemas being compared cannot come from an unrelated installed extension. Adjust paths only if a direct invocation uses a different compiled-op directory.
7. Run the validation ladder
At minimum run the focused tests and then the broader relevant suite. Include:
- output shape/dtype/device;
- mutation/alias behavior;
- fake/meta behavior where registered;
- depth zero, middle, and final legal depth;
- Q and QP;
- singleton row/digit;
- functional and in-place variants;
- CPU/CUDA parity for every shared schema;
- CPU thread-count coverage for parallel kernels;
- current-stream execution and asynchronous error localization on CUDA;
- multiple
logNvalues and NTT policies; - chained evaluator correctness;
- source build and installed wheel import.
8. Profile only after correctness
Profile the actual target shape and surrounding workload. Report whether a kernel change alters:
- launch count;
- global-memory traffic;
- register/shared-memory pressure;
- temporary/live memory;
- end-to-end operator or workload latency.
Do not claim an application speedup from a microbenchmark alone.
9. Keep generated and derived artifacts clean
Before review:
git status --short
git diff --check2
Ensure wrapper diffs are intentional, build products are not accidentally tracked, and source/build/wheel tests use matching commits.
10. Document the operation
Update:
- public docstring when the behavior is public;
- Developer Guide if architecture or invariants changed;
- tests describing edge cases;
- benchmark profile/report if performance policy changed;
- changelog/release notes when user-visible.
Record a native operation in a reviewable form such as:
Operation: mixed_radix_basis_extend_to_montgomery
Math:
Convert mixed-radix digits for one integer polynomial into residues
modulo every destination prime, preserving the polynomial element.
Input:
mixed_radix_components
shape [*batch, digit, coefficient]
signed integral dtype on one execution device
Tables and row mapping:
basis_extension_coefficients shape [digit - 1, destination_limb]
row r - 1 represents mixed-radix digit r >= 1
columns follow destination prime_ids
digit zero uses rns_params[R2] rather than a table row
Output:
shape [*batch, destination_limb, coefficient]
coefficient domain, Montgomery residues
Mutation and aliasing:
functional; output does not alias an input2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
The documented axes and table orientation must match the dispatcher schema and implementation. During review, trace the public semantic equation to each native tensor axis, verify prime-row and key-digit mappings, confirm rounding and residue-range laws against implementation and tests, and check mutation, aliasing, thread, and stream behavior. Use the definitions from Terminology and mathematical model.