keccak-fast.c

Minimal SIMD keccak implementation
git clone git://git.finwo.net/lib/keccak-fast.c
Log | Files | Refs | README | LICENSE

README.md (2780B)


      1 # keccak-fast
      2 
      3 Drop-in-speed replacement for `coruus/keccak-tiny`. SHA-3/SHAKE one-shots,
      4 TurboSHAKE128/256, and batch entry points. Plain C11, libc only, no build
      5 flags, runtime-dispatched, thread-safe.
      6 
      7 ## Use
      8 
      9 ```c
     10 #include <finwo/keccak-fast.h>
     11 
     12 kf_shake256(out, sizeof out, in, inlen);            /* 0, or -1 */
     13 kf_shake256_batch(count, in, inlen, out, outlen);   /* equal-length, packed */
     14 ```
     15 
     16 The one-shots (`kf_shake128`, `kf_shake256`, `kf_sha3_224`, `kf_sha3_256`,
     17 `kf_sha3_384`, `kf_sha3_512`, `kf_turboshake128`, `kf_turboshake256`) are
     18 `kf_hash_fn` table entries: a call is one indirect jump. Batch variants are
     19 `kf_<algo>_batch`, taking `count` equal-length messages packed back to back.
     20 TurboSHAKE is the same sponge over Keccak-p[1600,12] with domain `0x1F`, about
     21 1.9x faster than SHAKE; its batch entries use a separately registered 12-round
     22 permutation.
     23 
     24 ## Build
     25 
     26 ```
     27 dep add finwo/keccak-fast
     28 ```
     29 
     30 `export.mk` lists the sources; `.dep.export` installs `<finwo/keccak-fast.h>`.
     31 
     32 ## Backends
     33 
     34 | backend    | one-shot | batch |
     35 | ---------- | -------- | ----- |
     36 | scalar     | yes      | —     |
     37 | scalar+bmi | yes      | —     |
     38 | AVX2       | no       | x4    |
     39 | AVX-512    | no       | x8    |
     40 
     41 Backends register from constructors; the highest available priority wins
     42 (scalar 0, scalar+bmi 10, avx2 20, avx512 30). `kf_backend_name()` and
     43 `kf_batch_name()` report the active ones. `KECCAK_FAST_NO_BMI`,
     44 `KECCAK_FAST_NO_AVX2` and `KECCAK_FAST_NO_AVX512` drop those backends.
     45 
     46 The scalar backend is pinned to the x86-64 baseline
     47 (`no-avx,no-avx2,no-avx512f,no-bmi,no-bmi2`) so a consumer's `-march=native`
     48 cannot vectorize it or emit BMI. Its permutation keeps the 25 lanes in named
     49 locals, fuses theta into the previous round's chi, and unrolls the round loop
     50 4x; that beats both no unroll and a full 24x unroll. `scalar+bmi` is the same
     51 permutation with `bmi` enabled, which turns chi's `not`+`and` into `andn`; gcc
     52 also gets `rorx` for the rotates. About 20% faster than the donor.
     53 
     54 The AVX2 and AVX-512 batch kernels hold the 4 or 8 messages of a group one per
     55 vector lane, so the 25 state words live in 25 vector registers. The state is
     56 transposed in 8x8 (or 4x4) blocks with unpack and shuffle instructions at the
     57 sponge boundary. Their round bodies unroll theta and chi by hand so `-O2` emits
     58 the same code as `-O3`.
     59 
     60 ## Dev
     61 
     62 ```
     63 make test    # KAT and batch differential (TAP)
     64 make bench   # per-hash table: scalar, scalar+bmi, avx2, avx512 (+ donor)
     65 ```
     66 
     67 `bench/` fetches the donor via `dep install`. `OPT=-O2|-O3` (arch flags
     68 rejected), `BENCH_CPU=<n>` selects the core.
     69 
     70 ## License
     71 
     72 See `LICENSE.md`. The sponge is derived from `coruus/keccak-tiny` (David Leon
     73 Gil, CC0); the scalar permutation follows XKCP's public-domain `opt64`
     74 structure.
     75