README.md (2780B)
1 # keccak-fast 2 3 Drop-in-speed replacement for `coruus/keccak-tiny`. SHA-3/SHAKE one-shots, 4 TurboSHAKE128/256, and batch entry points. Plain C11, libc only, no build 5 flags, runtime-dispatched, thread-safe. 6 7 ## Use 8 9 ```c 10 #include <finwo/keccak-fast.h> 11 12 kf_shake256(out, sizeof out, in, inlen); /* 0, or -1 */ 13 kf_shake256_batch(count, in, inlen, out, outlen); /* equal-length, packed */ 14 ``` 15 16 The one-shots (`kf_shake128`, `kf_shake256`, `kf_sha3_224`, `kf_sha3_256`, 17 `kf_sha3_384`, `kf_sha3_512`, `kf_turboshake128`, `kf_turboshake256`) are 18 `kf_hash_fn` table entries: a call is one indirect jump. Batch variants are 19 `kf_<algo>_batch`, taking `count` equal-length messages packed back to back. 20 TurboSHAKE is the same sponge over Keccak-p[1600,12] with domain `0x1F`, about 21 1.9x faster than SHAKE; its batch entries use a separately registered 12-round 22 permutation. 23 24 ## Build 25 26 ``` 27 dep add finwo/keccak-fast 28 ``` 29 30 `export.mk` lists the sources; `.dep.export` installs `<finwo/keccak-fast.h>`. 31 32 ## Backends 33 34 | backend | one-shot | batch | 35 | ---------- | -------- | ----- | 36 | scalar | yes | — | 37 | scalar+bmi | yes | — | 38 | AVX2 | no | x4 | 39 | AVX-512 | no | x8 | 40 41 Backends register from constructors; the highest available priority wins 42 (scalar 0, scalar+bmi 10, avx2 20, avx512 30). `kf_backend_name()` and 43 `kf_batch_name()` report the active ones. `KECCAK_FAST_NO_BMI`, 44 `KECCAK_FAST_NO_AVX2` and `KECCAK_FAST_NO_AVX512` drop those backends. 45 46 The scalar backend is pinned to the x86-64 baseline 47 (`no-avx,no-avx2,no-avx512f,no-bmi,no-bmi2`) so a consumer's `-march=native` 48 cannot vectorize it or emit BMI. Its permutation keeps the 25 lanes in named 49 locals, fuses theta into the previous round's chi, and unrolls the round loop 50 4x; that beats both no unroll and a full 24x unroll. `scalar+bmi` is the same 51 permutation with `bmi` enabled, which turns chi's `not`+`and` into `andn`; gcc 52 also gets `rorx` for the rotates. About 20% faster than the donor. 53 54 The AVX2 and AVX-512 batch kernels hold the 4 or 8 messages of a group one per 55 vector lane, so the 25 state words live in 25 vector registers. The state is 56 transposed in 8x8 (or 4x4) blocks with unpack and shuffle instructions at the 57 sponge boundary. Their round bodies unroll theta and chi by hand so `-O2` emits 58 the same code as `-O3`. 59 60 ## Dev 61 62 ``` 63 make test # KAT and batch differential (TAP) 64 make bench # per-hash table: scalar, scalar+bmi, avx2, avx512 (+ donor) 65 ``` 66 67 `bench/` fetches the donor via `dep install`. `OPT=-O2|-O3` (arch flags 68 rejected), `BENCH_CPU=<n>` selects the core. 69 70 ## License 71 72 See `LICENSE.md`. The sponge is derived from `coruus/keccak-tiny` (David Leon 73 Gil, CC0); the scalar permutation follows XKCP's public-domain `opt64` 74 structure. 75