Add bfloat16 based dot and conversion with single/double

1. Added bfloat16 based dot as new API: shdot 2. Implemented generic kernel and cooperlake-specific (AVX512-BF16) kernel for shdot 3. Added 4 conversion APIs for bfloat16 data type <=> single/double: shstobf16 shdtobf16 sbf16tos dbf16tod shstobf16 -- convert single float array to bfloat16 array shdtobf16 -- convert double float array to bfloat16 array sbf16tos -- convert bfloat16 array to single float array dbf16tod -- convert bfloat16 array to double float array 4. Implemented generic kernels for all 4 conversion APIs, and cooperlake-specific kernel for shstobf16 and shdtobf16 5. Update level1 thread facilitate functions and macros to support multi-threading for these new APIs 6. Fix Cooperlake platform detection/specify issue when under dynamic-arch building 7. Change the typedef of bfloat16 from unsigned short to more strict uint16_t Signed-off-by: Chen, Guobing <guobing.chen@intel.com>
2020-08-27 06:42:28 +08:00
parent c7ef7174e4
commit deaeb6c5b8
31 changed files with 1389 additions and 79 deletions
--- a/cmake/kernel.cmake
+++ b/cmake/kernel.cmake
@@ -126,12 +126,14 @@ if (BUILD_HALF)
  set(SHAXPYKERNEL ../arm/axpy.c)
  set(SHAXPBYKERNEL ../arm/axpby.c)
  set(SHCOPYKERNEL ../arm/copy.c)
-  set(SHDOTKERNEL ../arm/dot.c)
+  set(SHDOTKERNEL ../x86_64/shdot.c)
  set(SHROTKERNEL ../arm/rot.c)
  set(SHSCALKERNEL ../arm/scal.c)
  set(SHNRM2KERNEL ../arm/nrm2.c)
  set(SHSUMKERNEL ../arm/sum.c)
  set(SHSWAPKERNEL ../arm/swap.c)
+  set(TOBF16KERNEL ../x86_64/tobf16.c)
+  set(BF16TOKERNEL ../x86_64/bf16to.c)
 endif ()
 endmacro ()