OpenBLAS

Files

T

Chen, Guobing deaeb6c5b8 Add bfloat16 based dot and conversion with single/double

1. Added bfloat16 based dot as new API: shdot
2. Implemented generic kernel and cooperlake-specific (AVX512-BF16) kernel for shdot
3. Added 4 conversion APIs for bfloat16 data type <=> single/double: shstobf16 shdtobf16 sbf16tos dbf16tod
     shstobf16 -- convert single float array to bfloat16 array
     shdtobf16 -- convert double float array to bfloat16 array
     sbf16tos  -- convert bfloat16 array to single float array
     dbf16tod  -- convert bfloat16 array to double float array
4. Implemented generic kernels for all 4 conversion APIs, and cooperlake-specific kernel for shstobf16 and shdtobf16
5. Update level1 thread facilitate functions and macros to support multi-threading for these new APIs
6. Fix Cooperlake platform detection/specify issue when under dynamic-arch building
7. Change the typedef of bfloat16 from unsigned short to more strict uint16_t

Signed-off-by: Chen, Guobing <guobing.chen@intel.com>

2020-09-04 02:31:25 +08:00

alpha

Add implementations of ssum/dsum and csum/zsum

2019-03-30 22:05:11 +01:00

arm

Use OPENBLAS_MAKE_COMPLEX_FLOAT on PPC only

2020-07-23 20:40:13 +00:00

arm64

ARM64: Add THUNDERX3T110 Target

2020-07-26 23:32:24 -07:00

generic

powerpc: Optimized SHGEMM kernel for POWER10