Martin Kroeker
aec353b5a7
Add a Windows/CL build to the Azure Ci configuration
2020-04-19 19:04:33 +02:00
Martin Kroeker
c62fbefad4
Merge pull request #2567 from xianyi/revert-2566-azurewin
...
Revert "Add Windows build job on Azure CI"
2020-04-19 19:01:58 +02:00
Martin Kroeker
04706e760d
Revert "Add Windows build job on Azure CI ( #2566 )"
...
This reverts commit e1e543b145
.
2020-04-19 19:00:37 +02:00
Martin Kroeker
e1e543b145
Add Windows build job on Azure CI ( #2566 )
...
* Add Windows-CL build job on Azure
2020-04-19 16:16:15 +02:00
Martin Kroeker
e55ec82bb9
Delete KERNEL.1004K
2020-04-19 15:44:30 +02:00
Martin Kroeker
7353ea5afc
Delete KERNEL.24K
2020-04-19 15:44:19 +02:00
Martin Kroeker
6a04efb122
Rename KERNEL files to include MIPS prefix
2020-04-19 15:43:54 +02:00
Martin Kroeker
5afb66812f
Update getarch.c
2020-04-19 14:55:31 +02:00
Martin Kroeker
0d18f231fc
Update getarch.c
2020-04-19 13:52:58 +02:00
Martin Kroeker
2f4a8e5bc4
Rename the FORCE entries for 24K and 1004K to include the MIPS prefix
2020-04-19 13:22:19 +02:00
Martin Kroeker
4f70512b97
Update kernel.cmake
2020-04-19 08:10:26 +02:00
Martin Kroeker
8792fc4d5f
Disable RPCC macro on MIPS24K
2020-04-19 07:21:48 +02:00
Martin Kroeker
577c5d9f8f
Update README.md
2020-04-19 06:54:52 +02:00
Martin Kroeker
6721f2750e
Update TargetList.txt
2020-04-19 06:51:57 +02:00
Martin Kroeker
b0b02a080d
Add compiler options for MIPS32 24K/1004K
2020-04-19 06:50:51 +02:00
Martin Kroeker
a1fc98dc57
rename 1004K, 24K to MIPS1004K, MIPS24K to avoid identifier naming problem
2020-04-18 23:50:23 +02:00
Martin Kroeker
d0737b0142
Update kernel.cmake
2020-04-18 21:36:28 +02:00
Martin Kroeker
7dbb59b256
Update common_macro.h
2020-04-18 21:34:14 +02:00
Martin Kroeker
00172d440b
Typo fix in MIPS24K addition
2020-04-18 21:16:49 +02:00
Martin Kroeker
d712ea724c
Add MIPS24K support
2020-04-18 21:10:18 +02:00
Martin Kroeker
61bbae3ac1
Handle MIPS24K like P5600
...
and allow enforcing TARGET=1004K as well (omission from earlier 1004K merge and later introduction of TARGET check)
2020-04-18 21:09:32 +02:00
Martin Kroeker
1c1ca2bc0a
Merge pull request #47 from xianyi/develop
...
rebase
2020-04-18 21:07:14 +02:00
Martin Kroeker
c7d668c248
Update common_macro.h
2020-04-18 16:04:38 +02:00
Martin Kroeker
a83a59b038
Use generic kernels for ishama,shasum,shdot,shrot
2020-04-18 15:53:51 +02:00
Martin Kroeker
0a19bd813c
Use generic codes for shamax and shcopy
2020-04-18 12:52:51 +02:00
Martin Kroeker
e7afe8a969
Define AXPBY_K fallback for float16
2020-04-18 11:10:15 +02:00
Martin Kroeker
f361de30a3
Use generic axpy.c for SHAXPY as x86 lacks saxpy.c
2020-04-18 11:07:16 +02:00
Martin Kroeker
9f6d6f6cb6
use saxpy.c instead of axpy.S for SHAXPY
2020-04-17 22:27:58 +02:00
Rajalakshmi Srinivasaraghavan
22bb50fb81
cmake fixes
2020-04-17 13:35:17 -05:00
Martin Kroeker
236a3d8ce6
Merge pull request #2563 from zelong-1024/develop
...
[OpenBLAS]: benchmark error of potrf
2020-04-16 11:45:32 +02:00
l00536773
6b7ef6543a
[OpenBLAS]: benchmark error of potrf
...
[description]: when the matrix size goes higher than 5800 during the cpotrf test, error info, such as "Potrf info = 5679", will be returned on ARM64 and x86 machines. Uplo = L & F.
[solution]: changed the func for building the matrix so that the complex Hermitian matrix can stay positive definite during the computation.
[dts]:
2020-04-16 10:55:10 +08:00
Rajalakshmi Srinivasaraghavan
67cc4b9e16
Fix warnings in clang and export symbol
2020-04-15 19:15:23 -05:00
Martin Kroeker
250e6f8039
Merge pull request #2557 from martin-frbg/dronebadge
...
Update and reformat README
2020-04-15 20:23:43 +02:00
Martin Kroeker
7a6d0016b0
Merge pull request #2556 from martin-frbg/epicdrone
...
Add a drone.io multithread test for x86_64
2020-04-15 20:23:17 +02:00
Martin Kroeker
e8e8a6e608
Restore USE_OPENMP in the x86 thread test
2020-04-15 19:26:12 +02:00
Martin Kroeker
579811fb6a
Move all 19.04-based jobs back to ubuntu 18.04
2020-04-15 17:38:33 +02:00
Rajalakshmi Srinivasaraghavan
a87793e03c
Fix DYNAMIC_ARCH compilation errors
2020-04-15 09:09:50 -05:00
Rajalakshmi Srinivasaraghavan
ac6a22ae78
Update header
2020-04-14 22:58:39 -05:00
Rajalakshmi Srinivasaraghavan
ff010f496e
Build shgemm for all architecture
2020-04-14 20:38:53 -05:00
Rajalakshmi Srinivasaraghavan
7eb55504b1
RFC : Add half precision gemm for bfloat16 in OpenBLAS
...
This patch adds support for bfloat16 data type matrix multiplication kernel.
For architectures that don't support bfloat16, it is defined as unsigned short
(2 bytes). Default unroll sizes can be changed as per architecture as done for
SGEMM and for now 8 and 4 are used for M and N. Size of ncopy/tcopy can be
changed as per architecture requirement and for now, size 2 is used.
Added shgemm in kernel/power/KERNEL.POWER9 and tested in powerpc64le and
powerpc64. For reference, added a small test compare_sgemm_shgemm.c to compare
sgemm and shgemm output.
This patch does not cover OpenBLAS test, benchmark and lapack tests for shgemm.
Complex type implementation can be discussed and added once this is approved.
2020-04-14 14:55:08 -05:00
Martin Kroeker
84a9614345
try x86_64 test without openmp
2020-04-14 19:18:35 +02:00
Martin Kroeker
b969533703
Add drone.io badge, mention EMAG8180 support, reformat the DYNAMIC_ARCH paragraph
2020-04-14 10:53:28 +02:00
Martin Kroeker
0f08f3efa6
Add a multithread test for x86_64
2020-04-13 22:46:12 +02:00
Martin Kroeker
c861b2a7bd
Merge pull request #2553 from martin-frbg/issue2444
...
Add a read memory barrier to the traversal of the buffer slot list
2020-04-13 21:28:59 +02:00
Martin Kroeker
cf62adffbb
Merge pull request #2555 from martin-frbg/issue1137
...
Handle unaligned data in the SSE2 copy kernel
2020-04-13 18:29:56 +02:00
Martin Kroeker
3eec7d382c
ARMV7 does not support DMB ISHLD, use DMB ISH
2020-04-13 15:56:31 +02:00
Martin Kroeker
5b0093b5fe
Convert aligned moves to unaligned
...
should have no performance impact on reasonably modern cpus and fixes occasional crashes in actual user code.
2020-04-13 14:58:52 +02:00
Martin Kroeker
f41600e66f
Add a read barrier in the traversing of the buffer list
...
Needed on systems with weak memory ordering - the inferior, partially working fix from #2544 was already removed in #2551
2020-04-13 12:34:02 +02:00
Martin Kroeker
f5efecb7ca
Add (empty) read barrier definition
2020-04-13 12:24:10 +02:00
Martin Kroeker
a52bdd9d7b
Add (empty) read barrier definition
2020-04-13 12:22:35 +02:00