Werner Saar
|
7316a87930
|
added optimized dswap kernel for POWER8
|
2016-03-25 14:35:43 +01:00 |
Werner Saar
|
0bff057a87
|
added optimized dcopy kernel for POWER8
|
2016-03-25 13:03:02 +01:00 |
wernsaar
|
7ee1d29dd4
|
Merge pull request #822 from wernsaar/develop
added optimized dscal kernel for POWER8
|
2016-03-25 10:15:51 +01:00 |
Werner Saar
|
1e6cf9808c
|
added optimized dscal kernel for POWER8
|
2016-03-25 09:42:08 +01:00 |
Ashwin Sekhar T K
|
278511ad2d
|
Cortex-A57: Fix clang compilation errors
|
2016-03-24 10:42:04 +05:30 |
Ashwin Sekhar T K
|
3b5ffb49d3
|
Cortex-A57: Improve DGEMM 8x4 Implementation
|
2016-03-24 10:25:18 +05:30 |
wernsaar
|
8519e4ed9f
|
Merge pull request #817 from wernsaar/develop
added optimized zaxpy kernel for POWER8
|
2016-03-23 13:37:04 +01:00 |
Werner Saar
|
55eda3813b
|
added optimized zaxpy kernel for POWER8
|
2016-03-23 11:20:23 +01:00 |
Zhang Xianyi
|
53bfc83c26
|
Update appveyor version.
|
2016-03-22 11:37:35 -04:00 |
Zhang Xianyi
|
13ca89f6f0
|
Merge pull request #813 from theoractice/develop
Fix access violation on Windows while static linking in MSVC
|
2016-03-22 11:31:37 -04:00 |
wernsaar
|
461cf9ea38
|
Merge pull request #814 from wernsaar/develop
added optimized daxpy kernel for POWER8
|
2016-03-22 15:24:59 +01:00 |
Werner Saar
|
0664ba4c97
|
added optimized daxpy kernel for POWER8
|
2016-03-22 14:50:03 +01:00 |
Theoractice
|
aa744dfa59
|
Update memory.c
|
2016-03-22 20:02:37 +08:00 |
theoractice
|
61cf8f74d9
|
Fix access violation on Windows while static linking
|
2016-03-22 19:14:54 +08:00 |
Theoractice
|
de202fa375
|
Merge pull request #1 from xianyi/develop
upd
|
2016-03-22 05:33:20 -05:00 |
wernsaar
|
6f93b53590
|
Merge pull request #812 from wernsaar/develop
added optimized sdot kernel for POWER8
|
2016-03-21 13:59:44 +01:00 |
Werner Saar
|
11c44dede1
|
added optimized sdot kernel for POWER8
|
2016-03-21 13:18:23 +01:00 |
wernsaar
|
f00d642592
|
Merge pull request #811 from wernsaar/develop
added optimized zdot kernel for POWER8
|
2016-03-21 10:48:41 +01:00 |
Werner Saar
|
9e4584d069
|
added optimized zdot kernel for POWER8
|
2016-03-21 10:12:07 +01:00 |
Zhang Xianyi
|
2a5679da5f
|
Merge branch 'release-0.2.17' into develop
|
2016-03-20 20:52:43 -04:00 |
Zhang Xianyi
|
a71e8c82f6
|
Fix change log typo.
|
2016-03-20 20:52:15 -04:00 |
Zhang Xianyi
|
9b987badb0
|
Merge branch 'master' into develop
Bump to 0.2.18.dev
Conflicts:
CMakeLists.txt
Makefile.rule
|
2016-03-20 20:48:21 -04:00 |
Zhang Xianyi
|
1619b2f3c8
|
Merge branch 'release-0.2.17'
|
2016-03-20 20:44:01 -04:00 |
Zhang Xianyi
|
4f3153395a
|
Update doc for 0.2.17.
|
2016-03-20 20:43:42 -04:00 |
Zhang Xianyi
|
d7a1a7ff2a
|
Merge branch 'release-0.2.17' into develop
|
2016-03-20 09:24:28 -04:00 |
Zhang Xianyi
|
308e6195b7
|
Refs #807. Enable BUILD_LAPACK_DEPRECATED=1 by default.
|
2016-03-20 09:22:56 -04:00 |
Zhang Xianyi
|
7a3d7b1f52
|
Merge pull request #808 from theoractice/develop
Fix a minor compiler error in VisualStudio with CMake
|
2016-03-20 09:07:47 -04:00 |
wernsaar
|
74cc2d6623
|
Merge pull request #809 from wernsaar/develop
Ref #795: added optimized ddot kernel for POWER8
|
2016-03-20 13:16:41 +01:00 |
theoractice
|
fc3a558515
|
Fix a minor compiler error in VisualStudio with CMake
|
2016-03-20 18:58:18 +08:00 |
Werner Saar
|
cd9fafc054
|
ddot for POWER8: updated licence information
|
2016-03-20 11:19:27 +01:00 |
Werner Saar
|
84b92e6373
|
added optimized ddot kernel for POWER8
|
2016-03-20 11:06:06 +01:00 |
wernsaar
|
c279a53ed8
|
Merge pull request #806 from wernsaar/develop
adding optimized single precision blas level3 kernels for POWER8
|
2016-03-18 12:46:16 +01:00 |
Werner Saar
|
e1df5a6e23
|
fixed sgemm- and strmm-kernel
|
2016-03-18 12:12:03 +01:00 |
Werner Saar
|
5c658f8746
|
add optimized cgemm- and ctrmm-kernel for POWER8
|
2016-03-18 08:17:25 +01:00 |
Zhang Xianyi
|
ec4390a967
|
Bump devlop version to 0.2.17.dev.
|
2016-03-15 14:52:01 -04:00 |
Zhang Xianyi
|
fced5744fb
|
Merge branch 'release-0.2.16'
|
2016-03-15 14:49:10 -04:00 |
Zhang Xianyi
|
8c0fb1258d
|
Update 0.2.16 doc
|
2016-03-15 14:48:41 -04:00 |
Zhang Xianyi
|
aae581d004
|
Merge branch 'develop' into release-0.2.16
|
2016-03-15 13:56:01 -04:00 |
Zhang Xianyi
|
e17303933a
|
Merge pull request #802 from ashwinyes/develop_20160314_dgemm_optimization
DGEMM Optimizations for Cortex-A57
|
2016-03-14 20:31:03 -04:00 |
Zhang Xianyi
|
f9226275f4
|
Merge pull request #801 from Keno/patch-3
Don't pass REALNAME to `.end`
|
2016-03-14 15:42:31 -04:00 |
Ashwin Sekhar T K
|
cf8c7e28b3
|
Update CONTRIBUTORS.md
|
2016-03-14 20:01:02 +05:30 |
Ashwin Sekhar T K
|
5ac02f6dc7
|
Optimize Dgemm 4x4 for Cortex A57
|
2016-03-14 19:35:23 +05:30 |
Ashwin Sekhar T K
|
7aa1ad4923
|
Functional Assembly Kernels for CortexA57
Adding functional (non-optimized) kernels for Cortex-A57
with the following layouts.
SGEMM - 16x4, 8x8
CGEMM - 8x4
DGEMM - 8x4, 4x8
|
2016-03-14 19:33:21 +05:30 |
Werner Saar
|
dcd15b546c
|
BUGFIX: KERNEL.POWER8
|
2016-03-14 14:36:59 +01:00 |
Werner Saar
|
96284ab295
|
added sgemm- and strmm-kernel for POWER8
|
2016-03-14 13:52:44 +01:00 |
Keno Fischer
|
d5e1255ca7
|
Don't pass REALNAME to `.end`
Putting the procedure there is an MSVC-ism, where it is optional. GCC silently ignores and Clang errors, so it is best to remove this.
|
2016-03-13 18:56:21 -04:00 |
Zhang Xianyi
|
587455868e
|
Merge pull request #800 from jeromerobert/smallscaling
Fix smallscaling compilation
|
2016-03-10 15:45:33 -05:00 |
Jerome Robert
|
323c237e7b
|
Fix smallscaling compilation
Also revert 0bbca5e
|
2016-03-10 20:24:41 +01:00 |
Werner Saar
|
faa5e2e5e3
|
FIX: forgot the add the files cgemv_n_4.c and cgemv_t_4.c
|
2016-03-10 11:10:38 +01:00 |
wernsaar
|
551fdf53e8
|
Merge pull request #799 from wernsaar/develop
Added optimized cgemv_n and cgemv_t kernels for bulldozer, piledriver…
|
2016-03-10 10:22:08 +01:00 |