OpenBLAS

Author	SHA1	Message	Date
maamountki	82124729af	Merge branch 'develop' into z14	2019-01-31 19:36:41 +02:00
maamountki	29416cb5a3	[ZARCH] Add Z13 version for max/min functions	2019-01-31 19:11:11 +02:00
maamountki	48b9b94f7f	[ZARCH] Improve loading performance for camax/icamax	2019-01-31 18:52:11 +02:00
Martin Kroeker	86a824c97f	Fix wrong comparison that made IMIN identical to IMAX as reported by aarnez in #1990	2019-01-31 15:27:21 +01:00
Martin Kroeker	808410c2c7	Fix wrong comparison that made IMIN identical to IMAX as suggested in #1990	2019-01-31 15:25:15 +01:00
maamountki	fcd814a8d2	[ZARCH] Fix bug in max/min functions	2019-01-29 17:59:38 +02:00
maamountki	dc4d3bccd5	[ZARCH] Fix icamax/icamin	2019-01-29 03:47:49 +02:00
maamountki	c7143c1019	[ZARCH] Fix iamax/imax single precision	2019-01-28 17:52:23 +02:00
maamountki	04873bb174	[ZARCH] Undo the last commit	2019-01-28 17:32:24 +02:00
maamountki	c8ef9fb220	[ZARCH] Fix bug in iamax/iamin/imax/imin	2019-01-28 17:16:18 +02:00
maamountki	b111829226	[ZARCH] Update max/min functions	2019-01-21 15:56:04 +02:00
Martin Kroeker	32b0f1168e	Fix declaration of input arguments in the Sandybridge GER microkernels (#1967 ) * Tag arguments 0 and 1 as both input and output	2019-01-18 08:11:39 +01:00
Martin Kroeker	b495e54310	Fix declaration of input arguments in the x86_64 SCAL microkernels (#1966 ) * Tag arguments 0 and 1 as both input and output (see #1964)	2019-01-18 08:11:07 +01:00
Martin Kroeker	d5e6940253	Fix declaration of input arguments in the x86_64 microkernels for DOT and AXPY (#1965 ) * Tag operands 0 and 1 as both input and output For #1964 (basically a continuation of coding problems first seen in #1292)	2019-01-17 23:20:32 +01:00
Ubuntu	43a4572038	crot fix	2019-01-17 14:45:31 +00:00
Abdelrauf	a034e65512	Merge branch 'develop' into develop	2019-01-16 19:25:13 +04:00
Ubuntu	8c3386be87	Added missing Blas1 single fp {saxpy, caxpy, cdot, crot(refactored version of srot),isamax ,isamin, icamax, icamin}, Fixed idamin,icamin choosing the first occurance index of equal minimals	2019-01-16 15:16:21 +00:00
maamountki	b815a04c87	[ZARCH] fix a bug in max/min functions	2019-01-15 21:04:22 +02:00
maamountki	1a7925b3a3	[ZARCH] Update dgemv_n_4.c	2019-01-11 17:43:11 +02:00
maamountki	406f835f00	[ZARCH] update cgemv_n_4.c	2019-01-11 17:39:17 +02:00
maamountki	621dedb37b	[ZARCH] Update cgemv_t_4.c	2019-01-11 17:37:11 +02:00
maamountki	b731e8246f	Update sgemv_t_4.c	2019-01-11 17:14:04 +02:00
maamountki	ecc31b743f	Update dgemv_t_4.c	2019-01-11 17:13:02 +02:00
maamountki	5d89d6b143	[ZARCH] fix sgemv_n_4.c	2019-01-11 17:08:24 +02:00
maamountki	67432b23c2	[ZARCH] fix cgemv_n_4.c	2019-01-11 16:44:46 +02:00
maamountki	be66f5d5c2	[ZARCH] fix data prefetch type in sdot	2019-01-09 16:50:07 +02:00
maamountki	c2ffef8156	[ZARCH] fix data prefetch type in ddot	2019-01-09 16:49:44 +02:00
maamountki	e7455f500c	[ZARCH] fix dsdot.c	2019-01-09 16:33:54 +02:00
maamountki	3eafcfa650	[ZARCH] fix cgemv_n_4.c	2019-01-09 07:43:45 +02:00
maamountki	94cd946b96	[ZARCH] fix cgemv_n_4.c	2019-01-04 17:45:56 +02:00
maamountki	1aa840a0a2	[ZARCH] fix sgemv_t_4.c	2019-01-04 01:38:18 +02:00
Arjan van de Ven	795285c587	Fix thinko in skylake beta handling casting ints is cheaper but it has a rounding, not memory casing effect, resulting in invalid outcome	2018-12-24 18:49:50 +00:00
Arjan van de Ven	d321448a63	dgemm: use dgemm_ncopy_8_skylakex.c also for Haswell The dgemm_ncopy_8_skylakex.c code is not avx512 specific and gives a nice performance boost for medium sized matrices	2018-12-16 23:09:22 +00:00
Arjan van de Ven	c43331ad0a	dgemm: Use the skylakex beta function also for haswell it's more efficient for certain tall/skinny matrices	2018-12-16 23:09:17 +00:00
Martin Kroeker	c4e23dd016	Update Makefile	2018-12-16 18:14:40 +01:00
Martin Kroeker	cfc4acc221	typo	2018-12-16 16:19:51 +01:00
Martin Kroeker	545c2b1bbb	Add -mavx2 on Haswell only if the compiler supports it	2018-12-16 13:09:19 +01:00
Arjan van de Ven	69d206440a	Make the skylakex/haswell sgemm code compile and run even with compilers without avx2 support	2018-12-16 00:19:41 +00:00
Martin Kroeker	3843e3e017	use -maxv2 on haswell	2018-12-15 23:30:31 +01:00
Martin Kroeker	fbcb14a74b	should be core-avx2	2018-12-15 20:18:59 +01:00
Martin Kroeker	2a3190dc76	fix elseifeq and use older option core2-avx for compatibility	2018-12-15 20:17:44 +01:00
Martin Kroeker	1ebe5c0f49	Add -march=haswell to HASWELL part of DYNAMIC_ARCH build	2018-12-15 19:35:35 +01:00
Arjan van de Ven	0586899a10	Use sgemm_ncopy_4_skylakex.c also for Haswell sgemm_ncopy_4_skylakex.c uses SSE transpose operations where the real perf win happens; this also works great for Haswell. This gives double digit percentage gains on small and skinny matrices	2018-12-15 13:49:19 +00:00
Arjan van de Ven	00dc09ad19	Use the skylake sgemm beta code also for haswell with a few small changes it's possible to use the skylake sgemm code also for haswell, this gives a modest gain (10% range) for smallish matrixes but does wonders for very skinny matrixes	2018-12-15 13:49:13 +00:00
Arjan van de Ven	cdc668d82b	Add a "sgemm direct" mode for small matrixes OpenBLAS has a fancy algorithm for copying the input data while laying it out in a more CPU friendly memory layout. This is great for large matrixes; the cost of the copy is easily ammortized by the gains from the better memory layout. But for small matrixes (on CPUs that can do efficient unaligned loads) this copy can be a net loss. This patch adds (for SKYLAKEX initially) a "sgemm direct" mode, that bypasses the whole copy machinary for ALPHA=1/BETA=0/... standard arguments, for small matrixes only. What is small? For the non-threaded case this has been measured to be in the MNK = 28 * 512 * 512 range, while in the threaded case it's less, around MNK = 1 * 512 * 512	2018-12-13 13:47:31 +00:00
Martin Kroeker	87718807f0	Merge pull request #1910 from martin-frbg/issue1909 Fix for DYNAMIC_ARCH builds made on a AVX512-capable host	2018-12-12 14:56:25 +01:00
Martin Kroeker	51aec8e96b	make sure the added march=skylake-avx512 does not cause problems on Windows	2018-12-11 22:47:32 +01:00
Martin Kroeker	06f7d78d70	Add -march=skylake-avx512 to SkylakeX part of DYNAMIC_ARCH builds	2018-12-11 21:10:38 +01:00
Martin Kroeker	7639f2e1f0	Rewrite the conditional for OSX to fix cmake parsing on others The Makefile variable parser in utils.cmake currently does not handle conditionals. Having the definitions for non-OSX last will at least make cmake builds work again on non-OSX platforms.	2018-12-06 14:04:27 +01:00
Martin Kroeker	2fc712469d	Avoid creating spurious non-suffixed c/zgemm_kernels Plain cgemm_kernel and zgemm_kernel are not used anywhere, only cgemm_kernel_b etc. Needlessly building them (without any define like NN, CN, etc.) just happened to work on most platforms, but not on arm64. See #1870	2018-12-06 13:56:06 +01:00

1 2 3 4 5 ...

1103 Commits