bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
...
*	AMDGPU: Update for rsq intrinsic changes	Matt Arsenault	2016-07-15	1	-14/+12
\| \| \| \|	llvm-svn: 275622
*	AMDGPU: Add Clang Builtin for v_lerp_u8	Wei Ding	2016-07-15	1	-0/+2
\| \| \| \| \| \|	Differential Revision: http://reviews.llvm.org/D22380 llvm-svn: 275577
*	AMDGPU: Export workitem builtins	Jan Vesely	2016-07-10	1	-0/+28
\| \| \| \| \| \| \| \|	Reviewers: tstellardAMD Differential Revision: http://reviews.llvm.org/D20299 llvm-svn: 275030
*	[CodeGen] Use llvm::Type::getVectorNumElements instead of casting to ↵	Craig Topper	2016-07-08	1	-3/+2
\| \| \| \| \| \|	llvm::VectorType and calling getNumElements. This is equivalent and shorter. llvm-svn: 274823
*	[X86] Reuse existing lambda and remove unnecessary argument from vector cmp ↵	Craig Topper	2016-07-08	1	-16/+11
\| \| \| \| \| \|	builtin handling. NFC llvm-svn: 274821
*	[X86] Remove a couple calls to create V2F64 and V4F32 types for builtin ↵	Craig Topper	2016-07-08	1	-29/+17
\| \| \| \| \| \|	handling. Just get the type from the operand of the builtin instead. NFC llvm-svn: 274820
*	[X86] Use native IR for immediate values 0-7 of packed fp cmp builtins. This ↵	Craig Topper	2016-07-06	1	-0/+45
\| \| \| \| \| \|	makes them the same as what is done when using the SSE builtins for these same encodings. llvm-svn: 274608
*	[AVX512] Use the generic ctlz intrinsic to implement the vplzcntd/q builtins.	Craig Topper	2016-07-06	1	-0/+12
\| \| \| \|	llvm-svn: 274603
*	[OpenCL] An implementation of device side enqueue (DSE) from OpenCL v2.0 ↵	Anastasia Stulova	2016-07-05	1	-0/+142
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	s6.13.17. - Added new Builtins: enqueue_kernel, get_kernel_work_group_size and get_kernel_preferred_work_group_size_multiple. These Builtins use custom check to diagnose parameters of the passed Blocks i. e. variable number of 'local void*' type params, and check different overloads specified in Table 6.31 of OpenCL v2.0. - IR is generated as an internal library call for each OpenCL Builtin, reusing ObjC Block implementation. Review: http://reviews.llvm.org/D20249 llvm-svn: 274540
*	[OpenCL] Make OpenCL Builtins added according to the right version.	Anastasia Stulova	2016-07-04	1	-1/+1
\| \| \| \| \| \| \| \| \| \|	Currently we only have OpenCL 2.0 Builtins i.e. pipes or address space conversions. They have to be added only in the version 2.0 compilation mode to make the identifiers available for use in the other versions. Review: http://reviews.llvm.org/D20249 llvm-svn: 274509
*	[AVX512] Modify what indices we emit for the zero vector we use for zero ↵	Craig Topper	2016-07-04	1	-1/+1
\| \| \| \| \| \|	extension of the result of a v2i1 or v4i1 masked compare. This way we emit something that the backend easily interprets as a concatenation rather than a true shuffle. This delivers slightly better codegen with the current backend capabilities. llvm-svn: 274484
*	Emit more intrinsics for builtin functions	Matt Arsenault	2016-07-01	1	-39/+92
\| \| \| \| \| \| \| \| \| \| \| \|	This is important for building libclc. Since r273039 tests are failing due to now emitting calls to these functions instead of emitting the DAG node. The libm function names are implemented for OpenCL, and should call the locally defined versions, so -fno-builtin is used. The IR Some functions use the __builtins and expect the intrinsics to be emitted. Without this we end up with nobuiltin calls to intrinsics or to unsupported library calls. llvm-svn: 274370
*	[AVX512] Zero extend cmp intrinsic return value.	Igor Breger	2016-06-29	1	-2/+2
\| \| \| \| \| \|	Differential Revision: http://reviews.llvm.org/D21746 llvm-svn: 274110
*	AMDGPU: Add builtin to read exec mask	Matt Arsenault	2016-06-28	1	-4/+14
\| \| \| \|	llvm-svn: 273965
*	[AVX512] Replace masked integer cmp and ucmp builtins with native IR.	Craig Topper	2016-06-22	1	-7/+57
\| \| \| \|	llvm-svn: 273378
*	[X86][SSE4A] Use native IR for mask movntsd/movntss intrinsics.	Simon Pilgrim	2016-06-17	1	-0/+20
\| \| \| \| \| \|	Depends on llvm side commit r273002. llvm-svn: 273003
*	[ARM] Add mrrc/mrrc2 intrinsics and update existing mcrr/mcrr2 intrinsics.	Ranjeet Singh	2016-06-17	1	-0/+68
\| \| \| \| \| \| \| \| \| \|	Reapplying patch in r272777 which was reverted because the llvm patch which added support for generating the mcrr/mcrr2 instructions from the intrinsic was causing an assertion failure. This has now been fixed in llvm. llvm-svn: 272983
*	[x86] generate IR for AVX2 integer min/max builtins	Sanjay Patel	2016-06-16	1	-5/+17
\| \| \| \| \| \| \|	Sibling patch to r272932: http://reviews.llvm.org/rL272932 llvm-svn: 272933
*	[Builtin] Make __builtin_thread_pointer target-independent.	Marcin Koscielnicki	2016-06-16	1	-0/+7
\| \| \| \| \| \| \| \|	This is now supported for ARM, AArch64, PowerPC, SystemZ, SPARC, Mips. Differential Revision: http://reviews.llvm.org/D19589 llvm-svn: 272893
*	[x86] translate SSE packed FP comparison builtins to IR	Sanjay Patel	2016-06-15	1	-124/+74
\| \| \| \| \| \| \| \| \| \| \| \| \|	As noted in the code comment, a potential follow-on would be to remove the builtins themselves. Other than ord/unord, this already works as expected. Eg: typedef float v4sf __attribute__((__vector_size__(16))); v4sf fcmpgt(v4sf a, v4sf b) { return a > b; } Differential Revision: http://reviews.llvm.org/D21268 llvm-svn: 272840
*	[x86] generate IR for SSE integer min/max builtins	Sanjay Patel	2016-06-15	1	-0/+27
\| \| \| \| \| \| \|	Sibling patch to r272806: http://reviews.llvm.org/rL272806 llvm-svn: 272807
*	Reverting r272777 because one of the tests	Ranjeet Singh	2016-06-15	1	-68/+0
\| \| \| \| \| \| \|	added in the llvm patch is causing an assertion to fail. llvm-svn: 272790
*	[AVX512] Use native IR for mask pcmpeq/pcmpgt intrinsics.	Craig Topper	2016-06-15	1	-0/+49
\| \| \| \|	llvm-svn: 272787
*	[ARM] Add mrrc/mrrc2 intrinsics and update existing mcrr/mcrr2 intrinsics.	Ranjeet Singh	2016-06-15	1	-0/+68
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Patch adds intrinsics for mrrc/mrrc2. The intrinsics for mrrc/mrrc2 return a single uint64_t to represent two 32 bit values. The mcrr/mcrr2 intrinsic was changed to accept a single uint64_t instead of two 32 bit values as the input for consistency. Differential Revision: http://reviews.llvm.org/D21179 llvm-svn: 272777
*	Fix unused variable warning	Simon Pilgrim	2016-06-13	1	-51/+50
\| \| \| \|	llvm-svn: 272541
*	[Clang][X86] Convert non-temporal store builtins to generic ↵	Simon Pilgrim	2016-06-13	1	-63/+51
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	__builtin_nontemporal_store in headers We can now use __builtin_nontemporal_store instead of target specific builtins for naturally aligned nontemporal stores which avoids the need for handling in CGBuiltin.cpp The scalar integer nontemporal (unaligned) store builtins will have to wait as __builtin_nontemporal_store currently assumes natural alignment and doesn't accept the 'packed struct' trick that we use for normal unaligned load/stores. The nontemporal loads require further backend support before we can safely convert them to __builtin_nontemporal_load Differential Revision: http://reviews.llvm.org/D21272 llvm-svn: 272540
*	[CodeGen] Update to use an ArrayRef of uint32_t instead of int in calls to ↵	Craig Topper	2016-06-12	1	-10/+10
\| \| \| \| \| \|	CreateShuffleVector to match llvm interface change. llvm-svn: 272492
*	[X86] Handle AVX2 pslldqi and psrldqi intrinsics shufflevector creation ↵	Craig Topper	2016-06-09	1	-52/+0
\| \| \| \| \| \|	directly in the header file instead of in CGBuiltin.cpp. Simplify the sse2 equivalents as well. llvm-svn: 272246
*	[X86] Reuse the EmitX86Select routine to handle the select for masked ↵	Craig Topper	2016-06-09	1	-16/+7
\| \| \| \| \| \|	palignr too. llvm-svn: 272245
*	[AVX512] Emit select instruction instead of using x86 specific instrinsics.	Igor Breger	2016-06-08	1	-33/+59
\| \| \| \| \| \| \| \|	This will allow us to remove the x86 instrinics from the backend. Differential Revision: http://reviews.llvm.org/D21060 llvm-svn: 272141
*	[AVX512] Convert masked palignr builtins directly to native IR similar to ↵	Craig Topper	2016-06-06	1	-5/+23
\| \| \| \| \| \|	the other palignr builtins, but with a select to handle masking. llvm-svn: 271873
*	[AVX512] Convert masked load builtins to generic masked load intrinsics ↵	Craig Topper	2016-05-31	1	-0/+67
\| \| \| \| \| \| \| \|	instead of the x86 specific ones. This will allow the x86 intrinsics to be removed from the backend. llvm-svn: 271253
*	[AVX512] Emit generic masked store instrinsics instead of using x86 specific ↵	Craig Topper	2016-05-31	1	-0/+68
\| \| \| \| \| \| \| \|	intrinsics. This will allow us to remove the x86 instrinics from the backend. llvm-svn: 271246
*	[X86] Simplify alignr builtin support by recognizing that NumLaneElts is ↵	Craig Topper	2016-05-29	1	-9/+7
\| \| \| \| \| \|	always 16. NFC llvm-svn: 271176
*	[CodeGen] Use the ArrayRef form CreateShuffleVector instead of building ↵	Craig Topper	2016-05-29	1	-48/+39
\| \| \| \| \| \|	ConstantVectors or ConstantDataVectors and calling the other form. llvm-svn: 271165
*	AMDGPU: Add fract builtin	Matt Arsenault	2016-05-28	1	-0/+3
\| \| \| \|	llvm-svn: 271080
*	[CodeGen] Don't crash when sizeof(long) != 4 for some intrins	David Majnemer	2016-05-27	1	-6/+9
\| \| \| \| \| \| \| \| \| \|	_InterlockedIncrement and _InterlockedDecrement have 'long' in their prototypes. We assumed 'long' was the same size as an i32 which is incorrect for other targets. This fixes PR27892. llvm-svn: 270953
*	[OpenCL] Add to_{global\|local\|private} builtin functions.	Yaxun Liu	2016-05-20	1	-0/+23
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	OpenCL builtin functions to_{global\|local\|private} accepts argument of pointer type to arbitrary pointee type, and return a pointer to the same pointee type in different addr space, i.e. global gentype to_global(gentype p); It is not desirable to declare it as global void to_global(void ); in opencl header file since it misses diagnostics. This patch implements these builtin functions as Clang builtin functions. In the builtin def file they are defined to have signature void(void). When handling call expressions, their declarations are re-written to have correct parameter type and return type corresponding to the call argument. In codegen call to addr void to_addr(void) is generated with addrcasts or bitcasts to facilitate implementation in builtin library. Differential Revision: http://reviews.llvm.org/D19932 llvm-svn: 270261
*	Add all the avx512 flavors to __builtin_cpu_supports's list.	Benjamin Kramer	2016-05-20	1	-0/+21
\| \| \| \| \| \| \| \|	This is matching what trunk gcc is accepting. Also adds a missing ssse3 case. PR27779. The amount of duplication here is annoying, maybe it should be factored into a separate .def file? llvm-svn: 270224
*	[CUDA] Implement __ldg using intrinsics.	Justin Lebar	2016-05-19	1	-0/+45
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: Previously it was implemented as inline asm in the CUDA headers. This change allows us to use the [addr+imm] addressing mode when executing ld.global.nc instructions. This translates into a 1.3x speedup on some benchmarks that call this instruction from within an unrolled loop. Reviewers: tra, rsmith Subscribers: jhen, cfe-commits, jholewinski Differential Revision: http://reviews.llvm.org/D19990 llvm-svn: 270150
*	[WebAssembly] Rename memory_size intrinsic to current_memory	Derek Schuff	2016-05-02	1	-2/+2
\| \| \| \| \| \|	This follows the recent change in the wasm spec. llvm-svn: 268256
*	[AArch64] Fix D19098 fallout.	Marcin Koscielnicki	2016-04-19	1	-5/+0
\| \| \| \| \| \| \| \| \| \|	The intrinsic is now called llvm.thread.pointer, not llvm.aarch64.thread.pointer. Also, the code handling it in CGBuiltin.cpp is dead - it's already covered by GCCBuiltin. Remove it. Differential Revision: http://reviews.llvm.org/D19099 llvm-svn: 266817
*	[ARM NEON] Define vfms_f32 on ARM, and all vfms using vfma.	Ahmed Bougacha	2016-04-19	1	-16/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	r259537 added vfma/vfms to armv7, but the builtin was only lowered on the AArch64 side. Instead of supporting it on ARM, get rid of it. The vfms builtin lowered to: %nb = fsub float -0.0, %b %r = @llvm.fma.f32(%a, %nb, %c) Instead, define the operation in terms of vfma, and swap the multiplicands. It now lowers to: %na = fsub float -0.0, %a %r = @llvm.fma.f32(%na, %b, %c) This matches the instruction more closely, and lets current LLVM generate the "natural" operand ordering: fmls.2s v0, v1, v2 instead of the crooked (but equivalent): fmls.2s v0, v2, v1 Except for theses changes, assembly is identical. LLVM accepts both commutations, and the LLVM tests in: test/CodeGen/AArch64/arm64-fmadd.ll test/CodeGen/AArch64/fp-dp3.ll test/CodeGen/AArch64/neon-fma.ll test/CodeGen/ARM/fusedMAC.ll already check either the new one only, or both. Also verified against the test-suite unittests. llvm-svn: 266807
*	make __builtin_isfinite more efficient (PR27145)	Sanjay Patel	2016-04-07	1	-19/+12
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	isinf (is infinite) and isfinite should be implemented with the same function except we change the comparison operator. See PR27145 for more details: https://llvm.org/bugs/show_bug.cgi?id=27145 Ref: forked off of the discussion in D18513. Differential Revision: http://reviews.llvm.org/D18648 llvm-svn: 265675
*	NFC: make AtomicOrdering an enum class	JF Bastien	2016-04-06	1	-55/+52
\| \| \| \| \| \| \| \| \| \| \| \|	Summary: See LLVM change D18775 for details, this change depends on it. Reviewers: jyknight, reames Subscribers: cfe-commits Differential Revision: http://reviews.llvm.org/D18776 llvm-svn: 265569
*	AMDGPU: Add frexp_mant + frexp_exp builtins	Matt Arsenault	2016-03-30	1	-0/+8
\| \| \| \|	llvm-svn: 264960
*	Silencing warnings from MSVC 2015 Update 2. Both of these changes silence ↵	Aaron Ballman	2016-03-30	1	-1/+1
\| \| \| \| \| \|	"C4334 '<<': result of 32-bit shift implicitly converted to 64 bits (was 64-bit shift intended?)". NFC. llvm-svn: 264932
*	Add missing __builtin_bitreverse8	Matt Arsenault	2016-03-23	1	-0/+1
\| \| \| \| \| \|	Also add documentation for bitreverse builtins llvm-svn: 264203
*	[CUDA] Implement atomicInc and atomicDec builtins	Justin Lebar	2016-03-22	1	-0/+16
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	These functions cannot be implemented as atomicrmw or cmpxchg instructions, so they are implemented as a call to the NVVM intrinsics @llvm.nvvm.atomic.load.inc.32.p0i32 and @llvm.nvvm.atomic.load.dec.32.p0i32. Patch by Jason Henline. Reviewers: jlebar Differential Revision: http://reviews.llvm.org/D18322 llvm-svn: 264009
*	Preserve ExtParameterInfos into CGFunctionInfo.	John McCall	2016-03-11	1	-3/+1
\| \| \| \| \| \| \| \| \|	As part of this, make the function-arrangement interfaces a little simpler and more semantic. NFC. llvm-svn: 263191