bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
...
*	Do not attribute static allocas to the call site's DebugLoc.	Paul Robinson	2014-10-21	1	-0/+6
\| \| \| \| \| \| \| \| \| \| \| \| \|	When functions are inlined, instructions without debug information are attributed to the call site's DebugLoc. After inlining, inlined static allocas are moved to the caller's entry block, adjacent to the caller's original static alloca instructions. By retaining the call site's DebugLoc, these instructions could cause instructions that were subsequently inserted at the entry block to pick up the same DebugLoc. Patch by Wolfgang Pieb! llvm-svn: 220255
*	PR21202: Memory leak in Windows RWMutexImpl when using SRWLOCK	David Blaikie	2014-10-21	1	-5/+3
\| \| \| \|	llvm-svn: 220251
*	Make AsmPrinter::EmitLabelOffsetDifference a static helper and simplify.	Rafael Espindola	2014-10-21	2	-33/+27
\| \| \| \| \| \|	It had exactly one caller in a position where we know hasSetDirective is true. llvm-svn: 220250
*	[MCJIT] Temporarily revert r220245 - it broke several bots.	Lang Hames	2014-10-21	2	-8/+1
\| \| \| \| \| \|	(See e.g. http://bb.pgr.jp/builders/cmake-llvm-x86_64-linux/builds/17653) llvm-svn: 220249
*	Introduce enum values for previously defined metadata types. (NFC)	Philip Reames	2014-10-21	7	-14/+29
\| \| \| \| \| \| \| \| \| \| \|	Our metadata scheme lazily assigns IDs to string metadata, but we have a mechanism to preassign them as well. Using a preassigned ID is helpful since we get compile time type checking, and avoid some (minimal) string construction and comparison. This change adds enum value for three existing metadata types: + MD_nontemporal = 9, // "nontemporal" + MD_mem_parallel_loop_access = 10, // "llvm.mem.parallel_loop_access" + MD_nonnull = 11 // "nonnull" I went through an updated various uses as well. I made no attempt to get all uses; I focused on the ones which were easily grepable and easily to translate. For example, there were several items in LoopInfo.cpp I chose not to update. llvm-svn: 220248
*	Extend the verifier to validate range metadata on calls and invokes.	Philip Reames	2014-10-20	1	-49/+56
\| \| \| \| \| \|	Range metadata applies to loads, call, and invokes. We were validating that metadata applied to loads was correct according to the LangRef, but we were not validating metadata applied to calls or invokes. This change extracts the checking functionality to a common location, reuses it for all valid locations, and adds a simple test to ensure a misused range on a call gets reported. llvm-svn: 220246
*	[MCJIT] Make MCJIT honor symbol visibility settings when populating the global	Lang Hames	2014-10-20	2	-1/+8
\| \| \| \| \| \| \| \|	symbol table. Patch by Anthony Pesch. Thanks Anthony! llvm-svn: 220245
*	[X86] Fix a bug in the lowering of the mask of VSELECT.	Quentin Colombet	2014-10-20	1	-1/+6
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	X86 code to lower VSELECT messed a bit with the bits set in the mask of VSELECT when it knows it can be lowered into BLEND. Indeed, only the high bits need to be set for those and it optimizes those accordingly. However, when the mask is a compile time constant, the lowering will be handled by the generic optimizer and those modifications will generate bad code in the generic optimizer. This patch fixes that by preventing the optimization if the VSELECT will be handled by the generic optimizer. <rdar://problem/18675020> llvm-svn: 220242
*	Introduce a 'nonnull' metadata on Load instructions.	Philip Reames	2014-10-20	1	-0/+4
\| \| \| \| \| \| \| \| \|	The newly introduced 'nonnull' metadata is analogous to existing 'nonnull' attributes, but applies to load instructions rather than call arguments or returns. Long term, it would be nice to combine these into a single construct. The value of the load is allowed to vary between successive loads, but null is not a valid value to be loaded by any load marked nonnull. Reviewed by: Hal Finkel Differential Revision: http://reviews.llvm.org/D5220 llvm-svn: 220240
*	[X86] Memory folding for commutative instructions (updated)	Simon Pilgrim	2014-10-20	3	-59/+102
\| \| \| \| \| \| \| \| \| \| \| \|	This patch improves support for commutative instructions in the x86 memory folding implementation by attempting to fold a commuted version of the instruction if the original folding fails - if that folding fails as well the instruction is 're-commuted' back to its original order before returning. Updated version of r219584 (reverted in r219595) - the commutation attempt now explicitly ensures that neither of the commuted source operands are tied to the destination operand / register, which was the source of all the regressions that occurred with the original patch attempt. Added additional regression test case provided by Joerg Sonnenberger. Differential Revision: http://reviews.llvm.org/D5818 llvm-svn: 220239
*	ARM: rework Thumb1 frame index rewriting	Tim Northover	2014-10-20	5	-112/+28
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The previous code had a few problems, motivating the choices here. 1. It could create instructions clobbering CPSR, but the incoming MachineInstr didn't reflect this. A potential source of corruption. This is why the patch has a new PseudoInst for before lowering. 2. Similarly, there was some code to handle the incoming instruction not being ARMCC::AL, but this would have caused massive problems if it was actually invoked when a complex offset needing more than one instruction was requested. 3. It wasn't designed to handle unaligned pointers (or offsets). These should probably be minimised anyway, but the code needs to deal with them properly regardless. 4. It had some rather dubious ad-hoc code to avoid calling emitThumbRegPlusImmediate, a function which should be designed to do precisely this job. We seem to cover the common cases correctly now, and hopefully can enhance emitThumbRegPlusImmediate to handle any extra optimisations we need to add in future. llvm-svn: 220236
*	Be more specific about return type of MachOUniversalBinary::getObjectForArch	Alexey Samsonov	2014-10-20	1	-2/+2
\| \| \| \|	llvm-svn: 220230
*	Constify input argument of RelocVisitor and DWARFContext constructors. NFC.	Alexey Samsonov	2014-10-20	3	-3/+3
\| \| \| \|	llvm-svn: 220228
*	Moved out IIT_V64 from common values section.	Robert Khasanov	2014-10-20	1	-5/+5
\| \| \| \| \| \|	Thanks Juergen Ributzka for notice. llvm-svn: 220224
*	Fix Intrinsic::getType not working with vararg	Steven Wu	2014-10-20	1	-0/+6
\| \| \| \| \| \| \| \|	VarArg Intrinsic functions are encoded with "void" type as the last argument. Now Intrinsic::getType can correctly return all the intrinsic function type. llvm-svn: 220205
*	[Thumb2] RFE, SRS and "SUBS pc, lr" are undefined on v7M	Oliver Stannard	2014-10-20	1	-3/+5
\| \| \| \| \| \| \|	These instructions are related to the v7[AR] exception model, and are not defined on v7M. llvm-svn: 220204
*	Remove unnecessary else.	Sid Manning	2014-10-20	1	-4/+4
\| \| \| \|	llvm-svn: 220200
*	[ARM] Do not select SMULW[BT] or SMLAW[BT]	Oliver Stannard	2014-10-20	2	-30/+8
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The current instruction selection patterns for SMULW[BT] and SMLAW[BT] are incorrect. These instructions multiply a 32-bit and a 16-bit value (both signed) and return the top 32 bits of the 48-bit result. This preserves the 16 bits of overflow, whereas the patterns they currently match truncate the result to 16 bits then sign extend. To select these instructions, we would need to match an ISD::SMUL_LOHI, a sign extend, two shifts and an or. There is no way to match SMUL_LOHI in an instruction pattern as it defines multiple values, so this would have to be done in C++. I have raised http://llvm.org/bugs/show_bug.cgi?id=21297 to cover allowing correct selection of these instructions. This fixes http://llvm.org/bugs/show_bug.cgi?id=19396 llvm-svn: 220196
*	[Thumb] Fix crash in Thumb1RegisterInfo::rewriteFrameIndex	Oliver Stannard	2014-10-20	1	-0/+1
\| \| \| \| \| \| \| \| \|	This function can, for some offsets from the SP, split one instruction into two. Since it re-uses the original instruction as the first instruction of the result, we need ensure its result register is not marked as dead before we use it in the second instruction. llvm-svn: 220194
*	Switch the default DataLayout to be little endian, and make the variable	Chandler Carruth	2014-10-20	1	-5/+5
\| \| \| \| \| \| \| \| \|	be BigEndian so the default can continue to be zero-initialized. This is one of the prerequisites to making DataLayout a constant and always available part of every module. llvm-svn: 220193
*	Fix a miscompile introduced in r220178.	Chandler Carruth	2014-10-20	1	-4/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The original code had an implicit assumption that if the test for allocas or globals was reached, the two pointers were not equal. With my changes to make the pointer analysis more powerful here, I also had to guard against circumstances where the results weren't useful. That in turn violated the assumption and gave rise to a circumstance in which we could have a store with both the queried pointer and stored pointer rooted at the same alloca. Clearly, we cannot ignore such a store. There are other things we might do in this code to better handle the case of both pointers ending up at the same alloca or global, but it seems best to at least make the test explicit in what it intends to check. I've added tests for both the alloca and global case here. llvm-svn: 220190
*	IR: Replace DataLayout::RoundUpAlignment with RoundUpToAlignment	David Majnemer	2014-10-20	3	-8/+8
\| \| \| \| \| \|	No functional change intended, just cleaning up some code. llvm-svn: 220187
*	Fix a somewhat subtle pair of issues with JumpThreading I introduced in	Chandler Carruth	2014-10-20	1	-3/+6
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	r220178. First, the creation routine doesn't insert prior to the terminator of the basic block provided, but really at the end of the basic block. Instead, get the terminator and insert before that. The next issue was that we need to ensure multiple PHI node entries for a single predecessor re-use the same cast instruction rather than creating new ones. All of the logic here was without tests previously. I've reduced and added a test case from the test suite that crashed without both of these fixes. llvm-svn: 220186
*	Teach the load analysis driving core instcombine logic and other bits of	Chandler Carruth	2014-10-20	3	-23/+55
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	logic to look through pointer casts, making them trivially stronger in the face of loads and stores with intervening pointer casts. I've included a few test cases that demonstrate the kind of folding instcombine can do without pointer casts and then variations which obfuscate the logic through bitcasts. Without this patch, the variations all fail to optimize fully. This is more important now than it has been in the past as I've started moving the load canonicialization to more closely follow the value type requirements rather than the pointer type requirements and thus this needs to be prepared for more pointer casts. When I made the same change to stores several test cases regressed without logic along these lines so I wanted to systematically improve matters first. llvm-svn: 220178
*	Do a better and more complete job of preserving metadata when combining	Chandler Carruth	2014-10-19	1	-8/+58
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	loads. This handles many more cases than just the AA metadata, some of them suggested by Hal in his review of the AA metadata handling patch. I've tried to test this behavior where tractable to do so. I'll point out that I have specifically not included a test for debuginfo because it was going to require 2 or 3 times as much work to craft some input which would survive the "helpful" stripping of debug info metadata that doesn't match the desired schema. This is another good example of why the current state of write-ability for our debug info metadata is unacceptable. I spent over 30 minutes trying to conjure some test case that would survive, even copying from other debug info tests, but it always failed to survive with no explanation of why or how I might fix it. =[ llvm-svn: 220165
*	Move previously dead code to handle computing the known bits of an alias	Chandler Carruth	2014-10-19	1	-10/+11
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	up to where it actually works as intended. The problem is that a GlobalAlias isa GlobalValue and so the prior block handled all of the cases. This allows us to constant fold based on the actual constant expression in the global alias. As an example, see the last function in the newly added test case which explicitly aligns an unaligned pointer using constant expression math. Without this change, we fail to see that and fold an alignment test to zero. llvm-svn: 220164
*	InstCombine: (sub (or A B) (xor A B)) --> (and A B)	David Majnemer	2014-10-19	1	-0/+9
\| \| \| \| \| \| \| \| \| \| \|	The following implements the transformation: (sub (or A B) (xor A B)) --> (and A B). Patch by Ankur Garg! Differential Revision: http://reviews.llvm.org/D5719 llvm-svn: 220163
*	InstCombine: Optimize icmp eq/ne (shl Const2, A), Const1	David Majnemer	2014-10-19	2	-1/+50
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The following implements the optimization for sequences of the form: icmp eq/ne (shl Const2, A), Const1 Such sequences can be transformed to: icmp eq/ne A, (TrailingZeros(Const1) - TrailingZeros(Const2)) This handles only the equality operators for now. Other operators need to be handled. Patch by Ankur Garg! llvm-svn: 220162
*	Fix a long-standing miscompile in the load analysis that was uncovered	Chandler Carruth	2014-10-19	2	-12/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	by my refactoring of this code. The method isSafeToLoadUnconditionally assumes that the load will proceed with the preferred type alignment. Given that, it has to ensure that the alloca or global is at least that aligned. It has always done this historically when a datalayout is present, but has never checked it when the datalayout is absent. When I refactored the code in r220156, I exposed this path when datalayout was present and that turned the latent bug into a patent bug. This fixes the issue by just removing the special case which allows folding things without datalayout. This isn't worth the complexity of trying to tease apart when it is or isn't safe without actually knowing the preferred alignment. llvm-svn: 220161
*	Switch how the datalayout availability test is handled in this code to	Chandler Carruth	2014-10-19	1	-7/+20
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	make much more sense and in theory be more correct. If you trace the code alllll the way back to when it was first introduced, the comments make it slightly more clear what was going on here. At that time, the only way Base != V was if DL (then TD) was non-null. As a consequence, if DL was null, that meant we were loading directly from the alloca or global found above the test. After refactoring, this has become at least terribly subtle and potentially incorrect. There are many forms of pointer manipulation that can be traversed without DataLayout, and some of them would in fact change the size of object being loaded vs. allocated. Rather than this subtlety, I've hoisted the actual 'return true' bits into the code which actually found an alloca or global and based them on the loaded pointer being that alloca or global. This is both more clear and safer. I've also added comments about exactly why this set of predicates is used. I've also corrected a misleading comment about globals -- if overridden they may not just have a different size, they may be null and completely unsafe to load from! Hopefully this confuses the next reader a bit less. I don't have any test cases or anything, the patch is motivated strictly to improve the readability of the code. llvm-svn: 220156
*	Use triple predicate functions instead of checking values directly. NFC.	Bob Wilson	2014-10-19	1	-24/+7
\| \| \| \|	llvm-svn: 220155
*	Rename 'TD' to 'DL' in this function as the argument is now a DataLayout	Chandler Carruth	2014-10-18	1	-7/+7
\| \| \| \| \| \|	argument. llvm-svn: 220151
*	Fix the other comment to use modern doxygen style and be a bit more	Chandler Carruth	2014-10-18	1	-4/+8
\| \| \| \| \| \| \|	direct. Notably, comment on the fact that the loaded type is significant in that it determines how wide of an access must be safe. llvm-svn: 220150
*	More formatting cleanup brought to you by clang-format.	Chandler Carruth	2014-10-18	1	-5/+8
\| \| \| \|	llvm-svn: 220149
*	Clean up doxygen syntax and reword comments to flow better, have a brief	Chandler Carruth	2014-10-18	1	-15/+20
\| \| \| \| \| \|	section, and not have unfinished sentence fragments. llvm-svn: 220147
*	Clean up the formatting and trailing whitespace of a routine before	Chandler Carruth	2014-10-18	1	-16/+19
\| \| \| \| \| \|	editting it. llvm-svn: 220146
*	[PBQP] Replace the interference-constraints algorithm with a faster version	Lang Hames	2014-10-18	1	-16/+115
\| \| \| \| \| \| \| \|	loosely based on linear scan. On x86-64 this is good for a ~2% drop in compile time on the nightly test suite. llvm-svn: 220143
*	Preserve AA metadata when combining (cast (load (...))) -> (load (cast	Chandler Carruth	2014-10-18	1	-0/+3
\| \| \| \| \| \|	(...))). llvm-svn: 220141
*	[InstCombine] Do an about-face on how LLVM canonicalizes (cast (load	Chandler Carruth	2014-10-18	1	-72/+43
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	...)) and (load (cast ...)): canonicalize toward the former. Historically, we've tried to load using the type of the pointer, and tried to match that type as closely as possible removing as many pointer casts as we could and trading them for bitcasts of the loaded value. This is deeply and fundamentally wrong. Repeat after me: memory does not have a type! This was a hard lesson for me to learn working on SROA. There is only one thing that should actually drive the type used for a pointer, and that is the type which we need to use to load from that pointer. Matching up pointer types to the loaded value types is very useful because it minimizes the physical size of the IR required for no-op casts. Similarly, the only thing that should drive the type used for a loaded value is how that value is used! Again, this minimizes casts. And in fact, the only thing motivating types in any part of LLVM's IR are the types used by the operations in the IR. We should match them as closely as possible. I've ended up removing some tests here as they were testing bugs or behavior that is no longer present. Mostly though, this is just cleanup to let the tests continue to function as intended. The only fallout I've found so far from this change was SROA and I have fixed it to not be impeded by the different type of load. If you find more places where this change causes optimizations not to fire, those too are likely bugs where we are assuming that the type of pointers is "significant" for optimization purposes. llvm-svn: 220138
*	[llvm-objdump] Fix mach-o binding decompression error	Nick Kledzik	2014-10-18	1	-3/+3
\| \| \| \|	llvm-svn: 220119
*	[SROA] Change how SROA does vector-based promotion of allocas to handle	Chandler Carruth	2014-10-18	1	-44/+128
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	cases where the alloca type, the load types, and the store types used all disagree. Previously, the only way that vector-based promotion occured was if the alloca type was a vector type. This was one of the very few remaining uses of the alloca's type to guide SROA/mem2reg left in LLVM. It turns out it was a bad idea. The alloca type can change very easily based on the mixture of types loaded and stored to that alloca. We shouldn't be relying on it as a signal for very much. Instead, the source of truth should be loads and stores. We should canonicalize the loads and stores as much as possible and then rely on them exclusively in SROA. When looking and loads and stores, we may find many different candidate vector types. This change will let SROA try all of them to find a vector type which is a viable way to promote the entire alloca to a vector register. With this change, it becomes possible to do better canonicalization and optimization of loads and stores without breaking SROA in random ways, and that should allow fixing a core source of performance loss in hot numerical loops such as those in Eigen. llvm-svn: 220116
*	R600/SI: Add global atomicrmw xchg	Aaron Watry	2014-10-17	2	-0/+4
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220110
*	R600/SI: Add global atomicrmw xor	Aaron Watry	2014-10-17	2	-1/+4
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220109
*	R600/SI: Add global atomicrmw or	Aaron Watry	2014-10-17	2	-1/+4
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220108
*	R600/SI: Add global atomicrmw min/umin	Aaron Watry	2014-10-17	2	-2/+8
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220107
*	R600/SI: Add global atomicrmw max/umax	Aaron Watry	2014-10-17	2	-2/+8
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220106
*	R600/SI: Add global atomicrmw and	Aaron Watry	2014-10-17	2	-1/+4
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220105
*	R600/SI: Add global atomicrmw sub	Aaron Watry	2014-10-17	2	-1/+4
\| \| \| \| \| \| \| \|	v2: Add separate offset/no-offset tests Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Matt Arsenault <matthew.arsenault@amd.com> llvm-svn: 220104
*	[msan] Fix handling of byval arguments with large alignment.	Evgeniy Stepanov	2014-10-17	1	-1/+2
\| \| \| \| \| \| \|	MSan param-tls slots are 8-byte aligned. This change clips alignment of memcpy into param-tls to 8. llvm-svn: 220101
*	Check for dynamic alloca's when selecting lifetime intrinsics.	Pete Cooper	2014-10-17	1	-1/+7
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	TL;DR: Indexing maps with [] creates missing entries. The long version: When selecting lifetime intrinsics, we index the static alloca map with the AllocaInst we find for that lifetime. Trouble is, we don't first check to see if this is a dynamic alloca. On the attached example, this causes a dynamic alloca to create an entry in the static map, and returns 0 (the default) as the frame index for that lifetime. 0 was used for the frame index of the stack protector, which given that it now has a lifetime, is coloured, and merged with other stack slots. PEI would later trigger an assert because it expects the stack protector to not be dead. This fix ensures that we only get frame indices for static allocas, ie, those in the map. Dynamic ones are effectively dropped, which is suboptimal, but at least isn't completely broken. rdar://problem/18672951 llvm-svn: 220099