bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
*	Fix a typo in comment.	Jim Grosbach	2013-04-15	1	-1/+1
\| \| \| \|	llvm-svn: 179542
*	Add an option -vectorize-slp-aggressive for running the BB vectorizer. Make ↵	Nadav Rotem	2013-04-15	1	-1/+12
\| \| \| \| \| \|	-fslp-vectorize run the slp-vectorizer. llvm-svn: 179508
*	Rename the slp-vectorizer clang/llvm flags. No functionality change.	Nadav Rotem	2013-04-15	1	-3/+3
\| \| \| \|	llvm-svn: 179505
*	SLPVectorizer: Add support for vectorizing trees that start at compare ↵	Nadav Rotem	2013-04-15	1	-21/+40
\| \| \| \| \| \|	instructions. llvm-svn: 179504
*	Reorders two transforms that collide with each other	David Majnemer	2013-04-14	1	-8/+8
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	One performs: (X == 13 \| X == 14) -> X-13 <u 2 The other: (A == C1 \|\| A == C2) -> (A & ~(C1 ^ C2)) == C1 The problem is that there are certain values of C1 and C2 that trigger both transforms but the first one blocks out the second, this generates suboptimal code. Reordering the transforms should be better in every case and allows us to do interesting stuff like turn: %shr = lshr i32 %X, 4 %and = and i32 %shr, 15 %add = add i32 %and, -14 %tobool = icmp ne i32 %add, 0 into: %and = and i32 %X, 240 %tobool = icmp ne i32 %and, 224 llvm-svn: 179493
*	Miscellaneous cleanups for VecUtils.h	Benjamin Kramer	2013-04-14	1	-9/+6
\| \| \| \|	llvm-svn: 179483
*	SLP: Document the scalarization cost method.	Nadav Rotem	2013-04-14	1	-3/+10
\| \| \| \|	llvm-svn: 179479
*	SLPVectorizer: Add support for trees that don't start at binary operators, ↵	Nadav Rotem	2013-04-14	3	-7/+25
\| \| \| \| \| \|	and add the cost of extracting values from the roots of the tree. llvm-svn: 179475
*	SLPVectorizer: add initial support for reduction variable vectorization.	Nadav Rotem	2013-04-14	3	-7/+95
\| \| \| \|	llvm-svn: 179470
*	GlobalDCE: Fix an oversight in my last commit that could lead to crashes.	Benjamin Kramer	2013-04-13	1	-2/+2
\| \| \| \| \| \|	There is a Constant with non-constant operands: blockaddress. llvm-svn: 179460
*	Fix a scalability issue with complex ConstantExprs.	Benjamin Kramer	2013-04-13	1	-4/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This is basically the same fix in three different places. We use a set to avoid walking the whole tree of a big ConstantExprs multiple times. For example: (select cmp, (add big_expr 1), (add big_expr 2)) We don't want to visit big_expr twice here, it may consist of thousands of nodes. The testcase exercises this by creating an insanely large ConstantExprs out of a loop. It's questionable if the optimizer should ever create those, but this can be triggered with real C code. Fixes PR15714. llvm-svn: 179458
*	InstCombine: Check the operand types before merging fcmp ord & fcmp ord.	Benjamin Kramer	2013-04-12	1	-0/+3
\| \| \| \| \| \|	Fixes PR15737. llvm-svn: 179417
*	SLPVectorizer: add support for vectorization of diamond shaped trees. We now ↵	Nadav Rotem	2013-04-12	2	-46/+254
\| \| \| \| \| \|	perform a preliminary traversal of the graph to collect values with multiple users and check where the users came from. llvm-svn: 179414
*	Add debug prints.	Nadav Rotem	2013-04-12	1	-1/+5
\| \| \| \|	llvm-svn: 179412
*	Simplify (A & ~B) in icmp if A is a power of 2	David Majnemer	2013-04-12	1	-0/+9
\| \| \| \| \| \| \| \|	The transform will execute like so: (A & ~B) == 0 --> (A & B) != 0 (A & ~B) != 0 --> (A & B) == 0 llvm-svn: 179386
*	LoopVectorizer: integer division is not a reduction operation	Arnold Schwaighofer	2013-04-12	1	-2/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Don't classify idiv/udiv as a reduction operation. Integer division is lossy. For example : (1 / 2) * 4 != 4/2. Example: int a[] = { 2, 5, 2, 2} int x = 80; for() x /= a[i]; Scalar: x /= 2 // = 40 x /= 5 // = 8 x /= 2 // = 4 x /= 2 // = 2 Vectorized: <80, 1> / <2,5> //= <40,0> <40, 0> / <2,2> //= <20,0> 20*0 = 0 radar://13640654 llvm-svn: 179381
*	Optimize icmp involving addition better	David Majnemer	2013-04-11	1	-0/+49
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Allows LLVM to optimize sequences like the following: %add = add nsw i32 %x, 1 %cmp = icmp sgt i32 %add, %y into: %cmp = icmp sge i32 %x, %y as well as: %add1 = add nsw i32 %x, 20 %add2 = add nsw i32 %y, 57 %cmp = icmp sge i32 %add1, %add2 into: %add = add nsw i32 %y, 37 %cmp = icmp sle i32 %cmp, %x llvm-svn: 179316
*	Fix for wrong instcombine on vector insert/extract	Benjamin Kramer	2013-04-11	1	-0/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When trying to collapse sequences of insertelement/extractelement instructions into single shuffle instructions, there is one specific case where the Instruction Combiner wrongly updates the resulting Mask of shuffle indexes. The problem is in function CollectShuffleElments. If we have a sequence of insert/extract element instructions like the one below: %tmp1 = extractelement <4 x float> %LHS, i32 0 %tmp2 = insertelement <4 x float> %RHS, float %tmp1, i32 1 %tmp3 = extractelement <4 x float> %RHS, i32 2 %tmp4 = insertelement <4 x float> %tmp2, float %tmp3, i32 3 Where: . %RHS will have a mask of [4,5,6,7] . %LHS will have a mask of [0,1,2,3] The Mask of shuffle indexes is wrongly computed to [4,1,6,7] instead of [4,0,6,7]. When analyzing %tmp2 in order to compute the Mask for the resulting shuffle instruction, the algorithm forgets to update the mask index at position 1 with the index associated to the element extracted from %LHS by instruction %tmp1. Patch by Andrea DiBiagio! llvm-svn: 179291
*	[ASan] Allow disabling init-order checks for globals by source file name.	Alexey Samsonov	2013-04-11	1	-1/+2
\| \| \| \|	llvm-svn: 179280
*	Rename the C function to create a SLPVectorizerPass to something sane and ↵	Benjamin Kramer	2013-04-11	1	-2/+2
\| \| \| \| \| \|	expose it in the header file. llvm-svn: 179272
*	Make the SLP store-merger less paranoid about function calls. We check for ↵	Nadav Rotem	2013-04-10	1	-4/+0
\| \| \| \| \| \|	function calls when we check if it is safe to sink instructions. llvm-svn: 179207
*	We require DataLayout for analyzing the size of stores.	Nadav Rotem	2013-04-10	2	-1/+6
\| \| \| \|	llvm-svn: 179206
*	Change CloneFunctionInto to always clone Argument attributes induvidually,	Joey Gouly	2013-04-10	1	-22/+19
\| \| \| \| \| \| \|	rather than checking if the source and destination have the same number of arguments and copying the attributes over directly. llvm-svn: 179169
*	Fix some comment typos.	Bob Wilson	2013-04-09	1	-2/+2
\| \| \| \|	llvm-svn: 179132
*	Add support for bottom-up SLP vectorization infrastructure.	Nadav Rotem	2013-04-09	5	-0/+707
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This commit adds the infrastructure for performing bottom-up SLP vectorization (and other optimizations) on parallel computations. The infrastructure has three potential users: 1. The loop vectorizer needs to be able to vectorize AOS data structures such as (sum += A[i] + A[i+1]). 2. The BB-vectorizer needs this infrastructure for bottom-up SLP vectorization, because bottom-up vectorization is faster to compute. 3. A loop-roller needs to be able to analyze consecutive chains and roll them into a loop, in order to reduce code size. A loop roller does not need to create vector instructions, and this infrastructure separates the chain analysis from the vectorization. This patch also includes a simple (100 LOC) bottom up SLP vectorizer that uses the infrastructure, and can vectorize this code: void SAXPY(int x, int y, int a, int i) { x[i] = a * x[i] + y[i]; x[i+1] = a * x[i+1] + y[i+1]; x[i+2] = a * x[i+2] + y[i+2]; x[i+3] = a * x[i+3] + y[i+3]; } llvm-svn: 179117
*	Redo the fix Benjamin Kramer committed in r178793 about iterator ↵	Shuxin Yang	2013-04-08	1	-12/+14
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	invalidation in Reassociate. I brazenly think this change is slightly simpler than r178793 because: - no "state" in functor - "OpndPtrs[i]" looks simpler than "&Opnds[OpndIndices[i]]" While I can reproduce the probelm in Valgrind, it is rather difficult to come up a standalone testing case. The reason is that when an iterator is invalidated, the stale invalidated elements are not yet clobbered by nonsense data, so the optimizer can still proceed successfully. Thank Benjamin for fixing this bug and generously providing the test case. llvm-svn: 179062
*	Fix PR15674 (and PR15603): a SROA think-o.	Chandler Carruth	2013-04-07	1	-0/+1
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	The fix for PR14972 in r177055 introduced a real think-o in the store side, likely because I was much more focused on the load side. While we can arbitrarily widen (or narrow) a loaded value, we can't arbitrarily widen a value to be stored, as that changes the width of memory access! Lock down the code path in the store rewriting which would do this to only handle the intended circumstance. All of the existing tests continue to pass, and I've added a test from the PR. llvm-svn: 178974
*	Removed trailing whitespace.	Michael Gottesman	2013-04-05	1	-27/+27
\| \| \| \|	llvm-svn: 178932
*	An objc_retain can serve as a use for a different pointer.	Michael Gottesman	2013-04-05	1	-2/+3
\| \| \| \| \| \| \|	This is the counterpart to commit r160637, except it performs the action in the bottomup portion of the data flow analysis. llvm-svn: 178922
*	Properly model precise lifetime when given an incomplete dataflow sequence.	Michael Gottesman	2013-04-05	1	-6/+20
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The normal dataflow sequence in the ARC optimizer consists of the following states: Retain -> CanRelease -> Use -> Release The optimizer before this patch stored the uses that determine the lifetime of the retainable object pointer when it bottom up hits a retain or when top down it hits a release. This is correct for an imprecise lifetime scenario since what we are trying to do is remove retains/releases while making sure that no ``CanRelease'' (which is usually a call) deallocates the given pointer before we get to the ``Use'' (since that would cause a segfault). If we are considering the precise lifetime scenario though, this is not correct. In such a situation, we DO care about the previous sequence, but additionally, we wish to track the uses resulting from the following incomplete sequences: Retain -> CanRelease -> Release (TopDown) Retain <- Use <- Release (BottomUp) NOTE This patch looks large but the most of it consists of updating test cases. Additionally this fix exposed an additional bug. I removed the test case that expressed said bug and will recommit it with the fix in a little bit. llvm-svn: 178921
*	Tidy up a bit. No functional change.	Jim Grosbach	2013-04-05	9	-259/+261
\| \| \| \|	llvm-svn: 178915
*	Disable the optimization about promoting vector-element-access with symbolic ↵	Shuxin Yang	2013-04-05	1	-11/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	index. This optimization is unstable at this moment; it 1) block us on a very important application 2) PR15200 3) test6 and test7 in test/Transforms/ScalarRepl/dynamic-vector-gep.ll (the CHECK command compare the output against wrong result) I personally believe this optimization should not have any impact on the autovectorized code, as auto-vectorizer is supposed to put gather/scatter in a "right" way. Although in theory downstream optimizaters might reveal some gather/scatter optimization opportunities, the chance is quite slim. For the hand-crafted vectorizing code, in term of redundancy elimination, load-CSE, copy-propagation and DSE can collectively achieve the same result, but in much simpler way. On the other hand, these optimizers are able to improve the code in a incremental way; in contrast, SROA is sort of all-or-none approach. However, SROA might slighly win in stack size, as it tries to figure out a stretch of memory tightenly cover the area accessed by the dynamic index. rdar://13174884 PR15200 llvm-svn: 178912
*	Added two debug logging messages to VisitInstructionsTopDown to match ↵	Michael Gottesman	2013-04-05	1	-0/+4
\| \| \| \| \| \|	VisitInstructionsBottomUp. llvm-svn: 178895
*	Cleaned up whitespace and made debug logging less verbose.	Michael Gottesman	2013-04-05	1	-114/+95
\| \| \| \|	llvm-svn: 178893
*	LoopVectorizer: Pass OperandValueKind information to the cost model	Arnold Schwaighofer	2013-04-04	1	-2/+13
\| \| \| \| \| \| \| \| \| \| \| \|	Pass down the fact that an operand is going to be a vector of constants. This should bring the performance of MultiSource/Benchmarks/PAQ8p/paq8p on x86 back. It had degraded to scalar performance due to my pervious shift cost change that made all shifts expensive on x86. radar://13576547 llvm-svn: 178809
*	Reassociate: Avoid iterator invalidation.	Benjamin Kramer	2013-04-04	1	-7/+12
\| \| \| \| \| \| \| \|	OpndPtrs stored pointers into the Opnd vector that became invalid when the vector grows. Store indices instead. Sadly I only have a large testcase that only triggers under valgrind, so I didn't include it. llvm-svn: 178793
*	Refactored out the helper method FindPredecessorAutoreleaseWithSafePath from ↵	Michael Gottesman	2013-04-03	1	-25/+45
\| \| \| \| \| \| \| \|	ObjCARCOpt::OptimizeReturns. Now ObjCARCOpt::OptimizeReturns is easy to read and reason about. llvm-svn: 178715
*	Refactored out the helper function FindPredecessorRetainWithSafePath from ↵	Michael Gottesman	2013-04-03	1	-18/+32
\| \| \| \| \| \|	ObjCARCOpt::OptimizeReturns. llvm-svn: 178714
*	Small cleanups.	Michael Gottesman	2013-04-03	1	-14/+14
\| \| \| \| \| \| \| \|	Cleaned up trailing whitespace and added extra slashes in front of a function level comment so that it follow the convention of having 3 slashes. llvm-svn: 178712
*	Refactored out a part of ObjCARCOpt::OptimizeReturns into its own method ↵	Michael Gottesman	2013-04-03	1	-22/+33
\| \| \| \| \| \|	HasSafePathToPredecessorCall. llvm-svn: 178710
*	Removed an old comment.	Michael Gottesman	2013-04-03	1	-7/+0
\| \| \| \|	llvm-svn: 178709
*	Clean up arc annotations by moving the top/bottom BB annotations into ↵	Michael Gottesman	2013-04-03	1	-58/+46
\| \| \| \| \| \| \| \|	conditional macros that no-op in Release mode instead of #ifdef sections of the code. This is to follow the example of the DEBUG macro. llvm-svn: 178705
*	Remove an optimization where we were changing an objc_autorelease into an ↵	Michael Gottesman	2013-04-03	1	-16/+1
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	objc_autoreleaseReturnValue. The semantics of ARC implies that a pointer passed into an objc_autorelease must live until some point (potentially down the stack) where an autorelease pool is popped. On the other hand, an objc_autoreleaseReturnValue just signifies that the object must live until the end of the given function at least. Thus objc_autorelease is stronger than objc_autoreleaseReturnValue in terms of the semantics of ARC* implying that performing the given strength reduction without any knowledge of how this relates to the autorelease pool pop that is further up the stack violates the semantics of ARC. *Even though objc_autoreleaseReturnValue if you know that no RV optimization will occur is more computationally expensive. llvm-svn: 178612
*	Improved comment. No functionality change.	Michael Gottesman	2013-04-03	1	-1/+2
\| \| \| \|	llvm-svn: 178605
*	Use a worklist to avoid a sneaky iterator invalidation.	Bill Wendling	2013-04-02	1	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \|	The iterator could be invalidated when it's recursively deleting a whole bunch of constant expressions in a constant initializer. Note: This was only reproducible if `opt' was run on a `.bc' file. If `opt' was run on a `.ll' file, it wouldn't crash. This is why the test first pushes the `.ll' file through `llvm-as' before feeding it to `opt'. PR15440 llvm-svn: 178531
*	Correct assertion condition	Shuxin Yang	2013-04-01	1	-1/+1
\| \| \| \|	llvm-svn: 178484
*	Implement XOR reassociation. It is based on following rules:	Shuxin Yang	2013-03-30	1	-1/+325
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	rule 1: (x \| c1) ^ c2 => (x & ~c1) ^ (c1^c2), only useful when c1=c2 rule 2: (x & c1) ^ (x & c2) = (x & (c1^c2)) rule 3: (x \| c1) ^ (x \| c2) = (x & c3) ^ c3 where c3 = c1 ^ c2 rule 4: (x \| c1) ^ (x & c2) => (x & c3) ^ c1, where c3 = ~c1 ^ c2 It reduces an application's size (in terms of # of instructions) by 8.9%. Reviwed by Pete Cooper. Thanks a lot! rdar://13212115 llvm-svn: 178409
*	Add clang.arc.used to ModuleHasARC so ARC always runs if said call is ↵	Michael Gottesman	2013-03-29	1	-1/+2
\| \| \| \| \| \| \| \| \| \|	present in a module. clang.arc.used is an interesting call for ARC since ObjCARCContract needs to run to remove said intrinsic to avoid a linker error (since the call does not exist). llvm-svn: 178369
*	Removed trailing whitespace.	Michael Gottesman	2013-03-29	1	-15/+15
\| \| \| \|	llvm-svn: 178329
*	Removed dead code from ObjCARCOpts relating to tracking objc_retainBlocks ↵	Michael Gottesman	2013-03-28	1	-37/+6
\| \| \| \| \| \|	through the ARC Dataflow analysis. By the time we get to the ARC dataflow analysis, any objc_retainBlock calls are not optimizable. llvm-svn: 178306