bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
*	[WebAssembly] Only treat imports/exports as symbols when reading relocatable ↵	Sam Clegg	2017-09-06	6	-14/+21
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	object files This change only treats imported and exports functions and globals as symbol table entries the object has a "linking" section (i.e. it is relocatable object file). In this case all globals must be of type I32 and initialized with i32.const. This was previously being assumed but not checked for and was causing a failure on big endian machines due to using the wrong value of then union. See: https://bugs.llvm.org/show_bug.cgi?id=34487 Differential Revision: https://reviews.llvm.org/D37497 llvm-svn: 312674
*	Insert IMPLICIT_DEFS for undef uses in tail merging	Matthias Braun	2017-09-06	2	-3/+90
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Tail merging can convert an undef use into a normal one when creating a common tail. Doing so can make the register live out from a block which previously contained the undef use. To keep the liveness up-to-date, insert IMPLICIT_DEFs in such blocks when necessary. To enable this patch the computeLiveIns() function which used to compute live-ins for a block and set them immediately is split into new functions: - computeLiveIns() just computes the live-ins in a LivePhysRegs set. - addLiveIns() applies the live-ins to a block live-in list. - computeAndAddLiveIns() is a convenience function combining the other two functions and behaving like computeLiveIns() before this patch. Based on a patch by Krzysztof Parzyszek <kparzysz@codeaurora.org> Differential Revision: https://reviews.llvm.org/D37034 llvm-svn: 312668
*	Disable jump threading into loop headers	Krzysztof Parzyszek	2017-09-06	1	-14/+23
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Consider this type of a loop: for (...) { ... if (...) continue; ... } Normally, the "continue" would branch to the loop control code that checks whether the loop should continue iterating and which contains the (often) unique loop latch branch. In certain cases jump threading can "thread" the inner branch directly to the loop header, creating a second loop latch. Loop canonicalization would then transform this loop into a loop nest. The problem with this is that in such a loop nest neither loop is countable even if the original loop was. This may inhibit subsequent loop optimizations and be detrimental to performance. Differential Revision: https://reviews.llvm.org/D36404 llvm-svn: 312664
*	[x86] fix triple and regenerate checks for psubus; NFC	Sanjay Patel	2017-09-06	1	-65/+65
\| \| \| \| \| \| \| \|	Patch by Yulia Koval! Differential Revision: https://reviews.llvm.org/D37523 llvm-svn: 312662
*	[AMDGPU] Fixed encoding of v_pk_mul_f16 in fcanonicalize	Stanislav Mekhanoshin	2017-09-06	1	-6/+5
\| \| \| \| \| \|	Differential Revision: https://reviews.llvm.org/D37522 llvm-svn: 312660
*	[IfConversion] Remove kill flags from common instructions as well	Krzysztof Parzyszek	2017-09-06	1	-0/+34
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When if-converting a diamond, two separate blocks will be placed back to back to form a straight line code. To ensure correctness of the liveness information, any registers that are live in the second block should not be killed in the first block, even if they were in the original code. Additionally, when the two blocks share common instructions at the beginning, these instructions will not be duplicated, but only placed once, before both of the blocks. Since the function "isIdenticalTo" (as used here) ignores kill flags, the common initial code in one block may have a kill flag for a register that is live in the other block. Because the code that removes kill flags only runs for the non-common parts of the predicated blocks, a kill flag mismatch in the common code could still lead to a live register being killed prematurely. llvm-svn: 312654
*	Revert "[llvm-objcopy] Add support for relocations"	Petr Hosek	2017-09-06	2	-121/+0
\| \| \| \| \| \|	This reverts r312643 because it's failing on llvm-i686-linux-RA. llvm-svn: 312645
*	[Hexagon] Add option to generate calls to "abort" for "unreachable"	Krzysztof Parzyszek	2017-09-06	1	-0/+8
\| \| \| \|	llvm-svn: 312644
*	[llvm-objcopy] Add support for relocations	Petr Hosek	2017-09-06	2	-0/+121
\| \| \| \| \| \| \| \| \| \| \|	This change adds support for SHT_REL and SHT_RELA sections in llvm-objcopy. Patch by Jake Ehrlich Differential Revision: https://reviews.llvm.org/D36554 llvm-svn: 312643
*	[TailCall] Allow llvm.memcpy/memset/memmove to be tail calls when parent	Wei Mi	2017-09-06	1	-0/+24
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	function return the intrinsics's first argument. llvm.memcpy/memset/memmove return void but they will return the first argument after they are expanded as libcalls. Now if the parent function has any return value, llvm.memcpy cannot be turned into tail call after expansion. The patch is to handle that case in SelectionDAGBuilder so when caller function return the same value as the first argument of llvm.memcpy, tail call is allowed. Differential Revision: https://reviews.llvm.org/D37406 llvm-svn: 312641
*	[AMDGPU] Fix shouldClusterMemOps to process flat loads	Stanislav Mekhanoshin	2017-09-06	1	-0/+20
\| \| \| \| \| \| \| \|	Flat loads do not have vdata operand but have vdst instead. Differential Revision: https://reviews.llvm.org/D37502 llvm-svn: 312640
*	AMDGPU: Make worst-case assumption about the wait states in inline assembly	Nicolai Haehnle	2017-09-06	1	-0/+29
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: Mesa still uses a hack where empty inline assembly is used as a kind of optimization barrier. This exposed a problem where not enough wait states were inserted, because the hazard recognizer implicitly assumed that each inline assembly "instruction" has at least one wait state. Reviewers: arsenm Subscribers: kzhuravl, wdng, yaxunl, dstuttard, tpr, llvm-commits, t-tye Differential Revision: https://reviews.llvm.org/D37205 llvm-svn: 312635
*	[x86] Fix PR34377 by disabling cmov conversion when we relied on it	Chandler Carruth	2017-09-06	1	-4/+22
\| \| \| \| \| \| \| \| \| \| \|	performing a zext of a register. On the PR there is discussion of how to more effectively handle this, but this patch prevents us from miscompiling code. Differential Revision: https://reviews.llvm.org/D37504 llvm-svn: 312620
*	X86 Tests: Tidy up AVX512 conversion tests. NFC.	Zvi Rackover	2017-09-06	1	-227/+227
\| \| \| \| \| \|	Rename functions to a consistent format to make it easier to track coverage. llvm-svn: 312619
*	Updating a test reference for rL312608.	Jatin Bhateja	2017-09-06	1	-13/+6
\| \| \| \| \| \|	Differential Revision: https://reviews.llvm.org/D37501 llvm-svn: 312614
*	[PowerPC] Don't use xscvdpspn on the P7	Hal Finkel	2017-09-06	1	-0/+27
\| \| \| \| \| \| \|	xscvdpspn was not introduced until the P8, so don't use it on the P7. Fixes a regression introduced in r288152. llvm-svn: 312612
*	[X86] Allow cross-lane permutations for sub targets supporting AVX2.	Jatin Bhateja	2017-09-06	5	-677/+752
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: Most instructions in AVX work “in-lane”, that is, each source element is applied only to other elements of the same lane, thus a cross lane permutation is costly and needs more than one instrution. AVX2 includes instructions to perform any-to-any permutation of words over a 256-bit register and vectorized table lookup. This should also Fix PR34369 Differential Revision: https://reviews.llvm.org/D37388 llvm-svn: 312608
*	Use the section name if a STT_SECTION symbol has empty name.	Rafael Espindola	2017-09-06	2	-18/+44
\| \| \| \| \| \| \| \| \| \| \| \| \|	Without this we would have multiple relocations pointing to symbols with the same name: the empty string. There was no way for yaml2obj to be able to handle that. A more general solution would be to unique symbol names in a similar way to how we unique section names. In practice I think this covers all common cases and is a bit more user friendly than using names like sym1, sym2, sym3, etc. llvm-svn: 312603
*	[AMDGPU] Transform __read_pipe_* and __write_pipe_*	Yaxun Liu	2017-09-06	1	-18/+111
\| \| \| \| \| \| \| \| \|	When packet size equals packet align and is power of 2, transform __read_pipe* and __write_pipe* to specialized library function. Differential Revision: https://reviews.llvm.org/D36831 llvm-svn: 312598
*	[ValueTracking, InstCombine] canonicalize fcmp ord/uno with non-NAN ops to ↵	Sanjay Patel	2017-09-05	3	-20/+10
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	null constants This is a preliminary step towards solving the remaining part of PR27145 - IR for isfinite(): https://bugs.llvm.org/show_bug.cgi?id=27145 In order to solve that one more generally, we need to add matching for and/or of fcmp ord/uno with a constant operand. But while looking at those patterns, I realized we were missing a canonicalization for nonzero constants. Rather than limiting to just folds for constants, we're adding a general value tracking method for this based on an existing DAG helper. By transforming everything to 0.0, we can simplify the existing code in foldLogicOfFCmps() and pick up missing vector folds. Differential Revision: https://reviews.llvm.org/D37427 llvm-svn: 312591
*	[ARM] Make ARMExpandPseudo add implicit uses for predicated instructions	Eli Friedman	2017-09-05	1	-0/+75
\| \| \| \| \| \| \| \| \| \| \|	Missing these could potentially screw up post-ra scheduling. Issue found by inspection, so I don't have a real testcase. Included test just verifies the expected operands after expansion. Differential Revision: https://reviews.llvm.org/D35156 llvm-svn: 312589
*	obj2yaml: Print unique section names.	Rafael Espindola	2017-09-05	1	-0/+28
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Without this patch passing a .o file with multiple sections with the same name to obj2yaml produces a yaml file that yaml2obj cannot handle. This is pr34162. The problem is that when specifying, for example, the section of a symbol, we get only Section: foo and don't know which of the sections whose name is foo we have to use. One alternative would be to use section numbers. This would work, but the output from obj2yaml would be very inconvenient to edit as deleting a section would invalidate all indexes. Another alternative would be to invent a unique section id that would exist only on yaml. This would work, but seems a bit heavy handed. We could make the id optional and default it to the section name. Since in the last alternative the id is basically what this patch uses as a name, it can be implemented as a followup patch if needed. llvm-svn: 312585
*	[CodeView] Don't output S_UDTs for nested typedefs.	Zachary Turner	2017-09-05	1	-92/+127
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	S_UDT records are basically the "bridge" between the debugger's expression evaluator and the type information. If you type (Foo)nullptr into the watch window, the debugger looks for an S_UDT record named Foo. If it can find one, it displays your type. Otherwise you get an error. We have always understood this to mean that if you have code like this: struct A { int X; }; struct B { typedef A AT; AT Member; }; that you will get 3 S_UDT records. "A", "B", and "B::AT". Because if you were to type (B::AT)nullptr into the debugger, it would need to find an S_UDT record named "B::AT". But "B::AT" is actually the S_UDT record that would be generated if B were a namespace, not a struct. So the debugger needs to be able to distinguish this case. So what it does is: 1. Look for an S_UDT named "B::AT". If it finds one, it knows that AT is in a namespace. 2. If it doesn't find one, split at the scope resolution operator, and look for an S_UDT named B. If it finds one, look up the type for B, and then look for AT as one of its members. With this algorithm, S_UDT records for nested typedefs are not just unnecessary, but actually wrong! The results of implementing this in clang are dramatic. It cuts our /DEBUG:FASTLINK PDB sizes by more than 50%, and we go from being ~20% larger than MSVC PDBs on average, to ~40% smaller. It also slightly speeds up link time. We get about 10% faster links than without this patch. Differential Revision: https://reviews.llvm.org/D37410 llvm-svn: 312583
*	Revert "[Decompression] Fail gracefully when out of memory"	Vedant Kumar	2017-09-05	2	-15/+0
\| \| \| \| \| \| \| \| \| \| \| \| \|	This reverts commit r312526. Revert "Fix test/DebugInfo/dwarfdump-decompression-invalid-size.test" This reverts commit r312527. It causes an ASan failure: http://lab.llvm.org:8080/green/job/clang-stage2-cmake-RgSan_check/4150 llvm-svn: 312582
*	[InstCombine] add nnan tests; NFC	Sanjay Patel	2017-09-05	1	-0/+30
\| \| \| \| \| \| \|	As suggested in D37427, we could have a value tracking function and folds that use it to simplify these cases. llvm-svn: 312578
*	Add llvm.codeview.annotation to implement MSVC __annotation	Reid Kleckner	2017-09-05	2	-0/+108
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: This intrinsic represents a label with a list of associated metadata strings. It is modelled as reading and writing inaccessible memory so that it won't be removed as dead code. I think the intention is that the annotation strings should appear at most once in the debug info, so I marked it noduplicate. We are allowed to inline code with annotations as long as we strip the annotation, but that can be done later. Reviewers: majnemer Subscribers: eraman, llvm-commits, hiraditya Differential Revision: https://reviews.llvm.org/D36904 llvm-svn: 312569
*	[X86] Remove unnecessary (v4f32 (X86vzmovl (v4f32 (scalar_to_vector ↵	Craig Topper	2017-09-05	1	-4/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	FR32X)))) patterns We had already disabled the pattern for SSE4.1 and SSE4.2. But it got re-enabled for AVX and AVX512. With SSE41 we rely on a separate (v4f32 (X86vzmovl VR128)) pattern to select blendps with a xorps to create zeroess. And a separate (v4f32 (scalar_to_vector FR32X)) to select a COPY_TO_REG_CLASS to move FR32 to VR128 The same thing can happen for AVX with vblendps and those separate patterns already exist. For AVX512, (v4f32 (X86vzmov VR128)) will select a VMOVSS instruction instead of VBLENDPS due to their not being a EVEX VBLENDPS. This is what we were getting out of the larger pattern anyway. So the larger pattern is unneeded for AVX512 too. For SSE1-SSSE3 we can rely on (v4f32 (X86vzmov VR128)) selecting a MOVSS similar to AVX512. Again this is what the larger pattern did too. So the only real change here is that AVX1/2 now properly outputs a VBLENDPS during isel instead of a VMOVSS to match SSE41. Most tests didn't notice because the two address instruction pass knows how to turn VMOVSS into VBLENDPS to get an independent destination register. llvm-svn: 312564
*	AMDGPU: Fix not accounting for tail call resource usage	Matt Arsenault	2017-09-05	1	-0/+31
\| \| \| \| \| \| \| \|	If the only call in a function is a tail call, the function isn't considered to have a call since it's a type of return. llvm-svn: 312561
*	X86 Tests: Adding missing AVX512 fptoui coverage tests. NFC.	Zvi Rackover	2017-09-05	1	-0/+231
\| \| \| \| \| \|	Some of the cases show missing pattern i intend to fix shortly. llvm-svn: 312560
*	Split opt-remark YAML and opt output testing on this test	Adam Nemet	2017-09-05	1	-2/+5
\| \| \| \| \| \|	This prepares for https://reviews.llvm.org/D33514 llvm-svn: 312544
*	[AVX512] Remove patterns for (v8f32 (X86vzmovl (insert_subvector undef, ↵	Craig Topper	2017-09-05	1	-0/+1
\| \| \| \| \| \| \| \|	(v4f32 (scalar_to_vector FR32X:)), (iPTR 0)))) and the same for v4f64. We don't have this same pattern for AVX2 so I don't believe we should have it for AVX512. We also didn't have it for v16f32. llvm-svn: 312543
*	[AMDGPU] Added extra test checks to make D19325 diff clearer	Simon Pilgrim	2017-09-05	1	-5/+11
\| \| \| \|	llvm-svn: 312537
*	[X86] Limit store merge size when implicitfloat is enabled (PR34421)	Simon Pilgrim	2017-09-05	1	-0/+40
\| \| \| \| \| \| \| \|	As suggested by @niravd : https://bugs.llvm.org/show_bug.cgi?id=34421#c2 Differential Revision: https://reviews.llvm.org/D37464 llvm-svn: 312534
*	[X86] Regenerate scalar rotation tests	Simon Pilgrim	2017-09-05	2	-69/+207
\| \| \| \|	llvm-svn: 312530
*	[X86][AVX512] Use AVX512 attributes instead of -mcpu in vector shift tests	Simon Pilgrim	2017-09-05	9	-38/+76
\| \| \| \|	llvm-svn: 312529
*	[X86][AVX512] Use AVX512 attributes instead of -mcpu	Simon Pilgrim	2017-09-05	3	-8/+18
\| \| \| \|	llvm-svn: 312528
*	Fix test/DebugInfo/dwarfdump-decompression-invalid-size.test	Jonas Devlieghere	2017-09-05	1	-0/+2
\| \| \| \|	llvm-svn: 312527
*	[Decompression] Fail gracefully when out of memory	Jonas Devlieghere	2017-09-05	2	-0/+13
\| \| \| \| \| \| \| \| \| \| \| \|	This patch adds failing gracefully when running out of memory when allocating a buffer for decompression. This provides a work-around for: https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=3224 Differential revision: https://reviews.llvm.org/D37447 llvm-svn: 312526
*	[ARM] GlobalISel: Support global variables for RWPI	Diana Picus	2017-09-05	3	-12/+72
\| \| \| \| \| \| \| \| \|	In RWPI code, globals that are not read-only are accessed relative to the SB register (R9). This is achieved by explicitly generating an ADD instruction between SB and an offset that we either load from a constant pool or movw + movt into a register. llvm-svn: 312521
*	[InstCombine] Add test cases for folding (select (icmp ne/eq (and X, C1), ↵	Craig Topper	2017-09-05	1	-1/+800
\| \| \| \| \| \| \| \| \| \|	(bitwiseop Y, C2), Y -> (bitwiseop Y, (shl/shr (and X, C1), C3)) or similar. This is possible if C1 and C2 are both powers of 2. Or if binop is 'and' then ~C2 needs to be a power of 2. We already support this for 'or', but we should be able to support 'and' and 'xor'. This will be enhanced by D37274. llvm-svn: 312519
*	[PowerPC] eliminate redundant compare instruction	Hiroshi Inoue	2017-09-05	1	-0/+722
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	If multiple conditional branches are executed based on the same comparison, we can execute multiple conditional branches based on the result of one comparison on PPC. For example, if (a == 0) { ... } else if (a < 0) { ... } can be executed by one compare and two conditional branches instead of two pairs of a compare and a conditional branch. This patch identifies a code sequence of the two pairs of a compare and a conditional branch and merge the compares if possible. To maximize the opportunity, we do canonicalization of code sequence before merging compares. For the above example, the input for this pass looks like: cmplwi r3, 0 beq 0, .LBB0_3 cmpwi r3, -1 bgt 0, .LBB0_4 So, before merging two compares, we canonicalize it as cmpwi r3, 0 ; cmplwi and cmpwi yield same result for beq beq 0, .LBB0_3 cmpwi r3, 0 ; greather than -1 means greater or equal to 0 bge 0, .LBB0_4 The generated code should be cmpwi r3, 0 beq 0, .LBB0_3 bge 0, .LBB0_4 Differential Revision: https://reviews.llvm.org/D37211 llvm-svn: 312514
*	NewGVN: Fix PR 34430 - we need to look through predicateinfo copies to ↵	Daniel Berlin	2017-09-05	1	-0/+48
\| \| \| \| \| \|	detect self-cycles of phi nodes. We also need to not ignore certain types of arguments when testing whether the phi has a backedge or was originally constant. llvm-svn: 312510
*	NewGVN: Fix PR 34452 by passing instruction all the way down when we do ↵	Daniel Berlin	2017-09-05	1	-0/+49
\| \| \| \| \| \|	aggregate value simplification llvm-svn: 312509
*	[x86] add tests for vector store merge opportunity; NFC	Sanjay Patel	2017-09-04	1	-0/+139
\| \| \| \|	llvm-svn: 312504
*	[x86] auto-generate complete checks; NFC	Sanjay Patel	2017-09-04	1	-7/+21
\| \| \| \|	llvm-svn: 312503
*	[x86] add/regenerate complete checks; NFC	Sanjay Patel	2017-09-04	3	-78/+146
\| \| \| \|	llvm-svn: 312502
*	[x86] add test for unnecessary cmp + masked store; NFC	Sanjay Patel	2017-09-04	1	-0/+28
\| \| \| \| \| \| \| \| \|	As noted in PR11210: https://bugs.llvm.org/show_bug.cgi?id=11210 ...fixing this should allow us to eliminate x86-specific masked store intrinsics in IR. (Although more testing will be needed to confirm that.) llvm-svn: 312496
*	Revert "Re-enable "[MachineCopyPropagation] Extend pass to do COPY source ↵	Sam McCall	2017-09-04	73	-303/+320
\| \| \| \| \| \| \| \| \| \|	forwarding"" This crashes on boringSSL on PPC (will send reduced testcase) This reverts commit r312328. llvm-svn: 312490
*	Fix test/Transforms/GlobalOpt/integer-bool-dwarf	Strahinja Petrovic	2017-09-04	1	-11/+2
\| \| \| \| \| \| \| \| \|	This patch fixes regression related with integer-bool-dwarf test. Patch by Nikola Prica. llvm-svn: 312489
*	Update test for testing avx512	Michael Zuckerman	2017-09-04	1	-26/+26
\| \| \| \|	llvm-svn: 312487