bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
*	Revert "Add missing load/store flags to thumb2 instructions."	Pete Cooper	2015-07-16	1	-1/+1
\| \| \| \| \| \| \| \| \| \|	This reverts commit r242300. This is causing buildbot failures which we are investigating. I'll reapply once we know whats going on, but for now want to get the bots green. llvm-svn: 242428
*	Internalize: internalize comdat members as a group, and drop comdat on such ↵	Peter Collingbourne	2015-07-16	1	-0/+52
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	members. Internalizing an individual comdat group member without also internalizing the other members of the comdat can break comdat semantics. For example, if a module contains a reference to an internalized comdat member, and the linker chooses a comdat group from a different object file, this will break the reference to the internalized member. This change causes the internalizer to only internalize comdat members if all other members of the comdat are not externally visible. Once a comdat group has been fully internalized, there is no need to apply comdat rules to its members; later optimization passes (e.g. globaldce) can legally drop individual members of the comdat. So we drop the comdat attribute from all comdat members. Differential Revision: http://reviews.llvm.org/D10679 llvm-svn: 242423
*	Correct lowering of memmove in NVPTX	Eli Bendersky	2015-07-16	1	-22/+96
\| \| \| \| \| \| \| \| \| \|	This fixes https://llvm.org/bugs/show_bug.cgi?id=24056 Also a bit of refactoring along the way. Differential Revision: http://reviews.llvm.org/D11220 llvm-svn: 242413
*	[Codegen] Add intrinsics 'absdiff' and corresponding SDNodes for absolute ↵	James Molloy	2015-07-16	1	-0/+242
\| \| \| \| \| \| \| \| \| \| \| \| \|	difference operation This adds new intrinsics "*absdiff" for absolute difference ops to facilitate efficient code generation for "sum of absolute differences" operation. The patch also contains the introduction of corresponding SDNodes and basic legalization support.Sanity of the generated code is tested on X86. This is 1st of the three patches. Patch by Shahid Asghar-ahmad! llvm-svn: 242409
*	Fix memcheck interval ends for pointers with negative strides	Silviu Baranga	2015-07-16	1	-0/+89
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: The checking pointer grouping algorithm assumes that the starts/ends of the pointers are well formed (start <= end). The runtime memory checking algorithm also assumes this by doing: start0 < end1 && start1 < end0 to detect conflicts. This check only works if start0 <= end0 and start1 <= end1. This change correctly orders the interval ends by either checking the stride (if it is constant) or by using min/max SCEV expressions. Reviewers: anemet, rengolin Subscribers: rengolin, llvm-commits Differential Revision: http://reviews.llvm.org/D11149 llvm-svn: 242400
*	[X86] Test for r242395 (Fix emitPrologue() to make less assumptions about ↵	Michael Kuperstein	2015-07-16	1	-0/+19
\| \| \| \| \| \|	pushes) llvm-svn: 242399
*	[X86] Reapply r240257 : "Allow more call sequences to use push instructions ↵	Michael Kuperstein	2015-07-16	1	-40/+53
\| \| \| \| \| \| \| \| \| \| \|	for argument passing" This allows more call sequences to use pushes instead of movs when optimizing for size. In particular, calling conventions that pass some parameters in registers (e.g. thiscall) are now supported. This should no longer cause miscompiles, now that a bug in emitPrologue was fixed in r242395. llvm-svn: 242398
*	Revert "[X86] Allow more call sequences to use push instructions for ↵	Reid Kleckner	2015-07-16	1	-53/+40
\| \| \| \| \| \| \| \| \| \| \|	argument passing" It miscompiles some code and a reduced test case has been sent to the author. This reverts commit r240257. llvm-svn: 242373
*	Fix broken testcase from r242358.	Alex Lorenz	2015-07-16	1	-1/+1
\| \| \| \| \| \| \|	The testcase failed on non X86 targets, because I forgot to pass the '-march=x86-64' option into llc for one of the X86 specific tests. llvm-svn: 242370
*	[ARM] Define a subtarget feature that is used to avoid using movt/movw	Akira Hatanaka	2015-07-16	2	-5/+50
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	pairs for 32-bit immediates. This change is needed to avoid emitting movt/movw pairs when doing LTO and do so on a per-function basis. Out-of-tree projects currently using cl::opt option -arm-use-movt=0 or false to avoid emitting movt/movw pairs should make changes to add subtarget feature "+no-movt" (see the changes made to clang in r242368). rdar://problem/21529937 Differential Revision: http://reviews.llvm.org/D11026 llvm-svn: 242369
*	Trying to fix the windows bots.	Rafael Espindola	2015-07-16	1	-4/+4
\| \| \| \|	llvm-svn: 242367
*	Fix handling of relative paths in thin archives.	Rafael Espindola	2015-07-16	1	-3/+18
\| \| \| \| \| \|	The member has to end up with a path relative to the archive. llvm-svn: 242362
*	MIR Serialization: Serialize the jump table index operands.	Alex Lorenz	2015-07-15	2	-1/+183
\| \| \| \| \|	Reviewers: Duncan P. N. Exon Smith llvm-svn: 242358
*	MIR Serialization: Serialize the jump table info.	Alex Lorenz	2015-07-15	2	-0/+116
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	The jump table info is serialized using a YAML mapping that contains its kind and a YAML sequence of jump table entries. A jump table entry is a YAML mapping that has an ID and an inline YAML sequence of machine basic block references. The testcase 'CodeGen/MIR/X86/jump-table-info.mir' doesn't have any instructions because one of them contains a jump table index operand. The jump table index operands will be serialized in a follow up patch, and the appropriate instructions will be added to this testcase. Reviewers: Duncan P. N. Exon Smith llvm-svn: 242357
*	Add a test for r242281 from an old patch of mine.	Sean Silva	2015-07-15	1	-0/+37
\| \| \| \| \| \|	This isn't thorough, but should serve as a sanity check. llvm-svn: 242356
*	llvm-ar: Don't write the directory in the string table.	Rafael Espindola	2015-07-15	1	-3/+12
\| \| \| \| \| \| \|	We were already doing the right thing for short file names, but not long ones. llvm-svn: 242354
*	MIR Serialization: Serialize references from the stack objects to named allocas.	Alex Lorenz	2015-07-15	3	-6/+36
\| \| \| \| \| \| \| \| \|	This commit serializes the references to the named LLVM alloca instructions from the stack objects in the machine frame info. This commit adds a field 'Name' to the struct 'yaml::MachineStackObject'. This new field is used to store the name of the alloca instruction when the alloca is present and when it has a name. llvm-svn: 242339
*	Add a "debugger tuning" concept that allows us to fine-tune how we	Paul Robinson	2015-07-15	2	-1/+45
\| \| \| \| \| \| \| \| \| \| \|	emit debug info, according to the preferences of the different debuggers used on various targets. Darwin and FreeBSD default to tuning for LLDB; PS4 defaults to tuning for the SCE (Sony Computer Entertainment) debugger. All others default to GDB. Differential Revision: http://reviews.llvm.org/D8506 llvm-svn: 242338
*	Fix mergefunc infinite loop	JF Bastien	2015-07-15	1	-0/+40
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Self-referential constants containing references to a merged function no longer cause the MergeFunctions pass to infinite loop. Also adds a reproduction IR which would otherwise fail, which was isolated from a similar issue in Chromium. Author: jrkoenig Reviewers: nlewycky, jfb Subscribers: llvm-commits, nlewycky, jfb Differential Revision: http://reviews.llvm.org/D11208 llvm-svn: 242337
*	Handle the error of trying to convert a regular archive to a thin one.	Rafael Espindola	2015-07-15	1	-0/+14
\| \| \| \| \| \|	While at it, test that we can add to a thin archive. llvm-svn: 242330
*	Analyze recursive PHI nodes in BasicAA	Tobias Edler von Koch	2015-07-15	1	-0/+75
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: This patch allows phi nodes like %x = phi [ %incptr, ... ] [ %var, ... ] %incptr = getelementptr %x, 1 to be analyzed by BasicAliasAnalysis. In aliasPHI, we can detect incoming values that are recursive GEPs with a constant offset. Instead of trying to analyze a recursive GEP (and failing), we now ignore it and instead set the size of the memory referenced by the PHINode to UnknownSize. This represents all the possible memory locations the pointer represented by the PHINode could be advanced to by the GEP. For now, this new behavior is turned off by default to allow debugging of performance degradations seen with SPEC/x86 and Hexagon benchmarks. The flag -basicaa-recphi turns it on. Reviewers: hfinkel, sanjoy Subscribers: tobiasvk_caf, sanjoy, llvm-commits Differential Revision: http://reviews.llvm.org/D10368 llvm-svn: 242320
*	Revert "Look through PHIs to find additional register sources"	Bruno Cardoso Lopes	2015-07-15	1	-84/+0
\| \| \| \| \| \| \| \| \| \|	Likely broke compilation on ARM: http://lab.llvm.org:8011/builders/clang-native-arm-lnt/builds/13054 This reverts commit 131ce4a838c081516cbfed039fc986b33e3979d6. llvm-svn: 242310
*	Debug Info: Add basic support for external types references.	Adrian Prantl	2015-07-15	1	-0/+51
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This is a necessary prerequisite for bootstrapping the emission of debug info inside modules. - Adds a FlagExternalTypeRef to DICompositeType. External types must have a unique identifier. - External type references are emitted using a forward declaration with a DW_AT_signature([DW_FORM_ref_sig8]) based on the UID. http://reviews.llvm.org/D9612 llvm-svn: 242302
*	Add missing load/store flags to thumb2 instructions.	Pete Cooper	2015-07-15	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \| \|	These were the cause of a verifier error when building 7zip with -verify-machineinstrs. Running 'make check' with the verifier triggered the same error on the test here so i've updated the test to run the verifier on one of its runs instead of adding a new one. While looking at this code, there was a stale comment that these instructions were only used for disassembly. This probably used to be the case, but they are now used in the 'ARM load / store optimization pass' too. llvm-svn: 242300
*	Look through PHIs to find additional register sources	Bruno Cardoso Lopes	2015-07-15	1	-0/+84
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	- Teaches the ValueTracker in the PeepholeOptimizer to look through PHI instructions. - Add findNextSourceAndRewritePHI method to lookup into multiple sources returnted by the ValueTracker and rewrite PHIs with new sources. With these changes we can find more register sources and rewrite more copies to allow coaslescing of bitcast instructions. Hence, we eliminate unnecessary VR64 <-> GR64 copies in x86, but it could be extended to other archs by marking "isBitcast" on target specific instructions. The x86 example follows: A: psllq %mm1, %mm0 movd %mm0, %r9 jmp C B: por %mm1, %mm0 movd %mm0, %r9 jmp C C: movd %r9, %mm0 pshufw $238, %mm0, %mm0 Becomes: A: psllq %mm1, %mm0 jmp C B: por %mm1, %mm0 jmp C C: pshufw $238, %mm0, %mm0 Differential Revision: http://reviews.llvm.org/D11197 rdar://problem/20404526 llvm-svn: 242295
*	[PPC] Disassemble little endian ppc instructions in the right byte order	Benjamin Kramer	2015-07-15	1	-0/+664
\| \| \| \| \| \|	PR24122. The test is simply a byte swapped version of ppc64-encoding.txt. llvm-svn: 242288
*	[SDAG] Optimize unordered comparison in soft-float mode (patch by Anton ↵	Alexey Bataev	2015-07-15	5	-51/+158
\| \| \| \| \| \| \| \| \| \| \|	Nadolskiy) Current implementation handles unordered comparison poorly in soft-float mode. Consider (a ULE b) which is a <= b. It is lowered to (ledf2(a, b) <= 0 \|\| unorddf2(a, b) != 0) (in general). We can do better job by lowering it to (__gtdf2(a, b) <= 0). Such replacement is true for other CMP's (ult, ugt, uge). In general, we just call same function as for ordered case but negate comparison against zero. Differential Revision: http://reviews.llvm.org/D10804 llvm-svn: 242280
*	[PowerPC] Use the MachineCombiner to reassociate fadd/fmul	Hal Finkel	2015-07-15	1	-0/+188
\| \| \| \| \| \| \| \| \| \| \| \| \|	This is a direct port of the code from the X86 backend (r239486/r240361), which uses the MachineCombiner to reassociate (floating-point) adds/muls to increase ILP, to the PowerPC backend. The rationale is the same. There is a lot of copy-and-paste here between the X86 code and the PowerPC code, and we should extract at least some of this into CodeGen somewhere. However, I don't want to do that until this code is enhanced to handle FMAs as well. After that, we'll be in a better position to extract the common parts. llvm-svn: 242279
*	[AArch64] Fix problems in decoding generic MSR instructions	Petr Pavlu	2015-07-15	1	-0/+4
\| \| \| \| \| \| \| \| \|	Bitpatterns rejected by the decoder method of `MSR (immediate)` should be decoded as the `extended MSR (register)` instruction. Differential Revision: http://reviews.llvm.org/D7174 llvm-svn: 242276
*	[TableGen] Improve decoding options for non-orthogonal instructions	Petr Pavlu	2015-07-15	3	-0/+131
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When FixedLenDecoder matches an input bitpattern of form [01]+ with an instruction bitpattern of form [01?]+ (where 0/1 are static bits and ? are mixed/variable bits) it passes the input bitpattern to a specific instruction decoder method which then makes a final decision whether the bitpattern is a valid instruction or not. This means the decoder must handle all possible values of the variable bits which sometimes leads to opcode rewrites in the decoder method when the instructions are not fully orthogonal. The patch provides a way for the decoder method to say that when it returns Fail it does not necessarily mean the bitpattern is invalid, but rather that the bitpattern is definitely not an instruction that is recognized by the decoder method. The decoder can then try to match the input bitpattern with other possible instruction bitpatterns. For example, this allows to solve a situation on AArch64 where the `MSR (immediate)` instruction has form: 1101 0101 0000 0??? 0100 ???? ???1 1111 but not all values of the ? bits are allowed. The rejected values should be handled by the `extended MSR (register)` instruction: 1101 0101 000? ???? ???? ???? ???? ???? The decoder will first try to decode an input bitpattern that matches both bitpatterns as `MSR (immediate)` but currently this puts the decoder method of `MSR (immediate)` into a situation when it must be able to decode all possible values of the ? bits, i.e. it would need to rewrite the instruction to `MSR (register)` when it is not `MSR (immediate)`. The patch allows to specify that the decoder method cannot determine if the instruction is valid for all variable values. The decoder method can simply return Fail when it knows it is definitely not `MSR (immediate)`. The decoder will then backtrack the decoding and find that it can match the input bitpattern with the more generic `MSR (register)` bitpattern too. Differential Revision: http://reviews.llvm.org/D7174 llvm-svn: 242274
*	[X86][SSE] Added i686/SSE2 vector shift tests.	Simon Pilgrim	2015-07-15	3	-34/+987
\| \| \| \| \| \|	We were only testing on x86-64, but we should be ensuring decent code gen of i64 shifts on 32-bit targets. llvm-svn: 242273
*	AVX : Fix ISA disabling in case AVX512VL , some instructions should be ↵	Igor Breger	2015-07-15	1	-0/+215
\| \| \| \| \| \| \| \| \| \|	disabled only if AVX512BW present. Tests added. Differential Revision: http://reviews.llvm.org/D11122 llvm-svn: 242270
*	Initial support for writing thin archives.	Rafael Espindola	2015-07-15	1	-0/+11
\| \| \| \|	llvm-svn: 242269
*	Tidy-up test case from r242257.	Michael Zolotukhin	2015-07-15	1	-5/+8
\| \| \| \|	llvm-svn: 242268
*	[LoopUnrolling] Handle cast instructions.	Michael Zolotukhin	2015-07-15	1	-0/+94
\| \| \| \| \| \| \| \| \|	During estimation of unrolling effect we should be able to propagate constants through casts. Differential Revision: http://reviews.llvm.org/D10207 llvm-svn: 242257
*	WebAssembly: fix build breakage.	JF Bastien	2015-07-14	1	-3/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: processFunctionBeforeCalleeSavedScan was renamed to determineCalleeSaves and now takes a BitVector parameter as of rL242165, reviewed in http://reviews.llvm.org/D10909 WebAssembly is still marked as experimental and therefore doesn't build by default. It does, however, grep by default! I notice that processFunctionBeforeCalleeSavedScan is still mentioned in a few comments and error messages, which I also fixed. Reviewers: qcolombet, sunfish Subscribers: jfb, dsanders, hfinkel, MatzeB, llvm-commits Differential Revision: http://reviews.llvm.org/D11199 llvm-svn: 242242
*	[PowerPC] Support symbolic targets in patchpoints	Hal Finkel	2015-07-14	1	-0/+15
\| \| \| \| \| \| \|	Follow-up r235483, with the corresponding support in PPC. We use a regular call for symbolic targets (because they're much cheaper than indirect calls). llvm-svn: 242239
*	Accept lower case to handle windows error messages.	Rafael Espindola	2015-07-14	1	-1/+1
\| \| \| \|	llvm-svn: 242236
*	[InstCombine] Generalize sub of selects optimization to all BinaryOperators	David Majnemer	2015-07-14	1	-0/+10
\| \| \| \| \| \| \|	This exposes further optimization opportunities if the selects are correlated. llvm-svn: 242235
*	[PowerPC] Use the ABI indirect-call protocol for patchpoints	Hal Finkel	2015-07-14	3	-19/+31
\| \| \| \| \| \| \| \| \| \| \| \|	We used to take the address specified as the direct target of the patchpoint and did no TOC-pointer handling. This, however, as not all that useful, because MCJIT tends to create a lot of modules, and they have their own TOC sections. Thus, to call from the generated code to other generated code, you really need to switch TOC pointers. Make this work as expected, and under ELFv1, tread the address as the function descriptor address so that the correct TOC pointer can be loaded. llvm-svn: 242217
*	Add support for reading members out of thin archives.	Rafael Espindola	2015-07-14	1	-0/+6
\| \| \| \| \| \| \| \| \| \|	For now the Archive owns the buffers of the thin archive members. This makes for a simple API, but all the buffers are destructed only when the archive is destructed. This should be fine since we close the files after mmap so we should not hit an open file limit. llvm-svn: 242215
*	MIR Serialization: Serialize the machine basic block live in registers.	Alex Lorenz	2015-07-14	2	-0/+46
\| \| \| \|	llvm-svn: 242204
*	GVN: tolerate an instruction being replaced without existing in the leaderboard	Tim Northover	2015-07-14	1	-0/+29
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Sometimes an incidentally created instruction can duplicate a Value used elsewhere. It then often doesn't end up in the leader table. If it's later removed, we attempt to remove it from the leader table and segfault. Instead we should just ignore the removal request, which won't cause any problems. The reverse situation, where the original instruction is replaced by the new one (which you might think could leave the leader table empty) cannot occur, because the incidental instruction will never be found in the first place. llvm-svn: 242199
*	[PowerPC] Fix the PPCInstrInfo::getInstrLatency implementation	Hal Finkel	2015-07-14	7	-34/+41
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	PowerPC uses itineraries to describe processor pipelines (and dispatch-group restrictions for P7/P8 cores). Unfortunately, the target-independent implementation of TII.getInstrLatency calls ItinData->getStageLatency, and that looks for the largest cycle count in the pipeline for any given instruction. This, however, yields the wrong answer for the PPC itineraries, because we don't encode the full pipeline. Because the functional units are fully pipelined, we only model the initial stages (there are no relevant hazards in the later stages to model), and so the technique employed by getStageLatency does not really work. Instead, we should take the maximum output operand latency, and that's what PPCInstrInfo::getInstrLatency now does. This caused some test-case churn, including two unfortunate side effects. First, the new arrangement of copies we get from function parameters now sometimes blocks VSX FMA mutation (a FIXME has been added to the code and the test cases), and we have one significant test-suite regression: SingleSource/Benchmarks/BenchmarkGame/spectral-norm 56.4185% +/- 18.9398% In this benchmark we have a loop with a vectorized FP divide, and it with the new scheduling both divides end up in the same dispatch group (which in this case seems to cause a problem, although why is not exactly clear). The grouping structure is hard to predict from the bottom of the loop, and there may not be much we can do to fix this. Very few other test-suite performance effects were really significant, but almost all weakly favor this change. However, in light of the issues highlighted above, I've left the old behavior available via a command-line flag. llvm-svn: 242188
*	[Hexagon] Generate instructions for operations on predicate registers	Krzysztof Parzyszek	2015-07-14	2	-0/+49
\| \| \| \| \| \| \|	Convert logical operations on general-purpose registers to the correspon- ding operations on predicate registers. llvm-svn: 242186
*	[CodeGen] Force emission of personality directive if explicitly specified	Keno Fischer	2015-07-14	1	-0/+12
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: Before this change, personality directives were not emitted if there was no invoke left in the function (of course until recently this also meant that we couldn't know what the personality actually was). This patch forces personality directives to still be emitted, unless it is known to be a noop in the absence of invokes, or the user explicitly specified `nounwind` (and not `uwtable`) on the function. Reviewers: majnemer, rnk Subscribers: rnk, llvm-commits Differential Revision: http://reviews.llvm.org/D10884 llvm-svn: 242185
*	AMDGPU: Avoid using 64-bit shift for i64 (shl x, 32)	Matt Arsenault	2015-07-14	4	-16/+82
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This can be done only with moves which theoretically will optimize better later. Although this transform increases the instruction count, it should be code size / cycle count neutral in the worst VALU case. It also seems to slightly improve a couple of testcases due to other DAG combines this exposes. This is probably slightly worse for the SALU case, so it might be better to handle this during moveToVALU, although then you lose some simplifications like the load width reducing in the simple testcase. llvm-svn: 242177
*	AMDGPU/SI: Fix read2 merging into a super register.	Matt Arsenault	2015-07-14	6	-15/+273
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	If the read2 produced was supposed to be writing into a super register, it would use the wrong subregister indices. Fix this by inserting copies, so we only ever write to a vreg_64. Run the register coalescer again to clean this up, although this isn't ideal and often does result in an extra move. Also remove the assert that offset1 > offset0. There isn't a real reason to not allow this other than a minor convenience in the compiler, and it doesn't seem worth the effort of avoiding it. llvm-svn: 242174
*	Add missing builtins to the PPC back end for ABI compliance (vol. 4)	Nemanja Ivanovic	2015-07-14	1	-0/+30
\| \| \| \| \| \| \| \| \|	This patch corresponds to review: http://reviews.llvm.org/D11183 Back end portion of the fourth round of additions to altivec.h. llvm-svn: 242167
*	ARM: add at least one real test for r242123.	Tim Northover	2015-07-14	1	-0/+10
\| \| \| \| \| \| \| \|	The ones committed were orthogonal to the change and would have passed before that revision. What it did do was prevent an assertion failure when generating object files. llvm-svn: 242166