bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
...
*	80-column.	Rui Ueyama	2014-01-17	1	-3/+6
\| \| \| \|	llvm-svn: 199519
*	llvm-objdump/COFF: Print ordinal base number.	Rui Ueyama	2014-01-17	1	-0/+6
\| \| \| \|	llvm-svn: 199518
*	Add two new calling conventions for runtime calls	Juergen Ributzka	2014-01-17	6	-0/+31
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds two new target-independent calling conventions for runtime calls - PreserveMost and PreserveAll. The target-specific implementation for X86-64 is defined as following: - Arguments are passed as for the default C calling convention - The same applies for the return value(s) - PreserveMost preserves all GPRs - except R11 - PreserveAll preserves all GPRs and all XMMs/YMMs - except R11 Reviewed by Lang and Philip llvm-svn: 199508
*	[mips][msa] Correct pattern for LSA	Daniel Sanders	2014-01-17	1	-2/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: $rs and $rt were the wrong way round in the .td and the testcase wasn't strict enough to detect the mistake. Reviewers: matheusalmeida Reviewed By: matheusalmeida Differential Revision: http://llvm-reviews.chandlerc.com/D2554 llvm-svn: 199498
*	[mips] Split IIIdiv int II_DIV, II_DIVU, II_DDIV, and II_DDIVU	Daniel Sanders	2014-01-17	4	-12/+22
\| \| \| \| \| \|	No functional change since the InstrItinData's were duplicated llvm-svn: 199497
*	[mips][sched] Split IIImul and IIImult into subclasses.	Daniel Sanders	2014-01-17	4	-31/+46
\| \| \| \| \| \| \| \| \|	IIImul -> II_MUL IIImult -> II_MULT, II_MULTU, II_MADD, II_MADDU, II_MSUB, II_MSUBU, II_DMULT, II_DMULTU No functional change since the InstrItinData's have been duplicated. llvm-svn: 199495
*	[mips][sched] Split IIHiLo into II_MFHI_MFLO and II_MTHI_MTLO	Daniel Sanders	2014-01-17	2	-7/+10
\| \| \| \| \| \|	No functional change since the InstrItinData's have been duplicated. llvm-svn: 199493
*	Add MLA alias for ARMv4 support.	Renato Golin	2014-01-17	1	-9/+14
\| \| \| \| \| \| \| \| \| \|	Fix MLA defs to use register class GPRnopc. Add encoding tests for multiply instructions. (Alias for MUL/SMLAL/UMLAL added by r199026.) Patch by Zhaoshi. llvm-svn: 199491
*	[PM] [cleanup] Rename some of the Verifier's members, re-arrange them,	Chandler Carruth	2014-01-17	1	-273/+269
\| \| \| \| \| \| \| \| \| \| \|	and tweak comments prior to more invasive surgery. Also clean up some other non-doxygen comments, and run clang-format over the parts that are going to change dramatically in subsequent commits so that those don't get cluttered with formatting changes. No functionality changed. llvm-svn: 199489
*	[asan] extend asan-coverage (still experimental).	Kostya Serebryany	2014-01-17	1	-31/+48
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	- add a mode for collecting per-block coverage (-asan-coverage=2). So far the implementation is naive (all blocks are instrumented), the performance overhead on top of asan could be as high as 30%. - Make sure the one-time calls to __sanitizer_cov are moved to function buttom, which in turn required to copy the original debug info into the call insn. Here is the performance data on SPEC 2006 (train data, comparing asan with asan-coverage={0,1,2}): asan+cov0 asan+cov1 diff 0-1 asan+cov2 diff 0-2 diff 1-2 400.perlbench, 65.60, 65.80, 1.00, 76.20, 1.16, 1.16 401.bzip2, 65.10, 65.50, 1.01, 75.90, 1.17, 1.16 403.gcc, 1.64, 1.69, 1.03, 2.04, 1.24, 1.21 429.mcf, 21.90, 22.60, 1.03, 23.20, 1.06, 1.03 445.gobmk, 166.00, 169.00, 1.02, 205.00, 1.23, 1.21 456.hmmer, 88.30, 87.90, 1.00, 91.00, 1.03, 1.04 458.sjeng, 210.00, 222.00, 1.06, 258.00, 1.23, 1.16 462.libquantum, 1.73, 1.75, 1.01, 2.11, 1.22, 1.21 464.h264ref, 147.00, 152.00, 1.03, 160.00, 1.09, 1.05 471.omnetpp, 115.00, 116.00, 1.01, 140.00, 1.22, 1.21 473.astar, 133.00, 131.00, 0.98, 142.00, 1.07, 1.08 483.xalancbmk, 118.00, 120.00, 1.02, 154.00, 1.31, 1.28 433.milc, 19.80, 20.00, 1.01, 20.10, 1.02, 1.01 444.namd, 16.20, 16.20, 1.00, 17.60, 1.09, 1.09 447.dealII, 41.80, 42.20, 1.01, 43.50, 1.04, 1.03 450.soplex, 7.51, 7.82, 1.04, 8.25, 1.10, 1.05 453.povray, 14.00, 14.40, 1.03, 15.80, 1.13, 1.10 470.lbm, 33.30, 34.10, 1.02, 34.10, 1.02, 1.00 482.sphinx3, 12.40, 12.30, 0.99, 13.00, 1.05, 1.06 llvm-svn: 199488
*	[PM] Remove the preverifier and directly compute the DominatorTree for	Chandler Carruth	2014-01-17	2	-54/+27
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	the verifier after ensuring the CFG is at least usefully formed. This fixes a number of problems: 1) The PreVerifier was missing the controls the Verifier provides over how an invalid module is handled -- it just aborted the program! Now it uses the same logic as the Verifier which is significantly more library-friendly. 2) The DominatorTree used previously could have been cached and not updated due to bugs in prior passes and we would silently use the stale tree. This could cause dominance errors to not be as quickly diagnosed. 3) We can now (in the next patch) pull the functionality of the verifier apart from the pass infrastructure so that you can verify IR without having any form of pass manager. This in turn frees the code to share logic between old and new pass manager variants. Along the way I fixed at least one annoying bug -- the state for 'Broken' wasn't being cleared from run to run causing all functions visited after the first broken function to be marked as broken regardless of whether they were a problem. Fortunately, I don't really know much of a way to observe this peculiarity. In case folks are worried about the runtime cost, its negligible. I looked at running the entire regression test suite (which should be a relatively good use of the verifier) before and after but was unable to even measure the time spent on the verifier and there was no regresion from before to after. I checked both with debug builds and optimized builds. llvm-svn: 199487
*	[AArch64 NEON] Expand vector for UDIV/SDIV/UREM/SREM/FREM as neon doesn't ↵	Kevin Qin	2014-01-17	1	-0/+55
\| \| \| \| \| \|	support these operations. llvm-svn: 199485
*	Switch a few instructions to use RI instead I so they don't require REX_W to ↵	Craig Topper	2014-01-17	3	-19/+19
\| \| \| \| \| \|	be explicitly specified. llvm-svn: 199479
*	Add OpSize16 flags to 32-bit CRC32 instructions so they can be encoded ↵	Craig Topper	2014-01-17	1	-2/+2
\| \| \| \| \| \|	correctly in 16-bit mode. llvm-svn: 199478
*	Teach x86 asm parser to handle 'opaque ptr' in Intel syntax.	Craig Topper	2014-01-17	1	-0/+1
\| \| \| \|	llvm-svn: 199477
*	Teach X86 asm parser to understand 'ZMMWORD PTR' in Intel syntax.	Craig Topper	2014-01-17	1	-0/+1
\| \| \| \|	llvm-svn: 199476
*	Fix intel syntax for 64-bit version of FXSAVE/FXRSTOR to use '64' suffix ↵	Craig Topper	2014-01-17	1	-2/+2
\| \| \| \| \| \|	instead of 'q' llvm-svn: 199474
*	VEX_PREFIX_66 doesn't need to set the hasOpSize flag since VEX instructions ↵	Craig Topper	2014-01-17	1	-11/+0
\| \| \| \| \| \|	don't use the size fields it controls. llvm-svn: 199470
*	Replace duplicated code with a existing helper function.	Craig Topper	2014-01-17	1	-16/+1
\| \| \| \|	llvm-svn: 199468
*	[AArch64]Fix the problem can't select f16_to_f32 and f32_to_f16.	Hao Liu	2014-01-17	2	-0/+16
\| \| \| \| \| \| \|	Also add copy support for FPR16. Also add a missing test case file belongs to commit r197361. llvm-svn: 199463
*	[AArch64 NEON] Custom lower conversion between vector integer and vector ↵	Kevin Qin	2014-01-17	1	-0/+94
\| \| \| \| \| \|	floating point if element bit-width doesn't match. llvm-svn: 199462
*	[AArch64]Fix the problem can't select concat_vectors of two v1i32 types.	Hao Liu	2014-01-17	2	-12/+10
\| \| \| \| \| \|	Also fix the problem can't select scalar_to_vector from f32 to v2f32/v4f32. llvm-svn: 199461
*	Change inalloca rules to make it only apply to the last parameter	Reid Kleckner	2014-01-16	1	-24/+9
\| \| \| \| \| \| \| \| \| \| \|	This makes things a lot easier, because we can now talk about the "argument allocation", which allocates all the memory for the call in one shot. The only functional change is to the verifier for a feature that hasn't shipped yet. llvm-svn: 199434
*	[opt][PassInfo] Allow opt to run passes that need target machine.	Quentin Colombet	2014-01-16	2	-5/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When registering a pass, a pass can now specify a second construct that takes as argument a pointer to TargetMachine. The PassInfo class has been updated to reflect that possibility. If such a constructor exists opt will use it instead of the default constructor when instantiating the pass. Since such IR passes are supposed to be rare, no specific support has been added to this commit to allow an easy registration of such a pass. In other words, for such pass, the initialization function has to be hand-written (see CodeGenPrepare for instance). Now, codegenprepare can be tested using opt: opt -codegenprepare -mtriple=mytriple input.ll llvm-svn: 199430
*	Fix two cases where we could lose fast math flags when optimizing FADD ↵	Owen Anderson	2014-01-16	1	-4/+10
\| \| \| \| \| \|	expressions. llvm-svn: 199427
*	Fix an instance where we would drop fast math flags when performing an fdiv ↵	Owen Anderson	2014-01-16	1	-1/+3
\| \| \| \| \| \|	to reciprocal multiply transformation. llvm-svn: 199425
*	Fix a bug in InstCombine where we failed to preserve fast math flags when ↵	Owen Anderson	2014-01-16	1	-2/+5
\| \| \| \| \| \|	optimizing an FMUL expression. llvm-svn: 199424
*	llvm-objdump/COFF: Print DLL name in the export table header.	Rui Ueyama	2014-01-16	1	-1/+11
\| \| \| \|	llvm-svn: 199422
*	Teach InstCombine that (fmul X, -1.0) can be simplified to (fneg X), which ↵	Owen Anderson	2014-01-16	1	-0/+10
\| \| \| \| \| \|	LLVM expresses as (fsub -0.0, X). llvm-svn: 199420
*	Use static instead of anonymous namespace.	Rui Ueyama	2014-01-16	1	-8/+4
\| \| \| \|	llvm-svn: 199419
*	Reduce nesting.	Rui Ueyama	2014-01-16	1	-13/+11
\| \| \| \|	llvm-svn: 199418
*	Use the current local variable naming style.	Rui Ueyama	2014-01-16	1	-244/+242
\| \| \| \|	llvm-svn: 199417
*	Tweak the MCExternalSymbolizer to print references to C string literals	Kevin Enderby	2014-01-16	1	-2/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	with raw_ostream's write_escaped() method. For example darwin's otool(1) program that uses the llvm disassembler now produces disassembly like this: leaq 0x7b(%rip), %rdi ## literal pool for: "%f\ntoto\n" and not print the new lines which messes up the output. rdar://15145300 llvm-svn: 199407
*	[mips][sched] Removed IIXfer. No instructions use it.	Daniel Sanders	2014-01-16	1	-2/+0
\| \| \| \|	llvm-svn: 199403
*	[mips][sched] Put AND, OR, XOR, MOVT_I, and MOVF_I in the same itinerary ↵	Daniel Sanders	2014-01-16	1	-5/+5
\| \| \| \| \| \| \| \|	class as their non-microMIPS counterparts. No functional change since both classes have the same InstrItinData definition. llvm-svn: 199402
*	Add an emitRawComment function and use it to simplify some uses of EmitRawText.	Rafael Espindola	2014-01-16	5	-22/+22
\| \| \| \|	llvm-svn: 199397
*	[mips][sched] Split IIseb into II_SEB and II_SEH	Daniel Sanders	2014-01-16	4	-9/+11
\| \| \| \| \| \|	No functional change since there are no InstrItinData's. llvm-svn: 199396
*	[mips][sched] Split IILogic into II_AND, II_OR, II_XOR, II_ANDI, II_ORI, II_XORI	Daniel Sanders	2014-01-16	3	-14/+14
\| \| \| \| \| \| \| \|	This is necessary because the classes are shared between all implementations. No functional change since the InstrItinData's have been duplicated. llvm-svn: 199394
*	[mips][sched] Split IIArith in preparation for the first scheduler targeting ↵	Daniel Sanders	2014-01-16	5	-69/+154
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	a specific MIPS CPU. IIArith -> II_ADD, II_ADDU, II_AND, II_CL[ZO], II_DADDIU, II_DADDU, II_DROTR, II_DROTR32, II_DROTRV, II_DSLL, II_DSLL32, II_DSLLV, II_DSR[AL], II_DSR[AL]32, II_DSR[AL]V, II_DSUBU, II_LUI, II_MOV[ZFNT], II_NOR, II_OR, II_RDHWR, II_ROTR, II_ROTRV, II_SLL, II_SLLV, II_SR[AL], II_SR[AL]V, II_SUBU, II_XOR No functional change since the InstrItinData's have been duplicated. This is necessary because the classes are shared between all schedulers. Once this patch series is committed there will be an InstrItinClass for each mnemonic with minimal grouping. This does increase the size of the itinerary tables for each MIPS scheduler but we have a few options for dealing with that later. These options include reducing the number of classes once we see the best way to simplify them, or by extending tablegen to be able to compress the table by eliminating duplicates entries, etc. llvm-svn: 199391
*	[mips] Correct itin class for MULT_MM and MULTu_MM to IIImult.	Daniel Sanders	2014-01-16	1	-2/+2
\| \| \| \| \| \| \|	This matches the itin class used by the non-microMIPS equivalents of these instructions. llvm-svn: 199389
*	[mips] IIImult should have an InstrItinData in the generic scheduler. Used ↵	Daniel Sanders	2014-01-16	1	-0/+1
\| \| \| \| \| \| \| \| \| \| \| \| \|	the same one as for IIImul. Affects: DMULT, DMULTu, MADD, MADD_MM, MADDU, MADDU_MM, MSUB, MSUB_MM, MSUBU, MSUBU_MM, MULT, MULTu Does not affect MULT_MM, MULTu_MM since they are currently miscategorised as IIImul. llvm-svn: 199381
*	ReMat: fix overly cavalier attitude to sub-register indices	Tim Northover	2014-01-16	1	-24/+21
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	There are two attempted optimisations in reMaterializeTrivialDef, trying to avoid promoting the size of a register too much when rematerializing. Unfortunately, both appear to be flawed. First, we see if the original register would have worked, but this is inadequate. Consider: v1 = SOMETHING (v1 is QQ) v2:Q0 = COPY v1:Q1 (v1, v2 are QQ) ... uses of v2 In this case even though v2 could be used directly as the output of SOMETHING, this would set the wrong bits of the QQ register involved. The correct rematerialization must be: v2:Q0_Q1 = SOMETHING (v2 promoted to QQQ) ... uses of v2:Q1_Q2 For the second optimisation, if the correct remat is "v2:idx = SOMETHING" then we can't necessarily expect v2 itself to be valid for SOMETHING, but we do try to hunt for a class between v1 and v2 that works. Unfortunately, this is also wrong: v1 = SOMETHING (v1 is QQ) v2:Q0_Q1 = COPY v1 (v1 is QQ, v2 is QQQ) ... uses of v2 as a QQQ The canonical rematerialization here is "v2:Q0_Q1 = SOMETHING". However current logic would decide that v2 could be a QQ (no interest is taken in later uses). This patch, therefore, always accepts the widened register class without trying to be clever. Generally there is no penalty to this (e.g. in the common GR32 < GR64 case, expanding the width doesn't matter because it's not like you were going to do anything else with the high bits of a GR32 register). It can increase register pressure in cases like the ARM VFP regs though (multiple non-overlapping but equivalent subregisters). This situation can be spotted by the fact that both source and destination in the not-quite-coalesced pair have a sub-register index and rematerialisation is skipped in that situation. Unfortunately, no in-tree targets actually expose this as far as I can tell (there are so few isAsCheapAsAMove instructions for it to trigger on) so I've been unable to produce a test. It was exposed in our ARM64 SPEC tests though, and I will be adding a test there that we should be able to contribute soon(TM). rdar://problem/15775279 llvm-svn: 199376
*	[asan] Remove -fsanitize-address-zero-base-shadow command line	Evgeniy Stepanov	2014-01-16	1	-22/+14
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	flag from clang, and disable zero-base shadow support on all platforms where it is not the default behavior. - It is completely unused, as far as we know. - It is ABI-incompatible with non-zero-base shadow, which means all objects in a process must be built with the same setting. Failing to do so results in a segmentation fault at runtime. - It introduces a backward dependency of compiler-rt on user code, which is uncommon and complicates testing. This is the LLVM part of a larger change. llvm-svn: 199371
*	For ARM, fix assertuib failures for some ld/st 3/4 instruction with wirteback.	Jiangning Liu	2014-01-16	4	-11/+78
\| \| \| \|	llvm-svn: 199369
*	AVX-512: fixed a compare pattern	Elena Demikhovsky	2014-01-16	1	-6/+10
\| \| \| \|	llvm-svn: 199366
*	Copy segment register when optimizing to MOV8ao8/MOV16ao16/MOV32ao32.	Craig Topper	2014-01-16	1	-1/+2
\| \| \| \|	llvm-svn: 199365
*	Allow x86 mov instructions to/from memory with absolute address to be ↵	Craig Topper	2014-01-16	8	-54/+92
\| \| \| \| \| \|	encoded and disassembled with a segment override prefix. Fixes PR16962. llvm-svn: 199364
*	Use a slightly smaller hack.	Rafael Espindola	2014-01-16	1	-2/+1
\| \| \| \|	llvm-svn: 199363
*	llmv-objdump/COFF: Print export table contents.	Rui Ueyama	2014-01-16	1	-3/+95
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds the capability to dump export table contents. An example output is this: Export Table: Ordinal RVA Name 5 0x2008 exportfn1 6 0x2010 exportfn2 By adding this feature to llvm-objdump, we will be able to use it to check export table contents in LLD's tests. Currently we are doing binary comparison in the tests, which is fragile and not readable to humans. llvm-svn: 199358
*	CommentColumn is always 40. Simplify.	Rafael Espindola	2014-01-16	2	-2/+0
\| \| \| \|	llvm-svn: 199357