bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
...
*	PPC: Mark vector CC action for SETO and SETONE as Expand	Hal Finkel	2013-07-08	1	-0/+3
\| \| \| \| \| \| \| \|	Another bug found by llvm-stress! This fixes hitting llvm_unreachable("Invalid integer vector compare condition"); at the end of getVCmpInst in PPCISelDAGToDAG. llvm-svn: 185855
*	PPC: Mark vector FREM as Expand by default	Hal Finkel	2013-07-08	1	-0/+1
\| \| \| \| \| \| \|	Another bug found by llvm-stress! This fixes crashing with: LLVM ERROR: Cannot select: v4f32 = frem ... llvm-svn: 185840
*	[PowerPC] Fix PR16556 (handle undef ppcf128 in LowerFP_TO_INT).	Bill Schmidt	2013-07-08	1	-0/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	PPCTargetLowering::LowerFP_TO_INT() expects its source operand to be either an f32 or f64, but this is not checked. A long double (ppcf128) operand will normally be custom-lowered to a conversion to f64 in this context. However, this isn't the case for an UNDEF node. This patch recognizes a ppcf128 as a legal source operand for FP_TO_INT only if it's an undef, in which case it creates an undef of the target type. At some point we might want to do a wholesale custom lowering of ISD::UNDEF when the type is ppcf128, but it's not really clear that's a great idea, and probably more work than it's worth for a situation that only arises in the case of a programming error. At this point I think simple is best. The test case comes from PR16556, and is a crash-test only. llvm-svn: 185821
*	[PowerPC] Support @tls in the asm parser	Ulrich Weigand	2013-07-05	1	-1/+3
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This adds support for the last missing construct to parse TLS-related assembler code: add 3, 4, symbol@tls The ADD8TLS currently hard-codes the @tls into the assembler string. This cannot be handled by the asm parser, since @tls is parsed as a symbol variant. This patch changes ADD8TLS to have the @tls suffix printed as symbol variant on output too, which allows us to remove the isCodeGenOnly marker from ADD8TLS. This in turn means that we can add a AsmOperand to accept @tls marked symbols on input. As a side effect, this means that the fixup_ppc_tlsreg fixup type is no longer necessary and can be merged into fixup_ppc_nofixup. llvm-svn: 185692
*	Remove the EXCEPTIONADDR, EHSELECTION, and LSDAADDR ISD opcodes.	Jakob Stoklund Olesen	2013-07-04	1	-5/+0
\| \| \| \| \| \|	These exception-related opcodes are not used any longer. llvm-svn: 185625
*	Revert r185595-185596 which broke buildbots.	Jakob Stoklund Olesen	2013-07-04	1	-0/+5
\| \| \| \| \| \| \|	Revert "Simplify landing pad lowering." Revert "Remove the EXCEPTIONADDR, EHSELECTION, and LSDAADDR ISD opcodes." llvm-svn: 185600
*	Remove the EXCEPTIONADDR, EHSELECTION, and LSDAADDR ISD opcodes.	Jakob Stoklund Olesen	2013-07-03	1	-5/+0
\| \| \| \| \| \|	These exception-related opcodes are not used any longer. llvm-svn: 185596
*	[PowerPC] Always use mfocrf if available	Ulrich Weigand	2013-07-03	1	-6/+6
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When accessing just a single CR register, it is always preferable to use mfocrf instead of mfcr, if the former is available on the CPU. Current code makes that distinction in many, but not all places where a single CR register value is retrieved. One missing location is PPCRegisterInfo::lowerCRSpilling. To fix this and make this simpler in the future, this patch changes the bulk of the back-end to always assume mfocrf is available and simply generate it when needed. On machines that actually do not support mfocrf, the instruction is replaced by mfcr at the very end, in EmitInstruction. This has the additional benefit that we no longer need the MFCRpseud hack, since before EmitInstruction we always have a MFOCRF instruction pattern, which already models data flow as required. The patch also adds the MFOCRF8 version of the instruction, which was missing so far. Except for the PPCRegisterInfo::lowerCRSpilling case, no change in generated code intended. llvm-svn: 185556
*	The getRegForInlineAsmConstraint function should only accept MVT value types.	Chad Rosier	2013-06-22	1	-1/+1
\| \| \| \|	llvm-svn: 184642
*	[PowerPC] Rename some more VK_PPC_ enums	Ulrich Weigand	2013-06-21	1	-4/+4
\| \| \| \| \| \| \| \| \| \| \| \| \|	This renames more VK_PPC_ enums, to make them more closely reflect the @modifier string they represent. This also prepares for adding a bunch of new VK_PPC_ enums in upcoming patches. For consistency, some MO_ flags related to VK_PPC_ enums are likewise renamed. No change in behaviour. llvm-svn: 184547
*	[PowerPC] Expose some calling convention functions in PPCISelLowering.h.	Bill Schmidt	2013-06-12	1	-29/+14
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This is a preparatory patch for fast-isel support. The instruction selector will need to access some functions in PPCGenCallingConv.inc, which in turn requires several helper functions to be defined. These are currently defined near the only use of PCCGenCallingConv.inc, inside PPCISelLowering.cpp. This patch moves the declaration of the functions into the associated header file to provide the needed visibility. No functional change intended. llvm-svn: 183844
*	Don't cache the instruction and register info from the TargetMachine, because	Bill Wendling	2013-06-07	1	-5/+7
\| \| \| \| \| \| \| \|	the internals of TargetMachine could change. No functionality change intended. llvm-svn: 183494
*	Order CALLSEQ_START and CALLSEQ_END nodes.	Andrew Trick	2013-05-29	1	-7/+12
\| \| \| \| \| \| \| \| \| \| \| \|	Fixes PR16146: gdb.base__call-ar-st.exp fails after pre-RA-sched=source fixes. Patch by Xiaoyi Guo! This also fixes an unsupported dbg.value test case. Codegen was previously incorrect but the test was passing by luck. llvm-svn: 182885
*	PPC: Add a isConsecutiveLS utility function	Hal Finkel	2013-05-27	1	-2/+42
\| \| \| \| \| \| \| \| \| \| \| \| \|	isConsecutiveLS is a slightly more general form of SelectionDAG::isConsecutiveLoad. Aside from also handling stores, it also does not assume equality of the chain operands is necessary. In the case of the PPC backend, this chain condition is checked in a more general way by the surrounding code. Mostly, this part of the refactoring in preparation for supporting optimized unaligned stores. llvm-svn: 182723
*	Prefer to duplicate PPC Altivec loads when expanding unaligned loads	Hal Finkel	2013-05-26	1	-3/+79
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When expanding unaligned Altivec loads, we use the decremented offset trick to prevent page faults. Unfortunately, if we have a sequence of consecutive unaligned loads, this leads to suboptimal code generation because the 'extra' load from the first unaligned load can be combined with the base load from the second (but only if the decremented offset trick is not used for the first). Search up and down the chain, through loads and token factors, looking for consecutive loads, and if one is found, don't use the offset reduction trick. These duplicate loads are later combined to yield the desired sequence (in the future, we might want a more-powerful chain search, but that will require some changes to allow the combiner routines to access the AA object). This should complete the initial implementation of the optimized unaligned Altivec load expansion. There is some refactoring that should be done, but that will happen when the unaligned store expansion is added. llvm-svn: 182719
*	PPC: Combine duplicate (offset) lvsl Altivec intrinsics	Hal Finkel	2013-05-25	1	-1/+28
\| \| \| \| \| \| \| \| \|	The lvsl permutation control instruction is a function only of the alignment of the pointer operand (relative to the 16-byte natural alignment of Altivec vectors). As a result, multiple lvsl intrinsics where the operands differ by a multiple of 16 can be combined. llvm-svn: 182708
*	Track IR ordering of SelectionDAG nodes 2/4.	Andrew Trick	2013-05-25	1	-62/+61
\| \| \| \| \| \| \|	Change SelectionDAG::getXXXNode() interfaces as well as call sites of these functions to pass in SDLoc instead of DebugLoc. llvm-svn: 182703
*	PPC: Initial support for permutation-based unaligned Altivec loads	Hal Finkel	2013-05-24	1	-0/+129
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Altivec only directly supports aligned loads, but the loads have a strange property: If given an unaligned address, they truncate the address to the next lower aligned address, and load from there. This property, along with an extra load and some special-purpose permutation-control instructions that generate the appropriate permutations from the original unaligned address, allow efficient lowering of aligned loads. This code uses the trick explained in the Apple Velocity Engine optimization overview document to prevent the needed extra load from possibly causing a page fault if the original address happens to be aligned. As noted in the FIXMEs, there are several additional optimizations that can be performed to reduce the cost of these loads even more. These will be implemented in future commits. llvm-svn: 182691
*	Add LLVMContext argument to getSetCCResultType	Matt Arsenault	2013-05-18	1	-4/+4
\| \| \| \|	llvm-svn: 182180
*	[PowerPC] Use true offset value in "memrix" machine operands	Ulrich Weigand	2013-05-16	1	-102/+20
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This is the second part of the change to always return "true" offset values from getPreIndexedAddressParts, tackling the case of "memrix" type operands. This is about instructions like LD/STD that only have a 14-bit field to encode immediate offsets, which are implicitly extended by two zero bits by the machine, so that in effect we can access 16-bit offsets as long as they are a multiple of 4. The PowerPC back end currently handles such instructions by carrying the 14-bit value (as it will get encoded into the actual machine instructions) in the machine operand fields for such instructions. This means that those values are in fact not the true offset, but rather the offset divided by 4 (and then truncated to an unsigned 14-bit value). Like in the case fixed in r182012, this makes common code operations on such offset values not work as expected. Furthermore, there doesn't really appear to be any strong reason why we should encode machine operands this way. This patch therefore changes the encoding of "memrix" type machine operands to simply contain the "true" offset value as a signed immediate value, while enforcing the rules that it must fit in a 16-bit signed value and must also be a multiple of 4. This change must be made simultaneously in all places that access machine operands of this type. However, just about all those changes make the code simpler; in many cases we can now just share the same code for memri and memrix operands. llvm-svn: 182032
*	[PowerPC] Report true displacement value from getPreIndexedAddressParts	Ulrich Weigand	2013-05-16	1	-2/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	DAGCombiner::CombineToPreIndexedLoadStore calls a target routine to decompose a memory address into a base/offset pair. It expects the offset (if constant) to be the true displacement value in order to perform optional additional optimizations; in particular, to convert other uses of the original pointer into uses of the new base pointer after pre-increment. The PowerPC implementation of getPreIndexedAddressParts, however, simply calls SelectAddressRegImm, which returns a TargetConstant. This value is appropriate for encoding into the instruction, but it is not always usable as true displacement value: - Its type is always MVT::i32, even on 64-bit, where addresses ought to be i64 ... this causes the optimization to simply always fail on 64-bit due to this line in DAGCombiner: // FIXME: In some cases, we can be smarter about this. if (Op1.getValueType() != Offset.getValueType()) { - Its value is truncated to an unsigned 16-bit value if negative. This causes the above opimization to generate wrong code. This patch fixes both problems by simply returning the true displacement value (in its original type). This doesn't affect any other user of the displacement. llvm-svn: 182012
*	Implement PPC counter loops as a late IR-level pass	Hal Finkel	2013-05-15	1	-0/+57
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The old PPCCTRLoops pass, like the Hexagon pass version from which it was derived, could only handle some simple loops in canonical form. We cannot directly adapt the new Hexagon hardware loops pass, however, because the Hexagon pass contains a fundamental assumption that non-constant-trip-count loops will contain a guard, and this is not always true (the result being that incorrect negative counts can be generated). With this commit, we replace the pass with a late IR-level pass which makes use of SE to calculate the backedge-taken counts and safely generate the loop-count expressions (including any necessary max() parts). This IR level pass inserts custom intrinsics that are lowered into the desired decrement-and-branch instructions. The most fragile part of this new implementation is that interfering uses of the counter register must be detected on the IR level (and, on PPC, this also includes any indirect branches in addition to function calls). Also, to make all of this work, we need a variant of the mtctr instruction that is marked as having side effects. Without this, machine-code level CSE, DCE, etc. illegally transform the resulting code. Hopefully, this can be improved in the future. This new pass is smaller than the original (and much smaller than the new Hexagon hardware loops pass), and can handle many additional cases correctly. In addition, the preheader-creation code has been copied from LoopSimplify, and after we decide on where it belongs, this code will be refactored so that it can be explicitly shared (making this implementation even smaller). The new test-case files ctrloop-{le,lt,ne}.ll have been adapted from tests for the new Hexagon pass. There are a few classes of loops that this pass does not transform (noted by FIXMEs in the files), but these deficiencies can be addressed within the SE infrastructure (thus helping many other passes as well). llvm-svn: 181927
*	Implement the PowerPC system call (sc) instruction.	Bill Schmidt	2013-05-14	1	-0/+1
\| \| \| \| \| \|	Instruction added at request of Roman Divacky. Tested via asm-parser. llvm-svn: 181821
*	PPC64: Constant initializers with dynamic relocations go in .data.rel.ro.	Bill Schmidt	2013-05-13	1	-0/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This fixes warning messages observed in the oggenc application test in projects/test-suite. Special handling is needed for the 64-bit PowerPC SVR4 ABI when a constant is initialized with a pointer to a function in a shared library. Because a function address is implemented as the address of a function descriptor, the use of copy relocations can lead to problems with initialization. GNU ld therefore replaces copy relocations with dynamic relocations to be resolved by the dynamic linker. This means the constant cannot reside in the read-only data section, but instead belongs in .data.rel.ro, which is designed for constants containing dynamic relocations. The implementation creates a class PPC64LinuxTargetObjectFile inheriting from TargetLoweringObjectFileELF, which behaves like its parent except to place constants of this sort into .data.rel.ro. The test case is reduced from the oggenc application. llvm-svn: 181723
*	Remove unused isLegalAddressImmediate() method.	Roman Divacky	2013-05-08	1	-12/+0
\| \| \| \|	llvm-svn: 181452
*	Fix handling of anonymous aggregate parameters for powerpc*-apple-darwin8.	Bill Schmidt	2013-05-08	1	-4/+4
\| \| \| \| \| \| \| \|	This fixes bug 15821 similarly to the powerpc64-linux fix for bug 14779. Patch by David Fang. llvm-svn: 181449
*	Change commentary for PowerPC Boolean vector contents.	Bill Schmidt	2013-04-23	1	-1/+2
\| \| \| \| \| \|	No functional change intended. llvm-svn: 180131
*	DAGCombine should not aggressively fold SEXT(VSETCC(...)) into a wider ↵	Owen Anderson	2013-04-23	1	-1/+1
\| \| \| \| \| \| \| \| \|	VSETCC without first checking the target's vector boolean contents. This exposed an issue with PowerPC AltiVec where it appears it was setting the wrong vector boolean contents. The included change fixes the PowerPC tests, and was OK'd by Hal. llvm-svn: 180129
*	Cleanup and improve PPC fsel generation	Hal Finkel	2013-04-07	1	-7/+33
\| \| \| \| \| \| \| \| \| \| \| \| \|	First, we should not cheat: fsel-based lowering of select_cc is a finite-math-only optimization (the ISA manual, section F.3 of v2.06, makes this clear, as does a note in our own README). This also adds fsel-based lowering of EQ and NE condition codes. As it turned out, fsel generation was covered by a grand total of zero regression test cases. I've added some test cases to cover the existing behavior (which is now finite-math only), as well as the new EQ cases. llvm-svn: 179000
*	Enable early if conversion on PPC	Hal Finkel	2013-04-05	1	-22/+7
\| \| \| \| \| \| \| \| \| \| \| \| \|	On cores for which we know the misprediction penalty, and we have the isel instruction, we can profitably perform early if conversion. This enables us to replace some small branch sequences with selects and avoid the potential stalls from mispredicting the branches. Enabling this feature required implementing canInsertSelect and insertSelect in PPCInstrInfo; isel code in PPCISelLowering was refactored to use these functions as well. llvm-svn: 178926
*	Rename the current PPC BCL definition to BCLalways	Hal Finkel	2013-04-04	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \|	BCL is normally a conditional branch-and-link instruction, but has an unconditional form (which is used in the SjLj code, for example). To make clear that this BCL instruction definition is specifically the special unconditional form (which does not meaningfully take a condition-register input), rename it to BCLalways. No functionality change intended. llvm-svn: 178803
*	PPC: Improve code generation for mixed-precision reciprocal sqrt	Hal Finkel	2013-04-04	1	-0/+27
\| \| \| \| \| \| \| \|	The DAGCombine logic that recognized a/sqrt(b) and transformed it into a multiplication by the reciprocal sqrt did not handle cases where the sqrt and the division were separated by an fpext or fptrunc. llvm-svn: 178801
*	Cleanup PPC reciprocal-estimate functionality	Hal Finkel	2013-04-03	1	-58/+45
\| \| \| \| \| \| \|	Incorporating review feedback from Bill Schmidt on r178617. No functionality change intended. llvm-svn: 178672
*	Fix PR15632: No support for ppcf128 floating-point remainder on PowerPC.	Bill Schmidt	2013-04-03	1	-0/+1
\| \| \| \| \| \| \| \|	For this we need to use a libcall. Previously LLVM didn't implement libcall support for frem, so I've added it in the usual straightforward manner. A test case from the bug report is included. llvm-svn: 178639
*	Use PPC reciprocal estimates with Newton iteration in fast-math mode	Hal Finkel	2013-04-03	1	-2/+205
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When unsafe FP math operations are enabled, we can use the fre[s] and frsqrte[s] instructions, which generate reciprocal (sqrt) estimates, together with some Newton iteration, in order to quickly generate floating-point division and sqrt results. All of these instructions are separately optional, and so each has its own feature flag (except for the Altivec instructions, which are covered under the existing Altivec flag). Doing this is not only faster than using the IEEE-compliant fdiv/fsqrt instructions, but allows these computations to be pipelined with other computations in order to hide their overall latency. I've also added a couple of missing fnmsub patterns which turned out to be missing (but are necessary for good code generation of the Newton iterations). Altivec needs a similar fix, but that will probably be more complicated because fneg is expanded for Altivec's v4f32. llvm-svn: 178617
*	Fix PR15630: Replace faulty stdcx. with stwcx.	Bill Schmidt	2013-04-02	1	-1/+1
\| \| \| \| \| \| \| \| \| \|	When doing a partword atomic operation, a lwarx was being paired with a stdcx. instead of a stwcx. when compiling for a 64-bit target. The target has nothing to do with it in this case; we always need a stwcx. Thanks to Kai Nacke for reporting the problem. llvm-svn: 178559
*	Fix typo in PPCISelLowering	Hal Finkel	2013-04-02	1	-1/+1
\| \| \| \| \| \|	Thanks to Bill Schmidt for finding this in review of r178480. llvm-svn: 178521
*	Fix a bad assert in PPCTargetLowering	Hal Finkel	2013-04-01	1	-2/+2
\| \| \| \|	llvm-svn: 178489
*	Add more PPC floating-point conversion instructions	Hal Finkel	2013-04-01	1	-19/+77
\| \| \| \| \| \| \| \| \|	The P7 and A2 have additional floating-point conversion instructions which allow a direct two-instruction sequence (plus load/store) to convert from all combinations (signed/unsigned i32/i64) <--> (float/double) (on previous cores, only some combinations were directly available). llvm-svn: 178480
*	Add the PPC popcntw instruction	Hal Finkel	2013-04-01	1	-1/+1
\| \| \| \| \| \| \| \| \|	The popcntw instruction is available whenever the popcntd instruction is available, and performs a separate popcnt on the lower and upper 32-bits. Ignoring the high-order count, this can be used for the 32-bit input case (saving on the explicit zero extension otherwise required to use popcntd). llvm-svn: 178470
*	Treat PPCISD::STFIWX like the memory opcode that it is	Hal Finkel	2013-04-01	1	-2/+9
\| \| \| \| \| \| \| \| \| \| \|	PPCISD::STFIWX is really a memory opcode, and so it should come after FIRST_TARGET_MEMORY_OPCODE, and we should use DAG.getMemIntrinsicNode to create nodes using it. No functionality change intended (although there could be optimization benefits from preserving the MMO information). llvm-svn: 178468
*	Add the PPC lfiwax instruction	Hal Finkel	2013-03-31	1	-10/+33
\| \| \| \| \| \| \| \| \|	This instruction is available on modern PPC64 CPUs, and is now used to improve the SINT_TO_FP lowering (by eliminating the need for the separate sign extension instruction and decreasing the amount of needed stack space). llvm-svn: 178446
*	Cleanup PPC(64) i32 -> float/double conversion	Hal Finkel	2013-03-31	1	-15/+7
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The existing SINT_TO_FP code for i32 -> float/double conversion was disabled because it relied on broken EXTSW_32/STD_32 instruction definitions. The original intent had been to enable these 64-bit instructions to be used on CPUs that support them even in 32-bit mode. Unfortunately, this form of lying to the infrastructure was buggy (as explained in the FIXME comment) and had therefore been disabled. This re-enables this functionality, using regular DAG nodes, but only when compiling in 64-bit mode. The old STD_32/EXTSW_32 definitions (which were dead) are removed. llvm-svn: 178438
*	Implement FRINT lowering on PPC using frin	Hal Finkel	2013-03-29	1	-0/+49
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Like nearbyint, rint can be implemented on PPC using the frin instruction. The complication comes from the fact that rint needs to set the FE_INEXACT flag when the result does not equal the input value (and frin does not do that). As a result, we use a custom inserter which, after the rounding, compares the rounded value with the original, and if they differ, explicitly sets the XX bit in the FPSCR register (which corresponds to FE_INEXACT). Once LLVM has better modeling of the floating-point environment we should be able to (often) eliminate this extra complexity. llvm-svn: 178362
*	Remove the old CodePlacementOpt pass.	Benjamin Kramer	2013-03-29	1	-1/+0
\| \| \| \| \| \|	It was superseded by MachineBlockPlacement and disabled by default since LLVM 3.1. llvm-svn: 178349
*	Add PPC FP rounding instructions fri[mnpz]	Hal Finkel	2013-03-29	1	-0/+17
\| \| \| \| \| \| \| \| \|	These instructions are available on the P5x (and later) and on the A2. They implement the standard floating-point rounding operations (floor, trunc, etc.). One caveat: frin (round to nearest) does not implement "ties to even", and so is only enabled in fast-math mode. llvm-svn: 178337
*	Only enable 64-bit bswap DAG combines for PPC64	Hal Finkel	2013-03-28	1	-0/+2
\| \| \| \| \| \| \| \|	Compiling in 32-bit mode on a P7 would assert after 64-bit DAG combines were added for bswap with load/store. This is because these combines are really only valid in 64-bit mode, regardless of the CPU (and this was not being checked). llvm-svn: 178286
*	Fix bad indentation in r178276	Hal Finkel	2013-03-28	1	-2/+1
\| \| \| \| \| \|	Thanks to Bill Schmidt for pointing this out! llvm-svn: 178280
*	Add the PPC64 ldbrx/stdbrx instructions	Hal Finkel	2013-03-28	1	-3/+9
\| \| \| \| \| \| \| \|	These are 64-bit load/store with byte-swap, and available on the P7 and the A2. Like the similar instructions for 16- and 32-bit words, these are matched in the target DAG-combine phase against load/store-bswap pairs. llvm-svn: 178276
*	Add the PPC64 popcntd instruction	Hal Finkel	2013-03-28	1	-2/+8
\| \| \| \| \| \| \|	PPC ISA 2.06 (P7, A2, etc.) has a popcntd instruction. Add this instruction and tell TTI about it so that popcount-loop recognition will know about it. llvm-svn: 178233