bcm5719-llvm - Project Ortega BCM5719 LLVM

	Commit message (Collapse)	Author	Age	Files	Lines
*	[llvm-objcopy] %python wants to be in quotes, because it might contain spaces	Benjamin Kramer	2018-07-18	4	-4/+4
\| \| \| \|	llvm-svn: 337399
*	[MC] Fix nested macro body parsing	Nirav Dave	2018-07-18	1	-0/+13
\| \| \| \| \| \|	Add missing .rep case in nestlevel checking for macro body parsing. llvm-svn: 337398
*	[mips] Fix predicate for the MipsTruncIntFP pattern	Simon Atanasyan	2018-07-18	1	-1/+2
\| \| \| \| \| \| \| \| \|	This is a follow-up to the rL337171. This patch fixes regression introduced by the r337171 and enables MipsTruncIntFP pattern. Differential revision: https://reviews.llvm.org/D49469 llvm-svn: 337392
*	ARM: switch armv7em triple to hard-float defaults and libcalls.	Tim Northover	2018-07-18	2	-1/+37
\| \| \| \| \| \| \|	We were emitting incorrect calls to libm functions that LLVM had decided it knew about because the default is soft-float. llvm-svn: 337385
*	[AArch64][SVE] Asm: Support for unpredicated FP operations.	Sander de Smalen	2018-07-18	12	-3/+208
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds support for the following unpredicated floating-point instructions: FADD Floating point add FSUB Floating point subtract FMUL Floating point multiplication FTSMUL Floating point trigonometric starting value FRECPS Floating point reciprocal step FRSQRTS Floating point reciprocal square root step The instructions have the following assembly format: fadd z0.h, z1.h, z2.h and have variants for 16, 32 and 64-bit FP elements. llvm-svn: 337383
*	[NFC] Make a test more neat	Max Kazantsev	2018-07-18	1	-2/+5
\| \| \| \|	llvm-svn: 337379
*	[InstCombine] Re-commit: Fold 'check for [no] signed truncation' pattern	Roman Lebedev	2018-07-18	2	-28/+20
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: [[ https://bugs.llvm.org/show_bug.cgi?id=38149 \| PR38149 ]] As discussed in https://reviews.llvm.org/D49179#1158957 and later, the IR for 'check for [no] signed truncation' pattern can be improved: https://rise4fun.com/Alive/gBf ^ that pattern will be produced by Implicit Integer Truncation sanitizer, https://reviews.llvm.org/D48958 https://bugs.llvm.org/show_bug.cgi?id=21530 in signed case, therefore it is probably a good idea to improve it. The DAGCombine will reverse this transform, see https://reviews.llvm.org/D49266 This transform is surprisingly frustrating. This does not deal with non-splat shift amounts, or with undef shift amounts. I've outlined what i think the solution should be: ``` // Potential handling of non-splats: for each element: // * if both are undef, replace with constant 0. // Because (1<<0) is OK and is 1, and ((1<<0)>>1) is also OK and is 0. // * if both are not undef, and are different, bailout. // * else, only one is undef, then pick the non-undef one. ``` This is a re-commit, as the original patch, committed in rL337190 was reverted in rL337344 as it broke chromium build: https://bugs.llvm.org/show_bug.cgi?id=38204 and https://crbug.com/864832 Proofs that the fixed folds are ok: https://rise4fun.com/Alive/VYM Differential Revision: https://reviews.llvm.org/D49320 llvm-svn: 337376
*	[X86][SSE] Add extra scalar fop + blend tests for commuted inputs	Simon Pilgrim	2018-07-18	1	-24/+208
\| \| \| \| \| \|	While working on PR38197, I noticed that we don't make use of FADD/FMUL being able to commute the inputs to support the addps+movss -> addss style combine llvm-svn: 337375
*	Revert "[Sparc] Use the IntPair reg class for r constraints with value type f64"	Daniel Cederman	2018-07-18	1	-9/+0
\| \| \| \| \| \| \| \| \|	This reverts commit 55222c9183c6e07f53a54c4061677734f54feac1. I missed that this patch has a dependency on https://reviews.llvm.org/D49219 that has not been approved yet. llvm-svn: 337373
*	[AArch64][SVE] Asm: Support for UDOT/SDOT instructions.	Sander de Smalen	2018-07-18	4	-0/+180
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The signed/unsigned DOT instructions perform a dot-product on quadtuplets from two source vectors and accumulate the result in the destination register. The instructions come in two forms: Vector form, e.g. sdot z0.s, z1.b, z2.b - signed dot product on four 8-bit quad-tuplets, accumulating results in 32-bit elements. udot z0.d, z1.h, z2.h - unsigned dot product on four 16-bit quad-tuplets, accumulating results in 64-bit elements. Indexed form, e.g. sdot z0.s, z1.b, z2.b[3] - signed dot product on four 8-bit quad-tuplets with specified quadtuplet from second source vector, accumulating results in 32-bit elements. udot z0.d, z1.h, z2.h[1] - dot product on four 16-bit quad-tuplets with specified quadtuplet from second source vector, accumulating results in 64-bit elements. llvm-svn: 337372
*	[Sparc] Use the IntPair reg class for r constraints with value type f64	Daniel Cederman	2018-07-18	1	-0/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: This is how it appears to be handled in GCC and it prevents a "Unknown mismatch" error in the SelectionDAGBuilder. Reviewers: venkatra, jyknight, jrtc27 Reviewed By: jyknight, jrtc27 Subscribers: eraman, fedor.sergeev, jrtc27, llvm-commits Differential Revision: https://reviews.llvm.org/D49218 llvm-svn: 337370
*	[AArch64][SVE] Asm: Integer divide instructions.	Sander de Smalen	2018-07-18	8	-0/+212
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds the following predicated instructions: UDIV Unsigned divide active elements UDIVR Unsigned divide active elements, reverse form. SDIV Signed divide active elements SDIVR Signed divide active elements, reverse form. e.g. udiv z0.s, p0/m, z0.s, z1.s (unsigned divide active elements in z0 by z1, store result in z0) sdivr z0.s, p0/m, z0.s, z1.s (signed divide active elements in z1 by z0, store result in z0) llvm-svn: 337369
*	[NFC][InstCombine] i65 tests for 'check for [no] signed truncation' pattern	Roman Lebedev	2018-07-18	2	-0/+28
\| \| \| \| \| \| \| \|	Those initially broke chromium build: https://bugs.llvm.org/show_bug.cgi?id=38204 and https://crbug.com/864832 llvm-svn: 337364
*	[llvm-objdump] - Stop reporting bogus section IDs.	George Rimar	2018-07-18	1	-0/+25
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Imagine we have a file with few sections, and one of them is .foo with index N != 0. Problem is that when llvm-objdump is given a -section=.foo parameter it lists .foo as a section at index 0. That makes impossible to write test cases which needs to find the index of the particular section, while ignoring dumping of others. The patch fixes that. Differential revision: https://reviews.llvm.org/D49372 llvm-svn: 337361
*	[llvm-readobj] - Teach tool to dump objects with >= SHN_LORESERVE of sections.	George Rimar	2018-07-18	3	-0/+36
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	http://www.sco.com/developers/gabi/2003-12-17/ch4.eheader.html says that e_shnum and/or e_shstrndx may have special values if "the number of sections is greater than or equal to SHN_LORESERVE" or "the section name string table section index is greater than or equal to SHN_LORESERVE (0xff00)" Previously llvm-readobj was unable to dump such files, patch changes that. I had to add a precompiled test case because it does not seem possible to prepare a test using yaml2obj or llvm-mc (not clear how to make .shstrtab to have index >= SHN_LORESERVE). Differential revision: https://reviews.llvm.org/D49369 llvm-svn: 337360
*	Revert test changes part of "Revert "[InstCombine] Fold 'check for [no] ↵	Roman Lebedev	2018-07-18	2	-108/+108
\| \| \| \| \| \| \| \| \| \| \|	signed truncation' pattern"" We want the test to remain good anyway. I think the fix is incoming. This reverts part of commit rL337344. llvm-svn: 337359
*	[AArch64][SVE] Asm: Support for integer MUL instructions.	Sander de Smalen	2018-07-18	6	-0/+220
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds the following instructions: MUL - multiply vectors, e.g. mul z0.h, p0/m, z0.h, z1.h - multiply with immediate, e.g. mul z0.h, z0.h, #127 SMULH - signed multiply returning high half, e.g. smulh z0.h, p0/m, z0.h, z1.h UMULH - unsigned multiply returning high half, e.g. umulh z0.h, p0/m, z0.h, z1.h llvm-svn: 337358
*	[X86] Enable commuting of VUNPCKHPD to VMOVLHPS to enable load folding by ↵	Craig Topper	2018-07-18	2	-14/+13
\| \| \| \| \| \| \| \|	using VMOVLPS with a modified address. This required an annoying amount of tablegen multiclass changes to make only VUNPCKHPDZ128rr commutable. llvm-svn: 337357
*	[X86] Add test case for missed opportunity to commute vunpckhpd to enable ↵	Craig Topper	2018-07-18	1	-0/+16
\| \| \| \| \| \| \| \|	use of vmovlps to fold a load. We do this transform for SSE, but not AVX or AVX512VL. llvm-svn: 337356
*	[X86] Regenerate fma.ll checks using current version of the script which ↵	Craig Topper	2018-07-18	1	-348/+348
\| \| \| \| \| \|	produces different regular expressions on spills and reloads. NFC llvm-svn: 337354
*	[X86] Generate v2f64 X86ISD::UNPCKL/UNPCKH instead of ↵	Craig Topper	2018-07-18	1	-1/+1
\| \| \| \| \| \| \| \| \| \| \| \|	X86ISD::MOVLHPS/MOVHLPS for unary v2f64 {0,0} and {1,1} shuffles with SSE2. I'm trying to restrict the MOVLHPS/MOVHLPS ISD nodes to SSE1 only. With SSE2 we can use unpcks. I believe this will allow some patterns to be cleaned up to require fewer bitcasts. I've put in an odd isel hack to still select MOVHLPS instruction from the unpckh node to avoid changing tests and because movhlps is a shorter encoding. Ideally we'd do execution domain switching on this, but the operands are in the wrong order and are tied. We might be able to try a commute in the domain switching using custom code. We already support domain switching for UNPCKLPD and MOVLHPS. llvm-svn: 337348
*	Introduce codegen for the Signal Processing Engine	Justin Hibbits	2018-07-18	5	-10/+599
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: The Signal Processing Engine (SPE) is found on NXP/Freescale e500v1, e500v2, and several e200 cores. This adds support targeting the e500v2, as this is more common than the e500v1, and is in SoCs still on the market. This patch is very intrusive because the SPE is binary incompatible with the traditional FPU. After discussing with others, the cleanest solution was to make both SPE and FPU features on top of a base PowerPC subset, so all FPU instructions are now wrapped with HasFPU predicates. Supported by this are: * Code generation following the SPE ABI at the LLVM IR level (calling conventions) * Single- and Double-precision math at the level supported by the APU. Still to do: * Vector operations * SPE intrinsics As this changes the Callee-saved register list order, one test, which tests the precise generated code, was updated to account for the new register order. Reviewed by: nemanjai Differential Revision: https://reviews.llvm.org/D44830 llvm-svn: 337347
*	Complete the SPE instruction set patterns	Justin Hibbits	2018-07-18	2	-5/+670
\| \| \| \| \| \| \| \| \|	This is the lead-up to having SPE codegen. Add the rest of the instructions, along with MC tests. Differential Revision: https://reviews.llvm.org/D44829 llvm-svn: 337346
*	Revert "[InstCombine] Fold 'check for [no] signed truncation' pattern"	Bob Haarman	2018-07-18	2	-110/+116
\| \| \| \| \| \| \| \| \|	This reverts r337190 (and a few follow-up commits), which caused the Chromium build to fail. See https://bugs.llvm.org/show_bug.cgi?id=38204 and https://crbug.com/864832 llvm-svn: 337344
*	CodeGen: Don't create address significance table entries for thread-local ↵	Peter Collingbourne	2018-07-18	1	-0/+3
\| \| \| \| \| \| \| \| \| \|	variables. The presence of these symbols in the symbol table can cause symbol type mismatch errors (or undefined symbol errors on emulated TLS targets) and they can't be ICF'd anyway. llvm-svn: 337338
*	[X86] Remove the vector alignment requirement from the patterns added in ↵	Craig Topper	2018-07-17	1	-2/+2
\| \| \| \| \| \| \| \|	r337320. The resulting instruction will only load 64 bits so alignment isn't required. llvm-svn: 337334
*	CodeGen: Add a target option for emitting .addrsig directives for all ↵	Peter Collingbourne	2018-07-17	1	-0/+36
\| \| \| \| \| \| \| \|	address-significant symbols. Differential Revision: https://reviews.llvm.org/D48143 llvm-svn: 337331
*	MC: Implement support for new .addrsig and .addrsig_sym directives.	Peter Collingbourne	2018-07-17	2	-0/+91
\| \| \| \| \| \| \| \| \|	Part of the address-significance tables proposal: http://lists.llvm.org/pipermail/llvm-dev/2018-May/123514.html Differential Revision: https://reviews.llvm.org/D47744 llvm-svn: 337328
*	[X86] Add patterns for folding full vector load into MOVHPS and MOVLPS with ↵	Craig Topper	2018-07-17	1	-3/+1
\| \| \| \| \| \|	SSE1 only. llvm-svn: 337320
*	[X86] Add test case for missed opportunity to use MOVLPS on the SSE1 only ↵	Craig Topper	2018-07-17	1	-0/+12
\| \| \| \| \| \|	targets. llvm-svn: 337319
*	[InstCombine] Preserve debug value when simplifying cast-of-select	Vedant Kumar	2018-07-17	1	-0/+11
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	InstCombine has a cast transform that matches a cast-of-select: Orig = cast (Src = select Cond TV FV) And tries to replace it with a select which has the cast folded in: NewSel = select Cond (cast TV) (cast FV) The combiner does RAUW(Orig, NewSel), so any debug values for Orig would survive the transform. But debug values for Src would be lost. This patch teaches InstCombine to replace all debug uses of Src with NewSel (taking care of doing any necessary DIExpression rewriting). Differential Revision: https://reviews.llvm.org/D49270 llvm-svn: 337310
*	Remove an errant piece of !dbg metadata from a test, NFC	Vedant Kumar	2018-07-17	1	-1/+1
\| \| \| \|	llvm-svn: 337309
*	[x86/SLH] Flesh out the data-invariant instruction table a bit based on ↵	Chandler Carruth	2018-07-17	1	-2/+61
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	feedback from Craig. Summary: The only thing he suggested that I've skipped here is the double-wide multiply instructions. Multiply is an area I'm nervous about there being some hidden data-dependent behavior, and it doesn't seem important for any benchmarks I have, so skipping it and sticking with the minimal multiply support that matches what I know is widely used in existing crypto libraries. We can always add double-wide multiply when we have clarity from vendors about its behavior and guarantees. I've tried to at least cover the fundamentals here with tests, although I've not tried to cover every width or permutation. I can add more tests where folks think it would be helpful. Reviewers: craig.topper Subscribers: sanjoy, mcrosier, hiraditya, llvm-commits Differential Revision: https://reviews.llvm.org/D49413 llvm-svn: 337308
*	[llvm-mca][x86] Add extend, carry-flag and CMP instructions to general ↵	Simon Pilgrim	2018-07-17	10	-10/+1200
\| \| \| \| \| \|	x86_64 resource tests llvm-svn: 337306
*	[llvm-mca][x86] Add MOVBE resource tests to all supporting targets	Simon Pilgrim	2018-07-17	9	-0/+462
\| \| \| \| \| \|	SNB doesn't support MOVBE but the numbers in Generic (which use the SNB model) look sane. llvm-svn: 337305
*	[llvm-mca][x86] Add BSWAP resource tests	Simon Pilgrim	2018-07-17	10	-10/+80
\| \| \| \|	llvm-svn: 337302
*	[WebAssembly] Update WebAssemblyLowerEmscriptenEHSjLj to handle separate ↵	Sam Clegg	2018-07-17	2	-34/+34
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	compilation Previously we were assuming whole program compilation. Now that separate compilation is a thing we need to update this pass. Firstly, it can no longer assert on the existence of malloc and free. This functions might not be in the current translation unit. If we need them then we will generate not imports for them. Secondly the global helper function we create should be marked as weak since we will be generating a separate copy in each translation unit. Finally the names of the symbols used must be unique and fixed since they need to agree across translation units. Differential Revision: https://reviews.llvm.org/D49263 llvm-svn: 337301
*	[llvm-mca][x86] Add displacement-only and additional scale=1 LEA tests	Simon Pilgrim	2018-07-17	1	-1/+82
\| \| \| \|	llvm-svn: 337298
*	[llvm-mca][x86] Add LEA resource tests (PR32326)	Simon Pilgrim	2018-07-17	1	-0/+362
\| \| \| \| \| \|	Add llvm-mca tests demonstrating how LEA instructions are currently modelled. Once this is working on btver2 I'll copy the test file to the other target directories. llvm-svn: 337297
*	[AArch64][SVE]: Integer multiply-add/subtract instructions.	Sander de Smalen	2018-07-17	8	-0/+204
\| \| \| \| \| \| \| \| \| \|	This patch adds support for the following instructions: MLA mul-add, writing addend (Zda = Zda + Zn * Zm) MLS mul-sub, writing addend (Zda = Zda + -Zn * Zm) MAD mul-add, writing multiplicant (Zdn = Za + Zdn * Zm) MSB mul-sub, writing multiplicant (Zdn = Za + -Zdn * Zm) llvm-svn: 337293
*	[Mips][FastISel] Fix handling of icmp with i1 type	Petar Jovanovic	2018-07-17	2	-2/+15
\| \| \| \| \| \| \| \| \| \| \|	The Mips FastISel back-end does not extend i1 values while lowering icmp. Ensure that we bail into DAG ISel when handling this case. Patch by Dragan Mladjenovic. Differential Revision: https://reviews.llvm.org/D49290 llvm-svn: 337288
*	[IPSCCP] Run Solve each time we resolved an undef in a function.	Florian Hahn	2018-07-17	2	-4/+48
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Once we resolved an undef in a function we can run Solve, which could lead to finding a constant return value for the function, which in turn could turn undefs into constants in other functions that call it, before resolving undefs there. Computationally the amount of work we are doing stays the same, just the order we process things is slightly different and potentially there are a few less undefs to resolve. We are still relying on the order of functions in the IR, which means depending on the order, we are able to resolve the optimal undef first or not. For example, if @test1 comes before @testf, we find the constant return value of @testf too late and we cannot use it while solving @test1. This on its own does not lead to more constants removed in the test-suite, probably because currently we have to be very lucky to visit applicable functions in the right order. Maybe we manage to come up with a better way of resolving undefs in more 'profitable' functions first. Reviewers: efriedma, mssimpso, davide Reviewed By: efriedma, davide Differential Revision: https://reviews.llvm.org/D49385 llvm-svn: 337283
*	[AArch64][SVE] Asm: FP fused multiply-add/subtract instructions.	Sander de Smalen	2018-07-17	16	-0/+576
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds support for the following instructions: FMLA mul-add, writing addend (Zda = Zda + Zn * Zm) FNMLA negated mul-add, writing addend (Zda = -Zda + -Zn * Zm) FMLS mul-sub, writing addend (Zda = Zda + -Zn * Zm) FNMLS negated mul-sub, writing addend (Zda = -Zda + Zn * Zm) FMAD mul-add, writing multiplicant (Zdn = Za + Zdn * Zm) FNMAD negated mul-add, writing multiplicant (Zdn = -Za + -Zdn * Zm) FMSB mul-sub, writing multiplicant (Zdn = Za + -Zdn * Zm) FNMSB negated mul-sub, writing multiplicant (Zdn = -Za + Zdn * Zm) llvm-svn: 337282
*	[SLPVectorizer] Don't attempt horizontal reduction on pointer types (PR38191)	Simon Pilgrim	2018-07-17	1	-0/+128
\| \| \| \| \| \|	TTI::getMinMaxReductionCost typically can't handle pointer types - until this is changed its better to limit horizontal reduction to integer/float vector types only. llvm-svn: 337280
*	More fixes for subreg join failure in RegCoalescer	Tim Renouf	2018-07-17	1	-0/+319
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Summary: Part of the adjustCopiesBackFrom method wasn't correctly dealing with SubRange intervals when updating. 2 changes. The first to ensure that bogus SubRange Segments aren't propagated when encountering Segments of the form [1234r, 1234d:0) when preparing to merge value numbers. These can be removed in this case. The second forces a shrinkToUses call if SubRanges end on the copy index (instead of just the parent register). V2: Addressed review comments, plus MIR test instead of ll test Subscribers: MatzeB, qcolombet, nhaehnle Differential Revision: https://reviews.llvm.org/D40308 Change-Id: I1d2b2b4beea802fce11da01edf71feb2064aab05 llvm-svn: 337273
*	[AArch64][SVE] Asm: Support for predicated FP operations (FP immediate)	Sander de Smalen	2018-07-17	10	-0/+380
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch completes support for the following floating point instructions that take FP immediates: FADD* (addition) FSUB (subtract) FSUBR (subtract reverse form) FMUL* (multiplication) FMAX* (maximum) FMAXNM (maximum number) FMIN (maximum) FMINNM (maximum number) All operations are predicated and take a FP immediate operand, e.g. fadd z0.h, p0/m, z0.h, #0.5 fmin z0.s, p0/m, z0.s, #1.0 ^___________^ (tied) * Instructions added in a previous patch. llvm-svn: 337272
*	[NFC][testcases] add testcases for folding srem whose operands are negatived.	Chen Zheng	2018-07-17	1	-0/+49
\| \| \| \| \| \| \|	Finish same optimization for add instruction in D49216 and sdiv instruction in D49382. This patch is for srem instruction. llvm-svn: 337270
*	[llvm-objcopy] Run not with any python, but the python configured in lit.	Benjamin Kramer	2018-07-17	4	-4/+4
\| \| \| \|	llvm-svn: 337262
*	[AArch64][SVE] Asm: Support for predicated FP operations.	Sander de Smalen	2018-07-17	26	-0/+739
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds support for the following floating point instructions: FABD (absolute difference) FADD (addition) FSUB (subtract) FSUBR (subtract reverse form) FDIV (divide) FDIVR (divide reverse form) FMAX (maximum) FMAXNM (maximum number) FMIN (minimum) FMINNM (minimum number) FSCALE (adjust exponent) FMULX (multiply extended) All operations are predicated and binary form, e.g. fadd z0.h, p0/m, z0.h, z1.h ^___________^ (tied) Supporting 16, 32 and 64-bit FP elements. llvm-svn: 337259
*	[DAGCombiner] Call SimplifyDemandedVectorElts from EXTRACT_VECTOR_ELT	Simon Pilgrim	2018-07-17	9	-592/+317
\| \| \| \| \| \| \| \|	If we are only extracting vector elements via EXTRACT_VECTOR_ELT(s) we may be able to use SimplifyDemandedVectorElts to avoid unnecessary vector ops. Differential Revision: https://reviews.llvm.org/D49262 llvm-svn: 337258