@jashook Please assign this to me. I will start looking at this tomorrow.
cc/ @dotnet/arm64-contrib @dotnet/jit-contrib
Nice!
Shouldn't we be after System.Runtime.Intrinsics.Neon instead?
Shouldn't we be after System.Runtime.Intrinsics.Neon instead?
It is not clear to me.
Discussions with MS indicated that SIMD and intrinsics would both continue to be developed
Yes, that's the case. And we have existing SIMD tests and usage.
It does seem compelling to implement intrinsics first as it provides a nice way to test the emitters as they are written.
To me, this would simply argue that there is value in doing some development in parallel, such that the SIMD & "bare" intrinsics for the instructions they utilize could be implemented together.
I am making some progress.
Some things are relatively easy : Changes to import, simd.cpp, lowering, codegen, and emitters.
The parts I will find more challenging will be Local Variables, Stack Frame Allocation, Calling conventions. This code is less familiar and therefore seems more complex. There are also lots of special cases handled by #if which makes it difficult to understand.
Compiler::fgMorphArgs(GenTreeCall* call) is particularly nasty and looks like it is overdue for refactoring.
Looks like the ARM/ARM64 HFA code needs to be made aware of SIMD.
Great to hear it! The complexity of fgMorphArgs is all too well known. It has long been a goal to refactor it.
I have uploaded most of my work in progress. The code is in relatively good shape, but it is getting unmanageable to keep it all from being checked in. While all tests are not passing, I believe some of the code can be checked in because most of it is disabled by FEATURE_SIMD (directly or indirectly via varTypeIsSIMD).
I would appreciate a review.
Of the 116 SIMD tests, most are still failing.
The vast majority of the failing tests are failing with the first assert below. This and the next several assets seem to indicate I am doing something wrong with register allocation or codegen genConsume.
Since I do not understand the reason for the first assert it is difficult to debug.
Assert failure(PID 41903 [0x0000a3af], Thread: 41903 [0xa3af]): Assertion failed '((consume == 0) && (produce == 0)) || (ComputeAvailableSrcCount(tree) == consume)' in 'VectorSubTest`1[Single][System.Single]:VectorSub(float,float,float):int' (IL size 58)
File: /home/vmjenkins/workspace/Dotnet/build_and_test/src/jit/lsra.cpp Line: 3730
Image: /home/vmjenkins/workspace/Dotnet/build_and_test/bin/tests/pr1/Tests/coreoverlay/corerun
Assert failure(PID 41289 [0x0000a149], Thread: 41289 [0xa149]): Assertion failed '(candidates & allRegs(regType)) != RBM_NONE' in 'VectorExpTest`1[Int32][System.Int32]:VectorExp(struct,int,int,int):int' (IL size 148)
File: /home/vmjenkins/workspace/Dotnet/build_and_test/src/jit/lsra.cpp Line: 5339
Image: /home/vmjenkins/workspace/Dotnet/build_and_test/bin/tests/pr1/Tests/coreoverlay/corerun
Assert failure(PID 40290 [0x00009d62], Thread: 40290 [0x9d62]): Assertion failed 'regArgMaskLiveSave != regArgMaskLive' in 'System.Collections.Generic.List`1[Vector4][System.Numerics.Vector4]:Add(struct):this' (IL size 71)
File: /home/vmjenkins/workspace/Dotnet/build_and_test/src/jit/codegencommon.cpp Line: 5275
Image: /home/vmjenkins/workspace/Dotnet/build_and_test/bin/tests/pr1/Tests/coreoverlay/corerun
Assert failure(PID 39873 [0x00009bc1], Thread: 39873 [0x9bc1]): Assertion failed '(consume > 1) || (regType(store->gtOp1->TypeGet()) == regType(store->TypeGet()))' in 'ClassLibrary.test:convex_hull(ref)' (IL size 476)
File: /home/vmjenkins/workspace/Dotnet/build_and_test/src/jit/lsra.cpp Line: 3833
Image: /home/vmjenkins/workspace/Dotnet/build_and_test/bin/tests/pr1/Tests/coreoverlay/corerun
If anyone is interested, I have uploaded my working tree here
Progress continues. On my local tree ~50% of SIMD tests are passing.
There are still issues with ABI, lsra, liveness...
For the ABI, callers are passing the SIMD arguments correctly (same as underlying struct type). However managed callees are consuming the arguments incorrectly. I have been trying to find the relevant code, but so far it has eluded me. @dotnet/jit-contrib I am sure I will find the code eventually, but guidance would be appreciated.
Maybe lvaAssignVirtualFrameOffsetToArg?
I think I just found it. Morph was marking the arg as implicit byref, because Compiler::lvaIsMultiregStruct was not using varTypeIsStruct, so was failing on SIMD args.
My tip is approaching 90% pass rate on SIMD tests. All of my current changes are uploaded.
I currently have >90% of tests passing.
I have spent the last two days trying to fix asserts in fgMorphMultiregStructArg() on the following pattern which is causing 75% of the remaining failures.
[000013] --CXG------- | /--* CALL simd16 System.Numerics.Vector`1[Single][System.Single].op_Multiply,NA
[000016] ---XG------- arg0 | | +--* OBJ(16) simd16
[000015] ------------ | | | \--* ADDR byref
[000011] ------------ | | | \--* SIMD simd16 float init
[000010] ------------ | | | \--* CNS_DBL float 1.0000000000000000
[000012] ------------ arg1 | | \--* LCL_VAR float V01 arg1
[000020] -ACXG---R--- \--* ASG simd16 (copy)
[000018] D------N---- \--* LCL_VAR simd16 V03 tmp0
I am not sure if I need to
SIMD simd16 float init into a LCL_VAR (or how to do it)fgMorphMultiregStructArg() or how to do it.I will be offline next week. If there any ideas, I would appreciate advice
Initial implementation is complete.
The current open ARM64 SIMD pull requests allow all tests to pass in debug build (with and without SIMD enabled).
I have run JitStress1, JitStress2, JITMinOpts. All discovered SIMD issues have PRs.
GC stress 0xF also looks good.
I have uploaded the final optimization (getItem containment) planned for this initial implementation.
I expect to transition to working on intrinsics. Let me know if there are any other issues I missed.
Most helpful comment
Initial implementation is complete.
The current open ARM64 SIMD pull requests allow all tests to pass in debug build (with and without SIMD enabled).