In particular I need Math.Exp support, similar to how there is support for Math.Sqrt via the Vector.SquareRoot function.
It should be trivial to add support for Math.Exp and in fact I've done this with my local copy of System.Numerics.Vectors using SquareRoot as a guiding example.
My application now uses the new Vector.Exp function instead of me having to use the CopyTo function to unpack my Vector
May I check the changes in?
This seems like a useful addition; my initial impression is that we would probably accept something like this, depending on a couple things. Like any other API addition, we need to follow the standard procedure:
https://github.com/dotnet/corefx/wiki/API-Review-Process
I'm curious whether something like Vector.Exp can be meaningfully accelerated using any hardware instructions. There's an assumption that operations in the Vector class are hardware-accelerated, at least where reasonably possible. I'm not familiar with all of the existing hardware instructions, but I don't recall any particularly relevant to calculating exponentials. Then again, I'm not a compiler guy so there may be some clever tricks we can do here to squeeze some speed out.
instead of me having to use the CopyTo function to unpack my Vector to a double[] array, perform Math.Exp on each element of the array and then pack up into a new Vector.
If we provided better support for creating Vector
The implementation of Vector.SquareRoot (for type Double) is essentially:
else if (typeof(T) == typeof(Double))
{
Double* dataPtr = stackalloc Double[Count];
for (int g = 0; g < Count; g++)
{
dataPtr[g] = (Double)Math.Sqrt((Double)(object)value[g]);
}
return new Vector<T>(dataPtr);
As you can see, it simply makes a call to Math.Sqrt for each element of Vector
I've done exactly the same with my implementation of Vector.Exp (ie, forward on to Math.Exp for each element in the Vector
However, when my exe uses this implementation of Vector.Exp, the performance is worse than when it uses the CopyTo function to unpack a Vector to a double[] array, perform Math.Exp on each element of the array and then pack up into a new Vector.
Do you know why this might be? (I'm building a Release 64-bit version of both my exe and System.Numerics.Vectors.dll on Windows 8.1 using VS 2013 )
Yep, that looks like the most sensible way to implement it, I can't really figure a reason why it would perform worse than actually copying everything out to an array and operating on that. I'd have to investigate the scenario more closely.
But, regarding hardware acceleration: The JIT recognizes these special methods and handles them specially at runtime. What you see in the library code is basically the "slow path" IL implementation, which will only get used if JIT intrinsics are disabled or not available. The level of acceleration can vary between operations, for example some are only accelerated for certain data types, some might be emulated using other intrinsics, etc. I'm not sure how Square Root is implemented by the JIT. CopyTo(array) is treated specially, actually, so it may have something to do with the performance you are seeing.
If you're interested, the JIT code is also open, search around here for SIMD-related code:
I'm curious whether something like Vector.Exp can be meaningfully accelerated using any hardware instructions.
Yes and no. No, there's no single SSE/AVX/Neon instruction that can compute exp. Yes, you can still hardware accelerate exp by using existing SIMD instructions to compute exp for 2/4/8 values at the same time. For example, DirectXMath provides XMVectorExpE that does just that.
There are 2 questions that can be asked at this point:
Vector class offer SIMD operations that do not map directly to hardware SIMD instructions? In other words, should the Vector class be more like a general purpose math library or just a SIMD intrinsic library?Vector enough for someone to be able to implement exp and other math functions (cos, sin...) in a separate library?I raise a similar issue in dotnet/runtime#14289.
As for the questions raised above, in my opinion Vector
We absolutely want to enable those scenarios, since there's no other way to get them without JIT support. The related set of math and geometric functionality can be moved into separate types / namespaces. On that front I think exposing the bulk of what DirectXMath does would be extremely beneficial, but that functionality should be built on top of Vector
It seems to me like Vector ought to live in a more core namespace, as it's a fundamental JIT type. The math and geometric functionality can live in System.Numerics.Vectors as a helper library for people who need it. This separation will make it more clear where future functionality ought to live.
Hi, thanks for the quick responses all.
I'm not sure exactly what is meant by "Should the Vector class offer SIMD operations that do not map directly to hardware SIMD instructions". According to https://software.intel.com/sites/default/files/a6/22/18072-347603.pdf exp is an intrinsic.
At any rate, I see that this issue has been assigned to mellinoe, so just wondering what the plan is for adding support for functions like Math.Exp? Currently the use of CopyTo tends to negate any performance gains I get through RyuJIT Vectorization.
According to https://software.intel.com/sites/default/files/a6/22/18072-347603.pdf exp is an intrinsic.
What that doc shows is a "compiler intrinsic" - a C runtime function that is treated specially by C/C++ compilers. For example, during auto-vectorization Visual C++ can replace calls to exp with calls to a custom library function that computes 2/4/8 values at the same time. But there is no SSE/AVX instruction that computes exp.
Ok, I see. But it is still using SIMD instructions for Vectorization, so we would still see a 2/4/8 times speed up by using compiler intrinsics, not so?
In which case, I would think it makes sense to add support for Math.Exp to RyuJIT/System.Numerics.Vectors via these compiler intrinsics.
But it is still using SIMD instructions for Vectorization, so we would still see a 2/4/8 times speed up by using compiler intrinsics, not so?
Yes.
In which case, I would think it makes sense to add support for Math.Exp to RyuJIT/System.Numerics.Vectors via these compiler intrinsics.
Perhaps.
The problem is that what Vector offers now is mostly a 1:1 mapping to SSE/AVX instructions (I don't know ARM Neon so I can't comment on that). Max is maxps/maxpd, SquareRoot is sqrtps/sqrtpd and so on. There are some exceptions (for example DotProduct requires a sequence of SSE instructions if the CPU doesn't support SSE 4.1) but those are special cases generated by Vector's attempt to expose SIMD capabilities in a hardware independent manner.
While it is true that Vector currently offers a pretty straightforward 1:1 mapping to SSE/AVX instructions, that is largely due to the fact that this is a v1 implementation. I would really like to see a greater range of Math functions supported, building upon (and adding to) the existing intrinsics support in RyuJIT.
@ConradMellin, given that you already seem to have the changes ready to go: would you mind starting the API review process by listing the APIs?
@ConradMellin Do you have any further input on this issue? I think the initial idea is good, but we'd need a more concrete proposal for this to be actionable.
Not actionable now, closing.
We can reopen the issue when someone provides more details.
Most helpful comment
While it is true that Vector currently offers a pretty straightforward 1:1 mapping to SSE/AVX instructions, that is largely due to the fact that this is a v1 implementation. I would really like to see a greater range of Math functions supported, building upon (and adding to) the existing intrinsics support in RyuJIT.