Hi,
This is a general question about RyuJIT and the current careful maintained balance between throughput and generated code efficiency.
In the case of CoreRT, when using RyuJIT as a backend, where the JIT throughput is bit less important (still but...) it makes sense to unleash the compiler and let it optimize as much as it can...
I'm not sure that there is today an AOT mode for RyuJIT, but suppose that we could introduce such mode, I'm wondering what are the top optimizations/improvements that could be part of this mode?
category:proposal
theme:optimization
skill-level:expert
cost:extra-large
I'm not sure that there is today an AOT mode for RyuJIT
Not really. While the JIT has been used for AOT via NGEN for a long time it doesn't do anything special in this regard.
but suppose that we could introduce such mode, I'm wondering what are the top optimizations/improvements that could be part of this mode?
Well, what optimizations do you聽hope to see? There are zillions of things that a compiler could do but not everything is truly useful.
Anyway, one obvious thing that the JIT could do is to increase聽various limits it has (the number of tracked local variables, the number of tracked CSEs etc.).聽It may also be possible to run certain optimizations聽multiple times, sometimes聽this results in better code.
Sure, you could also do more fancy things like inteprocedural optimization but the current JIT has no mechanism for that and adding it would likely require a聽significant amount of聽work.
Well, what optimizations do you hope to see? There are zillions of things that a compiler could do but not everything is truly useful.
I don't know enough of all the internals of compilers or RyuJIT to expect anything ultra specific... that's why I'm asking experts here 馃槄
In the past I did run a micro-benchmark between RyuJIT and .NET Native but figures there are today likely irrelevant (for both...)
Anyway, one obvious thing that the JIT could do is to increase various limits it has (the number of tracked local variables, the number of tracked CSEs etc.). It may also be possible to run certain optimizations multiple times, sometimes this results in better code.
Typically this kind of improvements if they are just a switch of some constants would be great!
When looking at the sustained optimizations effort done on RyuJIT over the recent months, it looks like things have been improved on many fronts, that's impressive.
But more generally, regardless of the throughput restriction, what are the existing&remaining areas where RyuJIT is lagging behind a C++ compiler? (let's assume that we don't even talk about vectorized code... but regular things like register allocator, stack spilling...)
How much of the constraints imposed by EH and managed references are breaking optimizations in the code? Is there anything improvable there?
But more generally, regardless of the throughput restriction, what are the existing&remaining areas where RyuJIT is lagging behind a C++ compiler?
There are聽a lot of optimizations that are missing. Just a few examples:
x / 3 > 1 is x > 5 but the JIT doesn't know this.Some of these optimization may affect the compiler throughput but I think there still room for many of them in the existing JIT, with or without AOT. Some聽are missing聽simply because nobody had time to implement them or because nobody considered them useful (yet). Others are missing聽because聽they're perhaps more difficult to implement in the existing JIT design but adjusting the design doesn't necessarily imply that JIT throughput will聽be negatively affected.
but regular things like register allocator
I'm not exactly familiar with register allocators but graph coloring register allocators (typically used by C/C++ compilers) are slightly better than the linear scan allocator used by the JIT.聽AFAIK not a lot better聽so it may be worthless to create a new register allocator, it's聽anything but trivial. That said, the existing allocator could use some additional tweaks and tuning, it works pretty well but sometimes it does weird things.
How much of the constraints imposed by EH and managed references are breaking optimizations in the code? Is there anything improvable there?
EH can be problematic but C++ has exceptions too. One聽way or another exception handling is better avoided if you want fast code. There are some improvements that can be made and one such聽improvement is currently being worked on - inlining/eliminating finally blocks.
Probably the biggest problem related to managed references are聽GC write barriers. Sometimes they aren't needed but聽eliminating them requires interprocedural analysis and full AOT.
There's another constraint that hasn't been mentioned - the CLR memory model. The fact that all stores have release semantics is聽IMO pretty ugly as it can prevent optimizations. And it's problematic to change this since it can break existing code.
I don't know enough of all the internals of compilers or RyuJIT to expect anything ultra specific... that's why I'm asking experts here 馃槄
Well, I'm not a compiler expert but hopefully the information above is of some use聽:smile:
cc @russellhadley @cmckinsey @RussKeldorph @dotnet/jit-contrib
I'm also interested in exploiting RyuJIT in this way especially for devices with small memory and weak computation power, e.g. ARM devices.
I'm not exactly familiar with register allocators but graph coloring register allocators (typically used by C/C++ compilers) are slightly better than the linear scan allocator used by the JIT.
It seems from LLVM 3.0 they departed from linear scan to a global live range splitting approach (blog post and pdf presentation) with sometimes 10%+ performance benefits.
But incidentally, it seems that Microsoft Research has been working on a Hierachical Graph Coloring Register Allocation in LLVM
The tension between JIT and AOT is constantly on the minds of the JIT developers, and we continue to investigate new optimizations when possible. This discussion is not useful to keep open, so closing.
Most helpful comment
There are聽a lot of optimizations that are missing. Just a few examples:
x / 3 > 1isx > 5but the JIT doesn't know this.Some of these optimization may affect the compiler throughput but I think there still room for many of them in the existing JIT, with or without AOT. Some聽are missing聽simply because nobody had time to implement them or because nobody considered them useful (yet). Others are missing聽because聽they're perhaps more difficult to implement in the existing JIT design but adjusting the design doesn't necessarily imply that JIT throughput will聽be negatively affected.
I'm not exactly familiar with register allocators but graph coloring register allocators (typically used by C/C++ compilers) are slightly better than the linear scan allocator used by the JIT.聽AFAIK not a lot better聽so it may be worthless to create a new register allocator, it's聽anything but trivial. That said, the existing allocator could use some additional tweaks and tuning, it works pretty well but sometimes it does weird things.
EH can be problematic but C++ has exceptions too. One聽way or another exception handling is better avoided if you want fast code. There are some improvements that can be made and one such聽improvement is currently being worked on - inlining/eliminating
finallyblocks.Probably the biggest problem related to managed references are聽GC write barriers. Sometimes they aren't needed but聽eliminating them requires interprocedural analysis and full AOT.
There's another constraint that hasn't been mentioned - the CLR memory model. The fact that all stores have release semantics is聽IMO pretty ugly as it can prevent optimizations. And it's problematic to change this since it can break existing code.
Well, I'm not a compiler expert but hopefully the information above is of some use聽:smile: