Runtime: Optimize enum1.HasFlag(enum2) into inline bittest without boxing allocations when types are the same

Created on 9 Jun 2016  路  8Comments  路  Source: dotnet/runtime

Enum.HasFlag implementation requires boxing of both arguments. When the enum types are the same we can replace the call with an inline bittest (value & flag) == flag that avoids the boxing allocations.

There is a discussion on StackOverflow: http://stackoverflow.com/questions/7368652/what-is-it-that-makes-enum-hasflag-so-slow

Mono Mini JIT does a peep in their CIL reader: https://na01.safelinks.protection.outlook.com/?url=https%3a%2f%2ft.co%2fv7I7dOj58T&data=01%7c01%7crajak%40microsoft.com%7c9217e353e39449cee06408d38ed36c54%7c72f988bf86f141af91ab2d7cd011db47%7c1&sdata=m63WkuKl9u061UPSUNqxjpwCg7OwahypwxJnfX71vgA%3d

The IL pattern in RyuJIT importer typically looks like, though the loading of the Enum value onto the stack could be done with any of the various ld* operations.

IL_0018  06                ldloc.0     
IL_0019  8c 02 00 00 02    box          0x2000002
IL_001e  07                ldloc.1     
IL_001f  8c 02 00 00 02    box          0x2000002
IL_0024  28 04 00 00 0a    call         System.Enum.HasFlag

A morph peep on the trees would be straightforward, though it could also be possible through a series of applications of inlining, type-propagation, unreachable code-elimination, and copy-propagation to achieve the same effect with a more generic global approach.

area-CodeGen-coreclr enhancement optimization tenet-performance

Most helpful comment

@mikedn If the JIT were to do the optimization we could probably replace the current implementation with a full IL implementation to facilitate. As is now we need the specialization.

@GSPP The JIT team is happy to accept contributions. Work that is assigned to Future are the best candidates, but it's good practice to check with the team before taking one just to make sure its not on the near-term radar for the team.

In general, use the issue to propose an approach, get some feedback from the team that the design direction you propose is reasonable, and get frequent code reviews. For certain kinds of changes in the more complex areas of the JIT optimizer or code-generator, be prepared for unit tests, lots of testing, assembly diff inspections, presenting benchmark data, etc... Not trying to discourage anyone, but just be prepared for the quality-aspects that goes with touching the underlying execution system.

I agree this one might be an easy one for external contributors to poach.

All 8 comments

though it could also be possible through a series of applications of inlining,...

That would be nice but it won't work with the current implementation as it calls InternalHasFlag and that method is provided by the runtime.

I'm curious in general: Is the JIT team open to such peephole optimizations being contributed by the community? This seems like a great candidate. The feature is not essential and can be cut at any time.

@mikedn If the JIT were to do the optimization we could probably replace the current implementation with a full IL implementation to facilitate. As is now we need the specialization.

@GSPP The JIT team is happy to accept contributions. Work that is assigned to Future are the best candidates, but it's good practice to check with the team before taking one just to make sure its not on the near-term radar for the team.

In general, use the issue to propose an approach, get some feedback from the team that the design direction you propose is reasonable, and get frequent code reviews. For certain kinds of changes in the more complex areas of the JIT optimizer or code-generator, be prepared for unit tests, lots of testing, assembly diff inspections, presenting benchmark data, etc... Not trying to discourage anyone, but just be prepared for the quality-aspects that goes with touching the underlying execution system.

I agree this one might be an easy one for external contributors to poach.

@cmckinsey
In the backend (at least in the legacy) there seems to be a problem that the code is generated tree-by-tree without optimizing inter-tree code. Or one can say, without any post-tree peephole optimizations on the generated code.

Is there anything I'm missing perhaps in the new backend, or are there any general design directions for a framework for peepholes etc.?

Working on this now...

Have the basics working ... preview on this fork.

Plan to generalize it so we can call it again post-inline.

@terrajobst, the issue lost the right milestone and got Future while the issue was fixed in 2017.

We haven鈥檛 been consistent at marking fixed issues with a milestone that matches their release

Was this page helpful?
0 / 5 - 0 ratings

Related issues

omajid picture omajid  路  3Comments

bencz picture bencz  路  3Comments

iCodeWebApps picture iCodeWebApps  路  3Comments

GitAntoinee picture GitAntoinee  路  3Comments

aggieben picture aggieben  路  3Comments