Runtime: Add `GC.IsPinnedHeapObject(object obj)` API

Created on 13 Mar 2020  路  43Comments  路  Source: dotnet/runtime

GC.IsPinnedHeapObject(object obj)

Returns true if the obj is allocated on pinned object heap.

Caveat: we generally cannot know cheaply/precisely if object is pinned (thus this is not a IsPinnedObject), but we can tell if it resides in Pinned Object Heap.

===

There are obvious scenarios where this could be useful.
When trying to use pinned object heap in prototypes, it does not take long before you want to write code like:

if (GC.IsPinnedHeapObject(obj))
{
    .. something simple that relies on obj not relocating ..
}
else
{
   .. pinning by hand for compat or throw something, if this is unexpected...
}

or do something like:

Debug.Assert(GC.IsPinnedHeapObject(buffer), "buffer must be on pinned heap");
api-suggestion area-GC-coreclr

All 43 comments

I couldn't add an area label to this Issue.

Checkout this page to find out which area owner to ping, or please add exactly one area label to help train me in the future.

CC: @Maoni0 @terrajobst

How fast would it be to query this API? I'd assume it's something like return (objPtr >= POH_START) && (objPtr <= POH_END); with some minor overhead?

Also, as a side question: is the behavior of GC.GetGeneration(some_obj_from_poh) stable?

The simplest implementation is

// C++
segment = segment_of(obj); 
return segment != NULL && segment->flags & heap_segment_flags_poh;

It is more involved than a range check. There are multiple heaps and multiple segments within, not necessarily contiguous.

segment_of is a table lookup + some checks.
If you do IO after that, then it is very fast. Otherwise, it depends what you do.

You can cache the result too. It will not change for a given object.
There might be other implementations.

So at least for arbitrary inputs, it will still be cheaper to just pin. But for certain specialized cases, it may be useful to check and take it into account?

Right. For fast operations, it would make more sense to just pin on the stack /w fixed statement.

If operation is long(ish) or you need something that is pinned across calls or async operations, then permanently pinned buffer becomes more interesting.

And then, at some point you may find yourself in a need to check or assert that.

Right. For fast operations, it would make more sense to just pin on the stack /w fixed statement.

And do you expect to be materially cheaper than just always allocating the handle for async operation?

I can see this to be useful for asserts or error handling, but not much else.

Seems like there might be a scenario for a managed object to be allocated on the POH and for managed thread + native thread to have a reference to the backing memory simultaneously? Exposing this API if even for error detection in such a scenario could be useful.

Allocating a handle may results in extra code to ensure freeing that handle.

I think there are scenarios where knowing if the object is in POH heap could be useful.
I'd say it could be more useful/actionable API than GC.GetGeneration, for example.

I think there are scenarios where knowing if the object is in POH heap could be useful.

what are those scenarios aside from what Jan pointed out (asserts/err handling)?

I can think of a few cases, hypothetical of course.

  • Imagine an array pool that serves pinned buffers. It probably would not want to accept a regular array returned back. Pool may ignore ordinary arrays or throw something. Then, it is error checking as Jan mentioned.
  • Imagine you get an array/list/dictionary of buffers to share with native code. Then you need to pin each of the buffers. In some cases you may use Overlapped/AsyncPinnedHandle for the sideeffects of pinning multiple things at once.
    If all buffers are in POH, you do not need to worry. Just need to be sure the outer collection is rooted as long as buffers are in use.

Again, you may choose to not support non-POH case at all and throw. Then it becomes error checking.

Maybe add an API that exposes all GC related properties of an object in one call? The pinnedness of the heap, generation, alignment if supported. It could return a struct which would provide extensibility for additional properties to be added in the future.

Exposing a slew of GC-related properties at once assumes that they're cheap to query. That might be the case today, but I don't know if that would universally be the case moving forward for all information we might want to query about the GC or managed objects.

Good point. I just thought, let's return everything that becomes available after the expensive call to segment_of(obj).

my understanding of exposing a public API is it would need to have proven usage cases. my opinion is when we actually use the POH we'll see if there is a proven usage case and the API folks can chime in and see what they think.

Can we add it as internal for now and see if we can take advantage of it from internal code paths?

@Maoni0 @GrabYourPitchforks - Yes. It is very likely that we will need this API. As soon as we try an array pool that can rent pinned buffers, we will need an internal version.

Once we have internal use, it will be easier to explain why a public version might be ok as well.

As soon as we try an array pool that can rent pinned buffers, we will need an internal version.

The array pool cannot afford to check this for every return. It is too expensive to check.

We would need this only if we have some sort of slow diagnostic version of array pool that we do not have today.

BTW: Speaking about array pool: Would it make sense to just change the default array pool to return pinned arrays when the T does not contain references? It would make things simpler - for both the implementation and use.

I am thinking of the path where user returns the array back to the pool. Will we need to check?

Also I think the check is not very expensive and probably can be made cheaper, but will have to measure. It is all relative.

Arguably mischievous user could do a lot of bad things to the pool anyways.

Arguably mischievous user could do a lot of bad things to the pool anyways.

This lines up with earlier statements from this team about ArrayPool being an unsafe malloc / free equivalent. A "checked" pool implementation (off by default) could ensure things like arrays not being double-returned, but the default implementation would need to assume that the caller is well-behaved.

BTW: Speaking about array pool: Would it make sense to just change the default array pool to return pinned arrays when the T does not contain references? It would make things simpler - for both the implementation and use.

Yes.
We may even allow references, but perf impact on marking needs to be measured.

allow references

My gut-feel is that it is not worth it.

is array pool actually renting anything it didn't allocate out? if it allocates the buffers anyway, why not just allocate on the pinned heap to begin with (if desired)? would it have a mix of buffers from POH and the rest of the heap?

why not just allocate on the pinned heap to begin with (if desired)?

That is what Jan is suggesting.

would it have a mix of buffers from POH and the rest of the heap?

no, unless a confused/buggy user code returns an array that was not obtained from the pool.

Right now this is harmless. We only require that buffer is not used after returning and is not returned twice.

It is not harmless. The array pool is not going to like you if you give it an array that does not fit its bucket sizes.

indeed, ArgumentException_BufferNotFromPool

why not just allocate on the pinned heap to begin with (if desired)?

That is what Jan is suggesting.

I was asking if it's even possible today for the array pool to accept a buffer it did not allocate itself and you just answered.

I have not come across a pool implementation in any of our general framework assemblies that would accept a buffer it doesn't allocate itself.

Technically it would accept, but chances are low, since the length is required to be a power of 2.

It is a small change to enable POH in pool - https://github.com/dotnet/runtime/pull/33738

Technically it would accept, but chances are low, since the length is required to be a power of 2.

A lot of arrays are allocated with power-of-2 sizes. But as was mentioned, there are other, easier, more disastrous ways to mess up the pool, like double-returning the same array to it. Even so, I'm not sure how comfortable I'd be with consumers of ArrayPool<T>.Shared.Rent being able to rely on the returned array being permanently pinned. There's almost certainly code somewhere that's putting arrays not from the pool back into the pool, which if they're a power-of-2 actually works today, and they may even be doing it on purpose thinking they're doing a good thing ("I have this array I'm about to drop, might as well allow someone else to use it").

Re: the array pool, what would be the benefit for ArrayPool<T>.Shared.Rent returning pre-pinned buffers? Let's assume for now this is an implementation detail rather than a documented behavior people can rely on.

Is the main benefit that it could avoid managed heap fragmentation in case the array is passed through p/invoke or some other long-running pinning API?

Related to this issue: Memory<T> can be aware it is referring to a pre-pinned object. There is no API to retrieve this information.

The Memory<T> behavior is a little different than the pinned object heap feature. Wrapping a Memory<T> instance around a pre-pinned buffer doesn't attempt any sort of validation. It assumes the caller is operating inside a fixed block, or that a fixed GCHandle has been allocated, or that there's some other mechanism (like the POH) keeping the array pinned. Similarly, wrapping a Memory<T> around an arbitrary T[] shouldn't check whether the array is on the POH, as that check isn't free and the overwhelming majority of Memory<T> instances will never care since nobody calls Pin() on them.

Quite literally the one and only reason for the Memory<T> "pre-pinned" feature is to avoid creating the GCHandle when you call Memory<T>.Pin().

I looked at Memory<T>.Pin() case. It might seem attractive to avoid allocating a handle when not needed, but it is not trivial.

MemoryHandle would need to have a reference back to the object to guarantee the lifetime. The reference could be stored in _pinnable field if that changes its type to object.

Then another issue is that MemoryHandle would need a new constructor - to take that pre-pinned object.

After allocating from POH, MemoryMarshal.CreateFromPinnedArray can be used to store in Memory that this object is pre-pinned. And then a later check on the Memory can be performed in use-cases similar to that of GC.IsPinnedHeapObject.

I would like to put one 'upvote' for this feature, as I see beneficial to create a dedicated array pool, instead of trying to modify existing ones. So my dedicated PinnedArrayPool<T> would really could live with this additional overhead of the check during Return to guard its correctness.

A use-case for this API may be to check if SocketAsyncEventArgs.BufferList arrays are pinned, which avoids having to allocate a list of GCHandle to pin them.

Use-cases on Memory/MemoryHandle are equally interesting:

MemoryHandle handle = memory.Pin();
if (handle.IsPinning) // or !memory.Is(Pre)Pinned
  _handlesToDispose.Add(handle);
SomeFunction(handle.Pointer, memory.Length);

Use-cases on Memory/MemoryHandle are equally interesting

Only if the API is super cheap, which it sounds like it's not. memory.Pin already avoids pinning if the memory was created from MemoryMarshal.CreateFromPinnedArray; I expect it'd be better in such situations to try to use that when creating memory from already pinned arrays.

EDIT: Ah, I may have misunderstood your comment. Were you saying that it'd be interesting to add an IsPinning to MemoryHandle?

Were you saying that it'd be interesting to add an IsPinning to MemoryHandle?

Yes: if (handle.IsPinning) // or !memory.Is(Pre)Pinned

MemoryMarshal.IsPrePinnedArray<T>(ROM<T>) : bool or MemoryMarshal.TryGetPrePinnedPointer<T>(ROM<T>, out T*) : bool could be added if callers need this information. Though the existing Pin method should abstract this away most of the time.

Edit: that last overload could also theoretically work if we ever add the ability to instantiate a Mem<T> directly over a T*.

Was this page helpful?
0 / 5 - 0 ratings

Related issues

matty-hall picture matty-hall  路  3Comments

iCodeWebApps picture iCodeWebApps  路  3Comments

jamesqo picture jamesqo  路  3Comments

Timovzl picture Timovzl  路  3Comments

btecu picture btecu  路  3Comments