Runtime: Add exchange type for a scheduler

Created on 1 Feb 2017  路  14Comments  路  Source: dotnet/runtime

For pipelines, I'd like to be able to pass a Scheduler that determines where the continuation should run (could also be useful for other things). The closet thing we have today is a SynchronizationContext but it's really not appropriate for this. There's also a TaskScheduler but that's about scheduling tasks not arbitrary callbacks.

```C#
public interface IScheduler
{
void Run(Action action);
}

Not sure what other options we'd need here. this is similar to Java's executor https://docs.oracle.com/javase/7/docs/api/java/util/concurrent/Executor.html. 

/cc @stephentoub 

**EDIT: Current proposal**

```C#
public interface IScheduler
{
   void Schedule(Action<object> action, object state);
}

Proposed API implemenations

Scheduler

```C#
public abstract class Scheduler : IScheduler
{
public static IScheduler Default => ThreadPool.Global;
public static IScheduler Inline { get; }

public abstract void Schedule(Action action, object state);
}

The scheduler base class just has the default inline scheduler which runs actions inline. There are concerns about stack diving here but it's required for cases where the consumer wants to stay on the caller's thread.

**ThreadPool**

```C#
public class ThreadPool
{
    public static IScheduler Local { get; }  // QueueUserWorkItem (preferLocal:true)
    public static IScheduler Global { get; } // QueueUserWorkItem (preferLocal:false)
}

Here we're exposing 2 schedulers from the ThreadPool type. These map to QueueUserWorkItem with various arguments.

SynchronizationContext

```C#
public class SynchronizationContext : IScheduler
{
public void Schedule(Action action, object state)
{
Post(new SendOrPostCallback(action), state);
}
}

Exposing the SynchronizationContext as an IScheduler would call Post on the sync context. Since most of the "dispatchers" in .NET already implement a sync context (WPF, System.Web, WinForms, UWP, Xamarin etc) this would be the way to bridge the existing tech with this new interface.

**TaskScheduler**

It should be possible to run Task continuations on any arbitrary IScheduler. We should be able to convert use any IScheduler as a TaskScheduler. This means any thing that uses an TaskScheduler today would just work with the IScheduler via a simple factory.


```C#
public class TaskScheduler
{
    public static TaskScheduler CreateFromScheduler(IScheduler scheduler);
}

Potential APIs Consumers

Once we have the above types exposed we can start looking at places where using a scheduler might be beneficial. Generally, we can look at places that call directly into QueueUserWorkItem and determine if it makes sense to allow callers to specify which IScheduler should be used instead.

More specifically, here are some places it might be useful (we'd of course need to dig into the details more).

SocketAsyncEventArgs

SocketAsyncEventArgs has a callback that executes whenever any IO operation completes (read/write/connect/accept). Currently the thread that callback runs on is either a thread pool thread (in the linux case) or the IOCP completion thread on windows. There are some cases where the calling code wants fine grain control over where continuations run to avoid unnecessary context switches. In these cases it would be ideal if the SocketAsyncEventArgs exposed an IScheduler property that determined where completions ran.

```C#
public class SocketAsyncEventArgs
{
public IScheduler Scheduler { get; set; }
}

This would allow callers to specify where continuations should run (it could even be OS specific).

**Timer**

Timer callbacks are always scheduled on threadpool threads. Adding a scheduler would allow using the timer API but would also allow callers to control where callbacks execute.

```C#
public Timer(TimerCallback callback, object state, int dueTime, int period, IScheduler scheduler);
public Timer(TimerCallback callback, object state, long dueTime, long period, IScheduler scheduler);
public Timer(TimerCallback callback, object state, TimeSpan dueTime, TimeSpan period, IScheduler scheduler);
public Timer(TimerCallback callback, object state, uint dueTime, uint period, IScheduler scheduler);

Task continuations

Running task continuations on specific threads when using async/await syntax today isn't possible without changing the current task scheduler (AFAIK). Ideally you would be able to do something like:

```C#
await task.ContinueOn(scheduler);

Today Task has very specific knowledge about the SynchronizationContext capturing because it's the most common scenario. There are scenarios where the caller whats to continue on another scheduler or TaskScheduler and wants to specify that at the point of the await.

```C#
public class Task
{
      public TaskAwaiter ContinueOn(TaskScheduler scheduler);
}

CancellationToken

CancellationTokens allow for the registration of callbacks that execute of callbacks upon cancellation. A scheduler can be specified to determine where these callbacks should run. This would be more efficient than scheduling the call to Cancel to run on another thread since doing if efficiently requires knowledge that callbacks are registered (see https://github.com/dotnet/corefx/issues/23716 for a related issue). There could be an overload of cancel that takes an IScheduler.

C# public class CancellationTokenSource { public void Cancel(IScheduler scheduler); }

api-needs-work area-System.Threading.Tasks

Most helpful comment

cc @vancem

All 14 comments

As a point of comparison, Rx also has an IScheduler interface.

Differences from the proposed interface are:

  • It includes a notion of time (in the form of the Now property and two Schedule overloads). I always thought that was a weird design (when and where to schedule should be separate concerns) and I think it doesn't make much sense to include that here.
  • It supports passing state to the delegate (and provides convenience Action versions using extension methods). Could be worth considering, depending on how important performance is.
  • The delegate has an IScheduler parameter, which receives the scheduler. Not sure how useful that is.
  • It supports cancellation (using IDisposables instead of CancellationTokens, since it predates the Task Asynchronous Pattern).

Is supports cancellation (using IDisposables instead of CancellationTokens, since it predates the Task Asynchronous Pattern).

Not sure about cancelling scheduled work. I guess some implementations would just ignore it.

It supports passing state to the delegate (and provides convenience Action versions using extension methods). Could be worth considering, depending on how important performance is.

This makes sense. I was going to make that change after seeing if there was interest:

```C#
public interface IScheduler
{
void Schedule(Action action, object state);
}

```C#
public interface IScheduler
{
   void Schedule<TState>(Action<TState> action, TState state);
}

Updated with a more thorough proposal

A few thoughts:

  • Generics. For increased generality, I wonder whether the interface (not the method on the interface) should be generic, e.g. public interface IScheduler<TState> that works with Action<TState> rather than TState effectively being hardcoded to object. This would allow for schedulers that can support TState rather than object to do so, e.g. ThreadPool.GetScheduler<TState>() returning an IScheduler<TState>.
  • Naming. I'm concerned the name is too focused on "scheduling", and that it wouldn't work well for synchronous APIs that just want to invoke something. We could consider IActionInvoker or something like that. This also avoids a conflict with existing interfaces like that used by Rx.
  • Continuations. The proposal suggests adding a ContinueOn method. We've suffered grief over the years for having ConfigureAwait(bool) rather than just having a AvoidCapturingContext() method or something like that, but the main reason ConfigureAwait was named that way was to support overloads with other ways of configuring... it'd be a shame to abandon that now when we could actually benefit from the naming. So if we wanted such an overload, I'd suggest ConfigureAwait(IScheduler), especially since accepting such a scheduler logically replaces the bool.
  • Scheduler. I'm not sure why this class is proposed as abstract. I also don't see a need for it at all. If Scheduler.Default is just ThreadPool.Global, devs can be explicit and just say ThreadPool.Global; we wouldn't change where it referred to, anyway. I'm also concerned about exposing an "inline" scheduler. If you expect that this interface would be used to schedule work in locations where currently code does ThreadPool.QueueUserWorkItem or Task.Run, often in those locations it's being done very explicitly to avoid synchronous invocation, where synchronous invocation would actually cause functional/behavioral problems (e.g. running while holding locks, running in a location that would likely lead to iterated stack dives, running in places where invariants are broken, running on a calling thread thats a critical resource, etc.); we should not make that easy.
  • SocketAsyncEventArgs. Seems like it'd be easy to cause problems here. On Windows, currently the callback just runs on the I/O completion port thread that picks up the I/O completion; any additional scheduling here is a) overhead and b) something that could be done by the callback itself if desired. Further, I'd need to re-look at the implementation, but it's possible we'd need to allocate an object here to pass as the object state to the scheduler. I get the desire here for Linux, though it's banking on an implementation detail that could very likely change to put it into a similar situation as for Windows. Maybe it's a good idea; maybe not. If the Linux implementation did stay the way it is, it's another example where an inline scheduler could be bad. cc: @geoffkizer.
  • Timer. These APIs highlight some of discords we'll have in augmenting existing APIs, namely that it's defined in terms of TimerCallback while this scheduler expects an Action<object>, and that'll make various parts of the implementation more expensive, if nothing else double the delegate invocations. We would want to explore the costs of that vs having the API just differ from the other overloads and take an Action<object>.
  • Exceptions. What are the exceptional semantics of this interface? Are exceptions allowed to propagate out of the Schedule call such that all call sites need to expect / deal with it? Would emerging exceptions fail fast?
  • Interface overhead. This adds an interface invocation in places that wouldn't otherwise have one, e.g. a call site that was using ThreadPool.QueueUserWorkItem now using _scheduler.Scheduler. While such overhead might normally not be an issue, the APIs we're talking about here are potentially used on very hot paths. Before we add an interface like this, someone should prototype and measure to make sure it's not actually a measurable overhead. It may not be, we should just be sure.

Generics. For increased generality, I wonder whether the interface (not the method on the interface) should be generic, e.g. public interface IScheduler\

I don't like the generics idea. It's viral and most of the time it'll be a reference type so we're not saving anything. It also forces all of the places that expose the schedulers to be generic (or expose methods like you suggest for ThreadPool). I don't see the benefit.

Naming. I'm concerned the name is too focused on "scheduling", and that it wouldn't work well for synchronous APIs that just want to invoke something. We could consider IActionInvoker or something like that. This also avoids a conflict with existing interfaces like that used by Rx.

What if it's an IExecutor then? The reason I like scheduling is because that's what we'll use it for. We can iterate on the naming. Proposals so far:

  • IExecutor
  • IActionInvoker (conflicts with ASP.NET MVC 馃槃 )
  • IScheduler

Continuations. The proposal suggests adding a ContinueOn method. We've suffered grief over the years for having ConfigureAwait(bool) rather than just having a AvoidCapturingContext() method or something like that, but the main reason ConfigureAwait was named that way was to support overloads with other ways of configuring... it'd be a shame to abandon that now when we could actually benefit from the naming. So if we wanted such an overload, I'd suggest ConfigureAwait(IScheduler), especially since accepting such a scheduler logically replaces the bool.

Sure. We're already down this rabbit hole I guess.

Scheduler. I'm not sure why this class is proposed as abstract. I also don't see a need for it at all. If Scheduler.Default is just ThreadPool.Global, devs can be explicit and just say ThreadPool.Global; we wouldn't change where it referred to, anyway.

I was mimicking the TaskScheduler APIs here. It's also a nice place to hang other prominent properties or defaults (like the inline scheduler). With these known scheduelrs, maybe it's also possible to perform optimizations (like skipping the interface dispatch if the scheduler is inline). This is something we do all over the place in the BCL for the TPL and the ExecutionContext.

I'm also concerned about exposing an "inline" scheduler. If you expect that this interface would be used to schedule work in locations where currently code does ThreadPool.QueueUserWorkItem or Task.Run, often in those locations it's being done very explicitly to avoid synchronous invocation, where synchronous invocation would actually cause functional/behavioral problems (e.g. running while holding locks, running in a location that would likely lead to iterated stack dives, running in places where invariants are broken, running on a calling thread thats a critical resource, etc.); we should not make that easy.

That's fine as a general concern and if there are places where it would actually break we can refuse to expose it but part of the point of this interface existing is that we want to give the caller more control over where callbacks execute. If that's an advanced scenario, then hide the extensibility behind more advanced options. We can crawl before we run here. This reminds me of the concerns listed here https://github.com/dotnet/corefx/issues/2454 and I generally agree. I just hate the fact that we have 2 hardcoded modes today:

  • TheadPool
  • Current thread (inline)

You should also be able to pass a TaskScheduler to a TaskCompletionSource (I'll add that to the proposal).

SocketAsyncEventArgs. Seems like it'd be easy to cause problems here. On Windows, currently the callback just runs on the I/O completion port thread that picks up the I/O completion; any additional scheduling here is a) overhead and b) something that could be done by the callback itself if desired. Further, I'd need to re-look at the implementation, but it's possible we'd need to allocate an object here to pass as the object state to the scheduler. I get the desire here for Linux, though it's banking on an implementation detail that could very likely change to put it into a similar situation as for Windows. Maybe it's a good idea; maybe not. If the Linux implementation did stay the way it is, it's another example where an inline scheduler could be bad. cc: @geoffkizer.

We should measure the overhead, my gut tells me it won't matter at all (we do this today with Kestrel). This could be a case where the code gets optimized if the Inline scheduler is chosen (skip the interface dispatch and just call inline). It also means we'd be able to avoid double dispatch in some cases.

Timer. These APIs highlight some of discords we'll have in augmenting existing APIs, namely that it's defined in terms of TimerCallback while this scheduler expects an Action\

Sound good to me. It would be ideal to avoid allocations here. I'll add that to the proposal. Maybe the only overloads that take Action\

Exceptions. What are the exceptional semantics of this interface? Are exceptions allowed to propagate out of the Schedule call such that all call sites need to expect / deal with it? Would emerging exceptions fail fast?

We should be specific about what sorts of things would cause exceptions:

  • Failure to schedule the work (a bad custom scheduler?)
  • The inline scheduler executes a callback on the scheduler's thread and user code throws.

The easiest thing would be to do nothing and let the exception propagate. What happens today if the SynchronizationContext.Post implementation throws?

Interface overhead. This adds an interface invocation in places that wouldn't otherwise have one, e.g. a call site that was using ThreadPool.QueueUserWorkItem now using _scheduler.Scheduler. While such overhead might normally not be an issue, the APIs we're talking about here are potentially used on very hot paths. Before we add an interface like this, someone should prototype and measure to make sure it's not actually a measurable overhead. It may not be, we should just be sure.

2 things:

  • We should measure
  • Like I said before, maybe we can optimize known schedulers and avoid the overhead.

It's viral and most of the time it'll be a reference type so we're not saving anything. It also forces all of the places that expose the schedulers to be generic (or expose methods like you suggest for ThreadPool).

No, it doesn't. If the type doesn't want to be generic, it can just use object, in which case it's no worse than if object were hardcoding in the interface:

public class TaskSchduler : IScheduler<object>

But if someone can be generic and wants to be on the state, they have the option.

You don't see the benefit. I see potential benefit, and I don't see the downside.

I was mimicking the TaskScheduler APIs here.

TaskScheduler is abstract because you derive from it to override. Here there's an interface; we don't need both an interface and an abstract class. I also don't see the value of the statics; if something wants to optimize for a Default that just returns ThreadPool.Global, it can just check ThreadPool.Global. I also think that if we're talking about such a general concept and an executor that's not tied to any particular framework, "default" is confusing, and suggests the default of the target API (e.g. "sockets, do whatever your default scheduling is") rather than some specific implementation like the thread pool. In fact, we'd probably need such a concept, so maybe we allow null to mean that.

if there are places where it would actually break we can refuse to expose it

There are.

What happens today if the SynchronizationContext.Post implementation throws?

It's not supposed to. And if it does, it's up to the consumer, e.g. in async methods we re-propagate the exception on the thread pool to crash. But that also means all such call sites generally protect themselves, with a try/finally around the invocation, which often means a bit more overhead.

You should also be able to pass a TaskScheduler to a TaskCompletionSource

Needing to store such a scheduler reference is of course going to increase the size of objects like this.

All in all, I'm just kind of "meh" on this proposal. It's not making me think "oh my goodness we totally need to add this to the platform". But if others think it's important to add and if someone does the due diligence of experimenting with plumbing it through some of these consumers and verifying the impact is acceptable, I'm ok with it, with the details appropriately sorted out.

In thinking about this a bit more, I realize that some of my feedback was a bit contradictory. My concerns with Inline stem from my desire that a consumer would be able to trust that the schedule method would be asynchronous from the call site, in which case it really is more of a "scheduler" than an "invoker". So from a naming perspective, if such an interface were added, I think I'd want to see it codified that implementations should be asynchronous, and it could be named IActionScheduler or something like that.

Could we change IScheduler to an abstract class, hide the implementations (InlineScheduler and TaskRunScheduler) and expose the implementations as properties on the abstract class??

@KrzysztofCwalina I think that would be fine.

@stephentoub That name sounds good as well. I also added another API that I think would fit well for cancellation token callbacks.

@pakrym has changed it to an abstract class. Thanks!

cc @vancem

My 2 cents - if you're thinking about new scheduler API, could you take into consideration exposing something like thread affinity? This would help tremendously when building multithreaded stateful logic, eg:

It would be great to be able to tell scheduler "all executions of this piece of code should be run on the same CPU, I don't want you to do work stealing here". This could be done either as a param to Schedule method or as property in IThreadPoolWorkItem interface.

There are several reasons for that:

  • Actor systems like Akka or Orleans relly heavily on the concept, that single actor is working sequentially. With ability to specify thread/CPU affinity this becomes trivial.
  • Stateful systems like actor frameworks but also any form of stream pipelines (Rx, TPL Dataflow) often execute the same unit of work over and over with some minimal state attached. It's more effective to try to execute that UoW on the same CPU (given that its state is probably already loaded into corresponding L1-3 slot) than to copy the state to new CPU. Taking advantage of affinity/locality is even more important when we take NUMA into account.

Of course there are some considerations to be taken like potentially uneven split of work but IMO "with great power comes great responsibility" approach is a right choice here and we shouldn't block users to take control into their own hands, if they know what they do.

That requirement doesn鈥檛 gel with my idea of an abstraction. It sounds extremely implementation specific. Can you suggest what that API would look like and what the suggested implementations would do?

Unless you want to fully abstract existence of thread from the .NET API, I don't think it's really that implementation specific. We know that thread pool operates on threads. We know that careful CPU allignment can have practical implications on the systems performance and even behavior (I'm not talking only about predictable sequentiality of operations but also full system linearizability useful to eg. build a determininstic test cases for async code even when parallel operations are inbound).

Regarding suggested API, the core concept would be to introduce some int used to identify where (and if) a work item is supposed to be executed on the specific core/thread. If could be composed into some more general data structure:

public readonly struct ThreadPoolQueueOptions
{
    /// <summary>
    /// When non-specified (0), gives a thread pool a full freedom to choose the operating thread.
    /// When specified, AffinityId is used to determine, which thread to execute work item on.
    /// The most trivial implementation could be to simply determine a thread to execute by
    /// using <code>var thread = threads[option.AffinityId % threads.Length]</code>.
    /// </summary>
    public int AffinityId { get; }
}

class ThreadPool
{
    public static bool UnsafeQueueUserWorkItem (IThreadPoolWorkItem item, in ThreadPoolQueueOptions options)
}

This could be propagated up (under different name) to an IScheduler API.

Was this page helpful?
0 / 5 - 0 ratings