Runtime: Proposal: return random double within a specific range (inclusive)

Created on 15 Jun 2020  Â·  9Comments  Â·  Source: dotnet/runtime

Background and Motivation

The Random class is often used to add noise to data, so it would be nice if there was a method that could return a random double within a specified range. Noise can be both positive or negative, so the method should accommodate returning a value between two arbitrary doubles (both bounds are included). A good candidate for a default range would be [-1,1], because this allows you to scale the noise by a constant or an expression.

In video games, noise can be used to simulate environmental effects. For example, when registering a hit on the cue ball in a game of pool, you could add a tiny amount of noise to the angle in which the cue ball will travel to simulate an uneven table or specks of dust on the table cloth. Similarly, in a shooter game, you could add noise to where the bullets go after being fired from a gun, which is called bullet spread. The noise in these examples is not necessarily always symmetric around the unmodified values - the unmodified cue ball direction and the centre of the cross hairs respectively.

Proposed API

```c#
Random rng = new Random();
// Return random double between -1 and 1 (inclusive)
double noiseDefaultProposed = rng.NextDoubleRange();
// Return a random double between -2 and 4 (inclusive)
double noiseRangeProposed = rng.NextDoubleRange(-2, 4);

## Usage Examples

Adding noise to data becomes easier and more readable with this method. Noise can be positive or negative, but the maximum deviation is generally the same in both directions. Example 1a shows the syntax using `NextDoubleRange()` whereas example 1b shows the traditional way to do it.

Proposed:
```c#
// Example 1a: Proposed
Random rng = new Random();
double mean = 10;     // Example mean value.
double deviation = 5; // Maximum magnitude of the deviation

// Generate a random value between 5 and 15 (10 ± 5)
double noise = deviation * rng.NextDoubleRange();
double noisyValue = mean + noise;

Current:
```c#
// Example 1b: Current
Random rng = new Random();
double mean = 10; // Example mean value
double deviation = 5; // Maximum magnitude of the deviation

// Generate a random value between 5 and 15 (10 ± 5)
double noise1 = deviation * (2 * rng.NextDouble() - 1); // One possibility
double noisyValue1 = mean + noise1;
double noise2 = deviation * rng.Next(-1, 2) * rng.NextDouble(); // Another possibility
double noisyValue2 = mean + noise2;

If you want to add bias to the noise, i.e. you want the deviation to be positive more often than negative (or vice versa), simply change the minimum and maximum parameters to tweak the noise range. Example 2a shows the syntax using `NextDoubleRange(double minValue, double maxValue)`, which is far more concise and more readable than the traditional way shown in example 2b, which requires some algebraic manipulations to get the same result.

Proposed:
```c#
// Example 2a:
Random rng = new Random();
double mean = 10;  // Example mean
double lower = -2; // Lower bound on the noise
double upper = 4;  // Upper bound on the noise

// Generate a random value between 8 and 14 (at least 10 - 2 and at most 10 + 4)
double noise = rng.NextDoubleRange(lower, upper);
double noisyValue = mean + noise;

Current:
```c#
// Example 2b:
Random rng = new Random();
double mean = 10; // Example mean
double lower = -2; // Lower bound on the noise
double upper = 4; // Upper bound on the noise

// Generate a random value between 8 and 14 (10 - 2 or 10 + 4 respectively)
double noiseMean = (lower + upper) / 2.0;
double noiseDeviation = upper - noiseMean;
double noisyValue = mean + noiseMean + noiseDeviation * (2 * rng.NextDouble() - 1);
```
Of course, it should be noted that the asymmetric noise expression could have been written more concisely using a mean of 11 and a symmetric deviation of 3, but this means you actually get something similar to example 1a.

The main takeaway for example 2 is that you still need to write a whole lot of boilerplate code just to get a random double between two arbitrary bounds, which is easy to mess up.

Alternative Designs

The name NextDoubleRange is kind of long, so perhaps it can be called NextDoubles, although that would probably be too confusing since the method only returns a single double.

Risks

I doubt that this change will have any risks related to it.

api-suggestion area-System.Numerics

Most helpful comment

@tannergooding

Unlike integers it's incredibly hard to describe what a "random value between min and max" might look like, because the floating-point range is not evenly distributed.

I think this being a hard problem is an argument for why this functionality should be included in the framework: If the framework doesn't solve it, you're leaving it up to each individual user to solve it and there is a decent chance they're not going to do it well.

As an example, noisyValue2 in the original post has the following histogram and I don't think that was intended:

You might specify for example, 0.0f as the lower bound and 33554434 as the upper bound, but the upper bound is actually 33554432 due to the requested not being exactly representable and the actual being the nearest representable.

I don't see how this is a unique problem for sampling. For example, MathF.Abs(33554434f) returns 33554432, but that doesn't mean Abs is somehow wrong. I'd expect a similar behavior from this method.

Likewise, you might say 268435456‬ as the lower and 268435472 as the upper, expect a range of 16, but you actually have a range of 0 because they are the same number when represented as a float

At the same time, if you write, 268435472f - 268435456f, you might expect the result of 16, but you get 0. But that doesn't mean that - is wrong, or that the framework should not provide -.

All 9 comments

I think it would be more consistent if the added method was:

c# // Returns a random floating-point number that is greater than or equal to minValue, // and less than maxValue. public double NextDouble(double minValue, double maxValue);

Would this work for you?

Unlike integers it's incredibly hard to describe what a "random value between min and max" might look like, because the floating-point range is not evenly distributed.
Roughly half of all floating-point numbers fall between -1 and +1. Roughly the other half exist between +1 and MaxValue and between -1 and MinValue.

Outside of that, they also aren't evenly distributed between these ranges. There is notably a gap between +/-0 and Epsilon. Then, starting from Epsilon the gap between numbers doubles every power of two. There is also +Infinity, -Infinity, and 2^52 representations of NaN (2^23 for float).

The closest you could get to a "true random" double is to take the integral representation of minValue and maxValue and choose a value between these.

This, however, is also likely to give a non well-distributed random input. You might specify for example, 0.0f as the lower bound and 33554434 as the upper bound, but the upper bound is actually 33554432 due to the requested not being exactly representable and the actual being the nearest representable.
NOTE: I use float here as the numbers are smaller and easier to read, but the same issue exists for double, just with different values

Likewise, you might say 268435456‬ as the lower and 268435472 as the upper, expect a range of 16, but you actually have a range of 0 because they are the same number when represented as a float
Or you might choose 268435473 as the upper bound, and the range suddenly becomes 32 because the nearest representable float is now 268435488.

@svick - OP explicitly wants the upper bound inclusive as well.

@svick I would prefer that either both limits are inclusive or exclusive, because I bring up this issue to make the case for symmetric noise.
@tannergooding would it be possible then to define this method as taking a random whole multiple of epsilon (at least 0) above the minimum value but not exceeding the maximum value (included)?

@tannergooding

Unlike integers it's incredibly hard to describe what a "random value between min and max" might look like, because the floating-point range is not evenly distributed.

I think this being a hard problem is an argument for why this functionality should be included in the framework: If the framework doesn't solve it, you're leaving it up to each individual user to solve it and there is a decent chance they're not going to do it well.

As an example, noisyValue2 in the original post has the following histogram and I don't think that was intended:

You might specify for example, 0.0f as the lower bound and 33554434 as the upper bound, but the upper bound is actually 33554432 due to the requested not being exactly representable and the actual being the nearest representable.

I don't see how this is a unique problem for sampling. For example, MathF.Abs(33554434f) returns 33554432, but that doesn't mean Abs is somehow wrong. I'd expect a similar behavior from this method.

Likewise, you might say 268435456‬ as the lower and 268435472 as the upper, expect a range of 16, but you actually have a range of 0 because they are the same number when represented as a float

At the same time, if you write, 268435472f - 268435456f, you might expect the result of 16, but you get 0. But that doesn't mean that - is wrong, or that the framework should not provide -.

Related: https://github.com/dotnet/runtime/issues/32855, where we rejected an API RandomNumberGenerator.GetDouble().

But that use case is different than the one here, as RandomNumberGenerator must be unbiased, and Random is allowed to have bias. So IMO we'd be allowed to fudge a little in this API.

As an example, noisyValue2 in the original post has the following histogram and I don't think that was intended:

While this particular implementation has issues, it is something that is inherent with floating-point:

This SO post shows the distribution just between 0 and 128: https://stackoverflow.com/questions/7006510/density-of-floating-point-number-magnitude-of-the-number
image

The full distribution between MinValue and MaxValue is much more severe (and would take a while to generate).

@svick After seeing your findings about noisyValue2 I noticed I'd forgotten to add in the deviation factor, so I edited my original post. The distribution will most likely look the same but at least the resulting values will be between 5 and 15 rather than between 9 and 11.

Tagging subscribers to this area: @tannergooding
Notify danmosemsft if you want to be subscribed.

Was this page helpful?
0 / 5 - 0 ratings

Related issues

v0l picture v0l  Â·  3Comments

matty-hall picture matty-hall  Â·  3Comments

btecu picture btecu  Â·  3Comments

bencz picture bencz  Â·  3Comments

aggieben picture aggieben  Â·  3Comments