using System;
namespace Test
{
static class Program
{
static void Main()
{
Console.WriteLine(Console.InputEncoding);
Console.WriteLine(Console.OutputEncoding);
Console.WriteLine();
Console.Write("Text : ");
Console.WriteLine("Result : {0}", Console.ReadLine());
}
}
}
Example run:
$ dotnet run
System.Text.UTF8Encoding+UTF8EncodingSealed
System.Text.UTF8Encoding+UTF8EncodingSealed
Text : abcæøådef
Result : abc def
The non-ASCII characters are basically being replaced with NULs for some reason.
This happens in all terminals I've tried (CMD, PowerShell, Windows Terminal, mintty). I checked chcp in all conhost-based terminals, and it reported 65001 (UTF-8) everywhere. I've also enabled global UTF-8 in Windows region settings just for good measure (enabling/disabling it appears to make no difference).
What is fascinating here is that this only seems to happen in .NET Core processes. No other programs in any of the terminals I've tried have issues processing non-ASCII characters. For example, things like this work in all of them:
$ echo abcæøådef
abcæøådef
$ cat
abcæøådef
abcæøådef
What is even more fascinating is that if you P/Invoke ReadFile to read from standard input in the .NET Core program instead of using System.Console, you get the same issue: The read is successful but non-ASCII characters are just replaced with NULs.
So the question is: Why are .NET Core processes special? What does .NET Core do that seemingly makes ReadFile misbehave?
$ dotnet --info
.NET SDK (reflecting any global.json):
Version: 5.0.100-rc.1.20452.10
Commit: 473d1b592e
Runtime Environment:
OS Name: Windows
OS Version: 10.0.19042
OS Platform: Windows
RID: win10-x64
Base Path: C:\Program Files\dotnet\sdk\5.0.100-rc.1.20452.10\
Host (useful for support):
Version: 5.0.0-rc.1.20451.14
Commit: 38017c3935
Tagging subscribers to this area: @eiriktsarpalis, @jeffhandley
See info in area-owners.md if you want to be subscribed.
Can also reproduce with Chinese characters, but they are replaced with 0 instead of NUL.
Can also reproduce with .NET Framework 4.7.2
FWIW, I don't see this on my machine. I get this:
C:\Users\stoub\Desktop\tmp>dotnet run
System.Text.OSEncoding
System.Text.OSEncoding
Text : abcæøådef
Result : abcæoådef
with:
.NET SDK (reflecting any global.json):
Version: 5.0.100-rc.2.20480.7
Commit: 53e0c8c7f9
Runtime Environment:
OS Name: Windows
OS Version: 10.0.19042
OS Platform: Windows
RID: win10-x64
Base Path: C:\Program Files\dotnet\sdk\5.0.100-rc.2.20480.7\
I'd guess we have different code pages set globally by default somewhere, as Console uses the win32 GetConsoleCP/GetConsoleOutputCP functions on Windows to determine what encoding to use.
cc: @tarekgh, @krwq, @safern
I checked
chcpin allconhost-based terminals, and it reported 65001 (UTF-8) everywhere.
This should be the key point. Check "use UTF-8 for non-Unicode programs".
I checked
chcpin allconhost-based terminals, and it reported 65001 (UTF-8) everywhere.
This should be the key point. Check "use UTF-8 for non-Unicode programs".

(Opened a new sandbox to get the screenshot in English)

This happens at input side. Hard-coded strings can be outputted correctly.
I had seen this issue in Visual Studio elsewhere, but didn't collect which parts are influenced.
@huoyaoyuan I mentioned in the bug description that I have tried both with that setting off and on; it doesn't seem to make any difference. I think whatever you set with chcp will be what .NET cares about at the end of the day, in any case.
@stephentoub for context, my region settings are:

@alexrp could you please send the content of the registry key OEMCP under the HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage?
Also, after changing the Region Settings option to enable UTF-8, did you reboot the machine after that?
Another thing, what font you are using in the console too.
the content of the registry key
For me, it's 65001
did you reboot the machine after that
I had turned on this for 1 year
what font you are using in the console too
Just Consolas. It should be unrelated.
@alexrp could you please send the content of the registry key
OEMCPunder theHKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage?
65001
Also, after changing the Region Settings option to enable UTF-8, did you reboot the machine after that?
Yes.
Another thing, what font you are using in the console too.
I don't think it matters since the text gets garbled at input, not output, but: CMD and PowerShell use Lucida Console, Windows Terminal and mintty use Consolas.
Seems I can repro this but only when I change code page of console to 65001 but note the output for 437 is also not exactly the same as input (ø => o)
> dotnet run
System.Text.OSEncoding
System.Text.OSEncoding
Text : abcæøådef
Result : abcæoådef
> chcp
Active code page: 437
> chcp 65001
Active code page: 65001
> dotnet run
System.Text.UTF8Encoding+UTF8EncodingSealed
System.Text.UTF8Encoding+UTF8EncodingSealed
Text : abcæøådef
Result : abc def
Output for 437 is expected. Input redirection works as expected too:
>chcp 65001
Active code page: 65001
>echo abcæøådef | dotnet run
System.Text.UTF8Encoding+UTF8EncodingSealed
System.Text.UTF8Encoding+UTF8EncodingSealed
Text : Result : abcæøådef
So only direct read from console is broken.
This might be interesting
Code
using System;
using System.IO;
using System.Linq;
using System.Text;
namespace Test
{
static class Program
{
static void PrintHex(Span<byte> bytes)
{
foreach (byte x in bytes)
{
Console.Write($"{x:X2} ");
}
Console.WriteLine();
}
static void Main()
{
string problematic = @"abcæøådef";
Console.WriteLine(Console.InputEncoding);
Console.WriteLine(Console.OutputEncoding);
Console.WriteLine();
Console.WriteLine($"original: {problematic}");
Console.Write(" Text : ");
Console.WriteLine(" Result : {0}", Console.ReadLine());
Stream stdin = Console.OpenStandardInput();
Console.Write(" Text : ");
byte[] bytes = new byte[100];
int readBytes = stdin.Read(bytes);
Span<byte> input = new Span<byte>(bytes).Slice(0, readBytes);
Console.Write(" input: ");
PrintHex(input);
Console.Write("in. conv: ");
Console.Write(Console.InputEncoding.GetString(input));
Console.Write("original: ");
PrintHex(Console.InputEncoding.GetBytes(problematic));
}
}
}
System.Text.UTF8Encoding+UTF8EncodingSealed
System.Text.UTF8Encoding+UTF8EncodingSealed
original: abcæøådef
Text : abcæøådef
Result : abc def
Text : abcæøådef
input: 61 62 63 00 00 00 64 65 66 0D 0A
in. conv: abc def
original: 61 62 63 C3 A6 C3 B8 C3 A5 64 65 66
I suspect the problem might be us using ReadFile (useFileAPIs under debugger is true): https://github.com/dotnet/runtime/blob/master/src/libraries/System.Console/src/System/ConsolePal.Windows.cs#L1167
I see people report issues with that:
https://stackoverflow.com/questions/48176431/reading-utf-8-characters-from-console
https://github.com/microsoft/terminal/issues/4551#issuecomment-585487802
I think we should always go to the other code path and possibly do some conversion there (or hard-code the input encoding to whatever ReadConsole is using on Windows)
@danmosemsft would this meet the bar for 5.0/servicing? This has impact on every customer using console apps with non ASCII characters with .NET. I know this repros at minimum in 3.1 and likely lower as well.
@krwq I recommend we get the fix into 6.0.0. After that, if we receive enough reports of users being blocked by this bug, we'd consider down-level servicing. We would need validation from users who have encountered this that the behavior is indeed fixed with the 6.0.0 builds.
Another result in RC 2 with CP 932...


Most helpful comment
@krwq I recommend we get the fix into 6.0.0. After that, if we receive enough reports of users being blocked by this bug, we'd consider down-level servicing. We would need validation from users who have encountered this that the behavior is indeed fixed with the 6.0.0 builds.