Consider the following text with 2 lines (one with content, one blank):
asdf
Load it into a StringReader and do the following:
var reader = new System.IO.StringReader(text);
reader.ReadLine(); // gives 'asdf'
reader.ReadLine(); // gives null
I would expect that the second ReadLine would give an empty string because there is still an additional line in the text. null would mean that there is no additional lines in this text...
Is this behavior by-design? Seems a bit strange if it is.
Does an intermediary line that is blank get read properly?
Yes it does.
Example:
asdf
asdf
Gives you:
"asdf"
""
"asdf"
null
and one with two newlines at the end behaves strangely:
asdf
gets you:
"asdf"
""
null
Here's a fuller description of the various behaviours:
using System;
using System.IO;
using System.Collections.Generic;
using System.Text.RegularExpressions;
using System.Text;
namespace readline
{
class Program
{
private static readonly string[] s_testStrings = new []
{
"a\nb\nc",
"a\nb\nc\n",
"a\r\nb\r\nc",
"a\r\nb\r\nc\r\n",
"a\n\nb\nc",
"a\n\nb\nc\n",
"a\nb\nc\n\n",
"a\n\nb\n\nc\n\n"
};
private static readonly IReadOnlyDictionary<string, Func<string, string[]>> s_splitFuncs = new Dictionary<string, Func<string, string[]>>
{
{ nameof(SplitWithStringReader), SplitWithStringReader },
{ nameof(SplitWithStringSplit), ((s) => SplitWithStringSplit(s)) },
{ $"{nameof(SplitWithStringSplit)}_RemoveEmpty", ((s) => SplitWithStringSplit(s, removeEmptyEntries: true)) },
{ nameof(SplitWithRegex), SplitWithRegex },
{ nameof(SplitManually), SplitManually }
};
static void Main(string[] args)
{
foreach (string str in s_testStrings)
{
string escapedStr = CSharpEscapeString(str);
Console.WriteLine($"==== {escapedStr} ====");
foreach (KeyValuePair<string, Func<string, string[]>> splitPair in s_splitFuncs)
{
string[] result = splitPair.Value(str);
string joinedResult = String.Join('/', result);
int lineCount = result.Length;
Console.WriteLine($"\t{splitPair.Key}:");
Console.WriteLine($"\t\t{joinedResult}");
Console.WriteLine($"\t\tCount: {lineCount}");
}
Console.WriteLine();
}
}
private static string[] SplitWithStringReader(string s)
{
var acc = new List<string>();
using (var sr = new StringReader(s))
{
string curr;
while ((curr = sr.ReadLine()) != null)
{
acc.Add(curr);
}
}
return acc.ToArray();
}
private static string[] SplitWithStringSplit(string s, bool removeEmptyEntries = false)
{
return s.Split(new [] { "\n", "\r\n" }, removeEmptyEntries ? StringSplitOptions.RemoveEmptyEntries : StringSplitOptions.None);
}
private static string[] SplitWithRegex(string s)
{
var notNewline = new Regex(@"\r?\n");
return notNewline.Split(s);
}
private static string[] SplitManually(string s)
{
var acc = new List<string>();
var curr = new StringBuilder();
bool sawCR = false;
for (int i = 0; i < s.Length; i++)
{
if (s[i] == '\r')
{
sawCR = true;
continue;
}
if (s[i] == '\n')
{
acc.Add(curr.ToString());
curr.Clear();
sawCR = false;
continue;
}
if (sawCR)
{
curr.Append('\r');
sawCR = false;
}
curr.Append(s[i]);
}
acc.Add(curr.ToString());
return acc.ToArray();
}
private static string CSharpEscapeString(string s)
{
var sb = new StringBuilder().Append('"');
foreach (char c in s)
{
sb.Append(CSharpEscapeChar(c));
}
sb.Append('"');
return sb.ToString();
}
private static string CSharpEscapeChar(char c)
{
switch (c)
{
case '\a':
return "\\a";
case '\b':
return "\\b";
case '\f':
return "\\f";
case '\n':
return "\\n";
case '\r':
return "\\r";
case '\t':
return "\\t";
case '\v':
return "\\v";
case '\\':
return "\\\\";
case '\0':
return "\\\0";
case '"':
return "\"";
default:
return c.ToString();
}
}
}
}
This program will print the following:
==== "a\nb\nc" ====
SplitWithStringReader:
a/b/c
Count: 3
SplitWithStringSplit:
a/b/c
Count: 3
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a/b/c
Count: 3
SplitManually:
a/b/c
Count: 3
==== "a\nb\nc\n" ====
SplitWithStringReader:
a/b/c
Count: 3
SplitWithStringSplit:
a/b/c/
Count: 4
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a/b/c/
Count: 4
SplitManually:
a/b/c/
Count: 4
==== "a\r\nb\r\nc" ====
SplitWithStringReader:
a/b/c
Count: 3
SplitWithStringSplit:
a/b/c
Count: 3
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a/b/c
Count: 3
SplitManually:
a/b/c
Count: 3
==== "a\r\nb\r\nc\r\n" ====
SplitWithStringReader:
a/b/c
Count: 3
SplitWithStringSplit:
a/b/c/
Count: 4
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a/b/c/
Count: 4
SplitManually:
a/b/c/
Count: 4
==== "a\n\nb\nc" ====
SplitWithStringReader:
a//b/c
Count: 4
SplitWithStringSplit:
a//b/c
Count: 4
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a//b/c
Count: 4
SplitManually:
a//b/c
Count: 4
==== "a\n\nb\nc\n" ====
SplitWithStringReader:
a//b/c
Count: 4
SplitWithStringSplit:
a//b/c/
Count: 5
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a//b/c/
Count: 5
SplitManually:
a//b/c/
Count: 5
==== "a\nb\nc\n\n" ====
SplitWithStringReader:
a/b/c/
Count: 4
SplitWithStringSplit:
a/b/c//
Count: 5
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a/b/c//
Count: 5
SplitManually:
a/b/c//
Count: 5
==== "a\n\nb\n\nc\n\n" ====
SplitWithStringReader:
a//b//c/
Count: 6
SplitWithStringSplit:
a//b//c//
Count: 7
SplitWithStringSplit_RemoveEmpty:
a/b/c
Count: 3
SplitWithRegex:
a//b//c//
Count: 7
SplitManually:
a//b//c//
Count: 7
Seems like the newline is considered part of the current line, it does not mean "there is another line behind".
Does it behave the same way on .NET Framework and .NET Core?
It is likely not something we can change due to compat ...
Confirmed that it's the same in Framework, via Windows PowerShell:
Add-type -TypeDefinition @"
using System;
public class Main2
{
public static void Run()
{
string x = "a\nb\nc\n";
var sr = new System.IO.StringReader(x);
var acc = new System.Collections.Generic.List<string>();
int count = 0;
string curr = null;
while ((curr = sr.ReadLine()) != null)
{
count++;
acc.Add(curr);
}
Console.WriteLine("Count: " + count.ToString());
Console.WriteLine("String: " + String.Join("/", acc.ToArray()));
}
}
"@
[Main2]::Run()
Count: 3
String: a/b/c
Echoing what @karelz said - it looks like the logic here is that a line is terminated by either a newline, or the end of the file.
With the current parsing users don't have to special case the last line. For example:
"line1\n" +
"line2\n" +
"line3\n"
I think that the intuitive thing to do is to parse that as three lines.
On the other hand, if you want an empty line at the end of the file, you can make that behavior explicit with an additional newline:
"line1\n" +
"line2\n" +
"line3\n" +
"\n"
It isn't perfect for all situations, but I think it's a reasonable behavior. Since it's consistent with .NET Framework I think we should probably leave it as-is.
This is, essentially, the way Unix views text files: all lines in a file properly end in a new-line character.
It would, essentially, vastly simplify parsing, because you just consider characters until you find the eol, and eof if after the last eol character.
Since this is a known behavior (and is something we probably can't change), I'm going to close this issue.
Thanks everyone for the info. Even though no changes are likely to come out of it, it was an informative thread.
@rmkerr "Known behaviour" might be a stretch since this information does not appear in the documentation -- we hit it the hard way. Is there somewhere we can open a PR to add a note in the docs?
Yep! You can always open a PR against the docs repo. The easiest way to do that is to go to the page you want to update on docs.microsoft.com, and click the edit button in the upper right hand corner. That should take you to the relevant file in github.
FWIW, I do think this behavior is fairly well documented in the remarks section of the StringReader.ReadLine documentation. If you find a way to make it clearer though, that would be a valuable contribution.
I figured that nothing would change here due to back-compat (makes sense) - How about an issue on the docs to call out this limitation.
Just so everyone understands what the use-case was... this is for PowerShell Editor Services which is used in the PowerShell extension for VSCode.
We read in the contents of files and break them down line-by-line. We were seeing crashes when a user had a newline at EOF because the size of the array that stored all the lines would be one off (because the new line at EOF would not be added to the array because it was null)
oh there were other comments lol
Opened an issue here so we're good: https://github.com/dotnet/docs/issues/8612