First of all, I'm hoping that this is the right place to post this.
Here is the original StackOverflow post.
There's this neat little MARK keyword in the PCRE Regex specification:
https://pcre.org/current/doc/html/pcre2syntax.html#SEC23
<?php
$string = 'true';
$matches = [];
preg_match('~(?|true(*:1)
|false(*:1)
|\d+(*:2))~x', $string, $matches);
// *MARK:1 stands for a boolean
// *MARK:2 stands for a integer literal
var_dump($matches);
//> array(2) {
//> [0]=> string(4) "true"
//> ["MARK"]=> string(1) "1"
//> }
$token = new Token(
$lexeme = $matches[0],
$type = $matches['MARK']
);
As you can see, MARK allows you to understand which group matched, not just if it was matched. In this example MARK would contain the token type which would make tokenizing mind blowingly elegant.
Unfortunately, right now I have to create capture groups and then manually check which capture group was matched when.
var pattern = @"(
(?<true>true)
|(?<false>false)
|(?<integer>\d+)
)";
var regexOptions = RegexOptions.ExplicitCapture | RegexOptions.IgnorePatternWhitespace;
var regex = new Regex(pattern, regexOptions);
var matches = regex.Matches("true");
foreach (Match match in matches)
{
int? mark = null;
if (match.Groups["true"].Success || match.Groups["false"].Success)
{
mark = 1;
}
else if (match.Groups["integer"].Success)
{
mark = 2;
}
}
I mean, it works. But this will have to be repeated for every symbol, every operator, everything. Since the regex engine already knows what happened it should be able to report it back.
Do you see any use in this feature?
Thanks @iluuu1994 for the suggestion. As we are currently busy with finishing up for 2.1 let me get back to this item later (in a few weeks).
@ViktorHofer 76 weeks and counting...
@Bartolomeus-649 I mean, I'm the only person who asked for it so probably not exactly a priority 馃槅
no one asked for foreach, but when it was there, everyone used it.
this is one of those many small things that makes a platform great instead of just good enough...they might not seem like a big thing or being a priority, but every time someone copies a regex from somewhere to .NET and it doesn't work, then guess what, it is .NET that doesn't "work" and get blamed.
When it comes to regex there's a de-facto standard, PCRE, and until you have implemented the sane and most frequently used parts of it, you're not really done and will always be considered "the lesser" option you'll have to live with.
There are open source implementations of PCRE for .NET, but it need to be part of the base library.
Thanks for bringing this up again. Regex is not a priority at the moment as we are working on finishing 3.0 and preparing for bigger 5.0 changes both infrastructure and API wise. If I understand correctly this isn't blocking anyone as a workaround is provided.
@Bartolomeus-649 as our code-base is open-source, you are welcome to send a PR. If you are willing to, we can discuss necessary changes.
@ViktorHofer Sorry, but I'm busy helping Microsofts customers, and trying to answer their questions on why just about everything new out from Microsoft doesn't work with their current Microsoft based IT infrastructure, which they have invested in for decades.