模式匹配中绕过可选字段:Perl正则提取球员信息异常排查
The issue with your original regex is that \d+?\S+? doesn’t properly handle the optional closing </td> tag after the player number—especially for single-digit numbers, where the lazy \d+? might not capture the full number, and \S+? can’t reliably skip between the number and the next <td> in both cases (with or without the closing tag).
Here’s the corrected regex that handles both normal and abnormal entries, without capturing the optional closing tag:
while($content =~ /<tr><td>\d+(?:<\/td>)?<td>(.*?)\s-.*?<\/td><td>(\d+?)<\/td>/g) { my $player = $1; my $total = $2; print "\nPlayer => $player Total => $total\n"; }
Key Changes Explained:
- Replaced
\d+?\S+?<td>with\d+(?:<\/td>)?<td>:\d+: Greedily matches one or more digits (works for single-digit, two-digit, etc., player numbers).(?:<\/td>)?: A non-capturing optional group (the?:avoids creating an extra capture) that matches the closing</td>tag only if it exists. The?makes this group optional, so it handles both cases where the tag is present or missing.<td>: Matches the opening tag for the player name column, which is consistent in both entry types.
Testing with Your Examples:
- Abnormal entry:
<tr><td>10<td>MANNY MACHADO - FA</td><td>37</td></tr>- The regex matches
<tr><td>10<td>, skips the missing closing tag, capturesMANNY MACHADOas the player, and37as the total.
- The regex matches
- Normal entry:
<tr><td>107</td><td>ALEDMYS DIAZ - HOU</td><td>18</td></tr>- The regex matches
<tr><td>107</td><td>, includes the closing tag in the optional group, capturesALEDMYS DIAZas the player, and18as the total.
- The regex matches
This approach explicitly accounts for the optional closing tag without unnecessary captures, and works reliably for all player number lengths.
内容的提问来源于stack exchange,提问作者James Finnegan
相关产品推荐
相关产品推荐

