如何用C#提取文本中两个全大写单词之间的子串?
提取两个全大写单词间的子串解决方案
给定字符串:
string str = "THIS is the first line in this PARAGRAPH";
要提取两个全大写单词之间的子串is the first line in this,以下是两种可行方案:
方案一:正则表达式(推荐)
正则可直接匹配首尾全大写单词,捕获中间内容,代码简洁高效:
using System.Text.RegularExpressions; string str = "THIS is the first line in this PARAGRAPH"; var matchResult = Regex.Match(str, @"^[A-Z]+\s+(.*?)\s+[A-Z]+$"); if (matchResult.Success) { string targetSubstring = matchResult.Groups[1].Value; Console.WriteLine(targetSubstring); // 输出:is the first line in this }
正则规则说明:
^[A-Z]+:匹配字符串开头的全大写单词(至少1个大写字母)\s+:匹配单词后的一个或多个空格(.*?):非贪婪模式捕获中间所有内容,避免误匹配结尾的大写单词\s+[A-Z]+$:匹配字符串结尾的全大写单词及前置空格
方案二:拆分单词数组处理
如果不想用正则,可拆分单词后定位首尾大写单词的位置,再拼接中间内容:
string str = "THIS is the first line in this PARAGRAPH"; string[] wordArray = str.Split(new[] {' '}, StringSplitOptions.RemoveEmptyEntries); // 找到第一个全大写单词的索引 int firstUpperPos = Array.FindIndex(wordArray, word => word.All(char.IsUpper)); // 找到最后一个全大写单词的索引 int lastUpperPos = Array.FindLastIndex(wordArray, word => word.All(char.IsUpper)); if (firstUpperPos != -1 && lastUpperPos != -1 && firstUpperPos < lastUpperPos) { string targetSubstring = string.Join(" ", wordArray.Skip(firstUpperPos + 1).Take(lastUpperPos - firstUpperPos - 1)); Console.WriteLine(targetSubstring); // 输出:is the first line in this }
为什么之前的方法无效?
str.All(char.IsUpper):该方法是检查整个字符串是否所有字符都是大写,你的字符串包含小写字母,无法用来定位单个大写单词。str.Contains:仅能判断字符串是否包含指定子串,无法精准定位首尾两个大写单词的位置,自然没法提取中间内容。
内容的提问来源于stack exchange,提问作者James
相关产品推荐
相关产品推荐

