如何提取两个子字符串间的文本?附字符串分割示例问询
提取两个子字符串之间文本的方法及示例解析
嘿,这问题我熟!提取两个子串之间的文本有几种实用方法,我结合你给的示例一步步讲清楚。
方法一:用字符串find()手动定位截取
这种方法适合标记固定、结构简单的场景,不用额外依赖库,直观易懂。
先看你的示例文本:
text = 'TEXT 1: Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry\'s standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing L...'
假设我们要提取TEXT 1: 之后,到It was popularised之前的内容,步骤如下:
- 先定义起始和结束标记:
start_mark = 'TEXT 1: ' end_mark = 'It was popularised'
- 定位起始位置:用
find()拿到起始标记的索引,再加上标记本身的长度,精准定位到要提取内容的开头; - 定位结束位置:直接用
find()拿到结束标记的索引,这就是要提取内容的结尾; - 截取中间文本,用
strip()去掉前后多余空格。
完整代码:
text = 'TEXT 1: Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry\'s standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing L...' start_mark = 'TEXT 1: ' end_mark = 'It was popularised' # 定位起始位置 start_idx = text.find(start_mark) + len(start_mark) # 定位结束位置 end_idx = text.find(end_mark) # 提取并处理文本 result = text[start_idx:end_idx].strip() print(result)
如果你的结束标记是文本末尾,直接把end_idx设为len(text)就行,省掉找结束标记的步骤。
方法二:用正则表达式灵活匹配
如果标记有变化(比如大小写不固定、有多种格式),或者需要一次性提取多个匹配结果,正则表达式会更强大。
还是用你的示例,提取TEXT 1: 到It was popularised之间的内容,代码如下:
import re text = 'TEXT 1: Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry\'s standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing L...' # 正则模式:起始标记 + 捕获组(非贪婪匹配任意内容) + 结束标记 pattern = r'TEXT 1: (.*?)It was popularised' # re.DOTALL 让.可以匹配换行符(如果文本有换行的话) match = re.search(pattern, text, re.DOTALL) if match: # 提取捕获组里的内容,再处理空格 result = match.group(1).strip() print(result)
这里的(.*?)是非贪婪匹配,意思是匹配到第一个结束标记就停止,避免把后面无关内容也包含进来。如果用(.*)贪婪模式,会匹配到文本里最后一个结束标记(如果有多个的话)。
总结一下
- 简单固定标记:用
find()方法,代码简洁,不需要额外库; - 复杂/多变标记:用正则表达式,灵活性拉满。
你可以根据实际需求选对应的方法哦!
内容的提问来源于stack exchange,提问作者Rade
相关产品推荐
相关产品推荐

