Python提取指定区间字符串:跳过首个OP取第二个OP间内容
问题与解决方法
问题描述
需要从文本20231006-OPUS solution _ cfhs0230.22d OP1696574785263-792.txt中提取cfhs0230.22d。尝试了以下Python代码,但因文本存在两个“OP”,当前代码取的是首个OP,无法得到目标内容。需求为跳过首个“OP”,提取“-”和第二个“OP”之间的字符串。
尝试的代码
mystr = "20231006-OPUS solution _ cfhs023 OP1696574785263-792" sub1 = "-" sub2 = "OP" #this has to be second OP idx1 = mystr.index(sub1) idx2 = mystr .index(sub2) print(idx1+ len(sub1) + 1,idx2) for nx in range(idx1+ len(sub1) + 1,idx2): print(nx)
解决方法
方法一:指定查找起始位置
利用str.index()的第二个参数,从第一个“OP”的下一个位置开始查找第二个“OP”,再提取目标内容:
mystr = "20231006-OPUS solution _ cfhs0230.22d OP1696574785263-792.txt" sub1 = "-" sub2 = "OP" # 获取第一个"-"的索引 dash_idx = mystr.index(sub1) # 获取第一个"OP"的索引 first_op_idx = mystr.index(sub2) # 从第一个"OP"之后开始找第二个"OP"的索引 second_op_idx = mystr.index(sub2, first_op_idx + len(sub2)) # 提取并去除前后多余空格 target = mystr[dash_idx + len(sub1):second_op_idx].strip() print(target) # 输出: cfhs0230.22d
方法二:正则表达式匹配
用正则表达式精准定位目标内容,适合更复杂的文本结构:
import re mystr = "20231006-OPUS solution _ cfhs0230.22d OP1696574785263-792.txt" # 正则规则:匹配"-"后跳过首个OP相关内容,捕获第二个OP前的目标字符串 pattern = r'-.*?OP.*?\s+(.*?)\s+OP' match_result = re.search(pattern, mystr) if match_result: print(match_result.group(1)) # 输出: cfhs0230.22d
内容的提问来源于stack exchange,提问作者Sagar Rawal
相关产品推荐
相关产品推荐

