如何用Python正则提取两个问题标记间的多个选中答案
Python 提取第1题中标记为“x”的答案解决方案
两步法(兼容所有Python版本)
这种方法先分离出第1题的答案区块,再提取所有标记为“x”的行。
import re text = """Question 1: What witch-like attributes do you have? Answer 1: x Hat o Pointy Nose x Float x Weigh more than a duck Question 2: Where could this coconut have come from? Answer 2: o It migrated x A European swallow carried it o An African swallow carried it x It doesn't matter """ # 第一步:提取第1题的完整答案区块 q1_match = re.search(r'Question 1:.*?Answer 1:\s*(.*?)\s*Question 2:', text, re.DOTALL) if q1_match: q1_answers_block = q1_match.group(1) # 第二步:提取所有以“x ”开头的行 x_answers = re.findall(r'^x\s+(.*)$', q1_answers_block, re.MULTILINE) print(x_answers) # 输出:['Hat', 'Float', 'Weigh more than a duck']
说明
- 分离第1题区块: 正则表达式
r'Question 1:.*?Answer 1:\s*(.*?)\s*Question 2:'使用re.DOTALL允许匹配换行符,捕获Answer 1:和Question 2:之间的所有内容。 - 提取标记为x的答案: 借助
re.MULTILINE,r'^x\s+(.*)$'匹配每一行以“x ”开头的内容,并捕获答案文本。
单正则表达式法(Python 3.6+)
Python 3.6及以上版本支持 \G 锚点(匹配上一次匹配的结束位置),无需先分离区块即可用单个正则提取所有目标答案。
import re text = """Question 1: What witch-like attributes do you have? Answer 1: x Hat o Pointy Nose x Float x Weigh more than a duck Question 2: Where could this coconut have come from? Answer 2: o It migrated x A European swallow carried it o An African swallow carried it x It doesn't matter """ x_answers = re.findall( r'(?:\G(?!^)|Question 1:)(?:(?!Question 1:|Question 2:)[\s\S])*?x\s+([\s\S]+?)(?=\s*(?:x|o|Question 2:))', text ) print(x_answers) # 输出:['Hat', 'Float', 'Weigh more than a duck']
说明
(?:\G(?!^)|Question 1:): 从Question 1:开始匹配,或在上一次匹配结束的位置继续匹配(确保始终处于第1题的范围内)。(?:(?!Question 1:|Question 2:)[\s\S])*?: 非贪婪匹配所有不属于新问题开头的字符。x\s+([\s\S]+?): 捕获“x ”之后的答案文本。(?=\s*(?:x|o|Question 2:)): 当遇到下一行以“x”、“o”开头或到达Question 2:时停止捕获。
内容的提问来源于stack exchange,提问作者Adam Brand
相关产品推荐
相关产品推荐

