使用正则捕获组与Pandas extractall提取句子列表元素的技术问询
Pandas正则提取列表元素实操方案
原始数据
| index | sentence |
|---|---|
| 0 | You can get cars, trucks, planes, and boats. |
| 1 | You can get the car, truck, and plane. |
| 2 | You should ignore this sentence. |
期望输出
| index | match | object |
|---|---|---|
| 0 | 0 | car |
| 1 | truck | |
| 2 | plane | |
| 3 | boat | |
| 1 | 0 | car |
| 1 | truck | |
| 2 | plane |
核心问题解答
1. 用后向断言避免捕获“the”
别用(?<=[Y|y]ou can get )这种错误写法([Y|y]是字符集,会误匹配|符号),改用包含可选“the”的后向断言:
(?<=You can get(?: the)?\s)
其中(?: the)?是非捕获可选组,匹配带空格的“the”,整个断言会定位到“You can get ”或“You can get the ”之后的位置,确保“the”不会混入提取结果。如果需要忽略大小写,调用extractall时加flags=re.IGNORECASE参数即可。
2. 单复数统一捕获
不用\w+(?=s)?这种逻辑混乱的写法,直接用非捕获组匹配末尾的s,同时捕获单词的单数核心:
(\w+?)(?:s\b)?
\w+?非贪婪匹配单词主体,避免把复数s包含进去;(?:s\b)?是可选的非捕获组,匹配单词结尾的s(\b确保是单词末尾的s);
不管是复数(cars)还是单数(car),都能提取出单数形式的核心词。如果要适配不规则复数(如bus→bus),可以调整为(\w+)(?<!s)s?,确保只有非s结尾的单词才去掉末尾s。
3. 直接提取单个元素,无需分步处理
完全可以一步写正则直接提取单个元素,不用先抓整段列表再拆分。利用extractall遍历所有匹配项的特性,结合多条件后向断言匹配元素的前置场景:
(?:(?<=You can get(?: the)?\s)|(?<=,\s)|(?<=and\s))(\w+?)(?:s\b)?
这个正则会匹配三种位置后的元素:
- 开头短语“You can get( the) ”之后;
- 逗号加空格之后;
- and加空格之后;
一次性把所有列表元素提取出来。
完整代码示例
import pandas as pd import re # 构造原始数据 data = { 'index': [0, 1, 2], 'sentence': [ "You can get cars, trucks, planes, and boats.", "You can get the car, truck, and plane.", "You should ignore this sentence." ] } df = pd.DataFrame(data).set_index('index') # 最终正则表达式 pattern = r'(?:(?<=You can get(?: the)?\s)|(?<=,\s)|(?<=and\s))(\w+?)(?:s\b)?' # 执行提取并重命名列 result = df['sentence'].str.extractall(pattern, flags=re.IGNORECASE).rename(columns={0: 'object'}) print(result)
运行后输出就是期望的表格结构,自动过滤掉不符合开头规则的句子(如index=2的句子)。
内容的提问来源于stack exchange,提问作者Jeff
相关产品推荐
相关产品推荐

