如何在含重复值的Pandas DataFrame中按Text列提取指定行
问题描述
我有如下Pandas DataFrame:
| id | left | top | width | height | Text |
|---|---|---|---|---|---|
| 1 | 12 | 34 | 12 | 34 | commercial |
| 2 | 99 | 42 | 99 | 42 | general |
| 3 | 1 | 47 | 9 | 4 | liability |
| 4 | 10 | 69 | 32 | 67 | commercial |
| 5 | 99 | 72 | 79 | 88 | available |
需要基于Text列提取特定行:使用正则匹配搜索类似“liability commercial”的关键词组,只要Text列的值匹配组内任意一个关键词,就提取对应行。比如输入该关键词组时,需提取第3、4行(结果如下),且要保留Text列的重复值行。
| id | left | top | width | height | Text |
|---|---|---|---|---|---|
| 3 | 1 | 47 | 9 | 4 | liability |
| 4 | 10 | 69 | 32 | 67 | commercial |
解决方案
可以通过拆分关键词组构建正则模式,结合Pandas的str.contains方法实现筛选:
import pandas as pd import re # 构造示例DataFrame df = pd.DataFrame({ 'id': [1,2,3,4,5], 'left': [12,99,1,10,99], 'top': [34,42,47,69,72], 'width': [12,99,9,32,79], 'height': [34,42,4,67,88], 'Text': ['commercial','general','liability','commercial','available'] }) # 定义目标关键词组 target_keywords = "liability commercial" # 转义关键词并拼接成正则匹配模式(匹配任意一个关键词) regex_pattern = '|'.join(re.escape(word) for word in target_keywords.split()) # 筛选匹配的行 filtered_df = df[df['Text'].str.contains(regex_pattern, regex=True)] print(filtered_df)
关键说明
re.escape(word):对每个关键词做转义处理,避免关键词中的特殊字符破坏正则匹配逻辑;'|'.join(...):将关键词用|拼接,形成“liability|commercial”的正则模式,实现“匹配任意一个关键词”的效果;df['Text'].str.contains(...):检查Text列每个值是否匹配正则模式,返回布尔索引;- 用布尔索引筛选DataFrame,得到所有符合条件的行,包含重复值的行也会被保留。
内容的提问来源于stack exchange,提问作者spectre
相关产品推荐
相关产品推荐

