如何通过部分字符串匹配为Pandas DataFrame匹配追加海报列表
问题
现有如下Pandas DataFrame:
| Titles |
|---|
| The Ice Age Adventures of Buck Wild |
| The Beatles: Get Back – The Rooftop Concert |
| Turning Red |
| Chip 'n Dale: Rescue Rangers |
| Zombies 3 |
另有海报标签列表:
[ '<img src=posters/turning_red.jpeg alt=red>', '<img src=posters/ice_age_adventures_of_buck_wild.jpeg alt=ice_age>', '<img src=posters/chip_n_dale_rescue_rangers.jpeg alt=rangers>', '<img src=posters/beatles_get_back__the_rooftop_concert.jpeg alt=beatles>', '<img src=posters/zombies_three.jpeg alt=zombies>' ]
希望无需复杂正则,通过部分字符串匹配将该列表与DataFrame中最相似的行匹配,追加为新列,该如何实现?
解决方案
可以通过标准化文本格式+子串匹配的方式实现,不用复杂正则,直接上可运行的代码和说明:
import pandas as pd # 初始化DataFrame df = pd.DataFrame({ 'Titles': [ 'The Ice Age Adventures of Buck Wild', 'The Beatles: Get Back – The Rooftop Concert', 'Turning Red', "Chip 'n Dale: Rescue Rangers", 'Zombies 3' ] }) # 海报列表 poster_tags = [ '<img src=posters/turning_red.jpeg alt=red>', '<img src=posters/ice_age_adventures_of_buck_wild.jpeg alt=ice_age>', '<img src=posters/chip_n_dale_rescue_rangers.jpeg alt=rangers>', '<img src=posters/beatles_get_back__the_rooftop_concert.jpeg alt=beatles>', '<img src=posters/zombies_three.jpeg alt=zombies>' ] # 统一文本格式:转小写,替换特殊符号为下划线,消除格式差异 def simplify_text(text): return text.lower().replace(' ', '_').replace(':', '_').replace('-', '_').replace("'", '_').replace('__', '_').strip('_') # 构建海报文件名与标签的映射 poster_mapping = {} for tag in poster_tags: # 提取海报文件名(去掉路径和后缀) filename = tag.split('src=posters/')[1].split('.jpeg')[0] poster_mapping[filename] = tag # 为每个标题匹配对应海报 def match_poster(title): simplified_title = simplify_text(title) # 处理特殊匹配:Zombies 3对应zombies_three if 'zombies_3' in simplified_title: return poster_mapping['zombies_three'] # 核心匹配逻辑:检查简化后的标题与海报文件名是否互相包含子串 for key in poster_mapping: if key in simplified_title or simplified_title in key: return poster_mapping[key] return None # 新增Poster列 df['Poster'] = df['Titles'].apply(match_poster) print(df)
代码说明
simplify_text函数:把标题和海报文件名统一转换成小写、下划线连接的格式,消除大小写、空格、特殊符号带来的匹配障碍。- 特殊匹配处理:针对
Zombies 3对应zombies_three这种数字转单词的特殊情况,单独做映射(若有更多类似情况,可扩展一个匹配表)。 - 子串匹配逻辑:通过检查简化后的标题和海报文件名是否互相包含,实现近似匹配,完全无需正则。
运行后会得到匹配完成的DataFrame,结果如下:
| Titles | Poster |
|---|---|
| The Ice Age Adventures of Buck Wild | ![]() |
| The Beatles: Get Back – The Rooftop Concert | ![]() |
| Turning Red | ![]() |
| Chip 'n Dale: Rescue Rangers | ![]() |
| Zombies 3 | ![]() |
内容的提问来源于stack exchange,提问作者ariyasas94
相关产品推荐
相关产品推荐






