You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过部分字符串匹配为Pandas DataFrame匹配追加海报列表

问题

现有如下Pandas DataFrame:

Titles
The Ice Age Adventures of Buck Wild
The Beatles: Get Back – The Rooftop Concert
Turning Red
Chip 'n Dale: Rescue Rangers
Zombies 3

另有海报标签列表:

[
'<img src=posters/turning_red.jpeg alt=red>', 
'<img src=posters/ice_age_adventures_of_buck_wild.jpeg alt=ice_age>', 
'<img src=posters/chip_n_dale_rescue_rangers.jpeg alt=rangers>', 
'<img src=posters/beatles_get_back__the_rooftop_concert.jpeg alt=beatles>', 
'<img src=posters/zombies_three.jpeg alt=zombies>'
]

希望无需复杂正则,通过部分字符串匹配将该列表与DataFrame中最相似的行匹配,追加为新列,该如何实现?

解决方案

可以通过标准化文本格式+子串匹配的方式实现,不用复杂正则,直接上可运行的代码和说明:

import pandas as pd

# 初始化DataFrame
df = pd.DataFrame({
    'Titles': [
        'The Ice Age Adventures of Buck Wild',
        'The Beatles: Get Back – The Rooftop Concert',
        'Turning Red',
        "Chip 'n Dale: Rescue Rangers",
        'Zombies 3'
    ]
})

# 海报列表
poster_tags = [
    '<img src=posters/turning_red.jpeg alt=red>', 
    '<img src=posters/ice_age_adventures_of_buck_wild.jpeg alt=ice_age>', 
    '<img src=posters/chip_n_dale_rescue_rangers.jpeg alt=rangers>', 
    '<img src=posters/beatles_get_back__the_rooftop_concert.jpeg alt=beatles>', 
    '<img src=posters/zombies_three.jpeg alt=zombies>'
]

# 统一文本格式:转小写,替换特殊符号为下划线,消除格式差异
def simplify_text(text):
    return text.lower().replace(' ', '_').replace(':', '_').replace('-', '_').replace("'", '_').replace('__', '_').strip('_')

# 构建海报文件名与标签的映射
poster_mapping = {}
for tag in poster_tags:
    # 提取海报文件名(去掉路径和后缀)
    filename = tag.split('src=posters/')[1].split('.jpeg')[0]
    poster_mapping[filename] = tag

# 为每个标题匹配对应海报
def match_poster(title):
    simplified_title = simplify_text(title)
    # 处理特殊匹配:Zombies 3对应zombies_three
    if 'zombies_3' in simplified_title:
        return poster_mapping['zombies_three']
    # 核心匹配逻辑:检查简化后的标题与海报文件名是否互相包含子串
    for key in poster_mapping:
        if key in simplified_title or simplified_title in key:
            return poster_mapping[key]
    return None

# 新增Poster列
df['Poster'] = df['Titles'].apply(match_poster)

print(df)

代码说明

  1. simplify_text函数:把标题和海报文件名统一转换成小写、下划线连接的格式,消除大小写、空格、特殊符号带来的匹配障碍。
  2. 特殊匹配处理:针对Zombies 3对应zombies_three这种数字转单词的特殊情况,单独做映射(若有更多类似情况,可扩展一个匹配表)。
  3. 子串匹配逻辑:通过检查简化后的标题和海报文件名是否互相包含,实现近似匹配,完全无需正则。

运行后会得到匹配完成的DataFrame,结果如下:

TitlesPoster
The Ice Age Adventures of Buck Wildice_age
The Beatles: Get Back – The Rooftop Concertbeatles
Turning Redred
Chip 'n Dale: Rescue Rangersrangers
Zombies 3zombies

内容的提问来源于stack exchange,提问作者ariyasas94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 20:12:46