You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不忽略空格遍历子串,匹配DataFrame中的多词条目?

问题描述

我有两个DataFrame:

  • df_globe(国家/地区对照表):
Country                Region
0                 Andorra                Europe
1                Andorran                Europe
2               Andorrano                Europe
3    United Arab Emirates                 MENA
4            Saudi Arabia                 MENA
5                    u.a.e                MENA
6 Democratic People's Republic of Korea,  East Asia
6               Puerto Rico,              Americas
..                    ...                 ...
539              Americas            Americas
540                  MENA               MENA
541            South Asia          South Asia
542    Sub-Saharan Africa  Sub-Saharan Africa
543               Pacific             Pacific
  • df_tweets(推特数据):
sentiment                   id                       date                                               text
0          0  1598814115664994307  2022-12-02 22:59:24+00:00  I think it is in South Asia
1          0  1598814115664994307  2022-12-02 22:59:24+00:00  Say hello to my friend in Puerto Rico

原代码将推特文本拆分为单个token逐个匹配,导致输出重复且仅匹配到单个词,结果如下:

sentiment                   id                       date                                   text    word
0         0  1598814115664994307  2022-12-02 22:59:24+00:00            I think it is in South Asia   South
0         0  1598814115664994307  2022-12-02 22:59:24+00:00            I think it is in South Asia   South
...
0         0  1598814115664994307  2022-12-02 22:59:24+00:00            I think it is in South Asia    Asia
1         0  1598814115664994307  2022-12-02 22:59:24+00:00  Say hello to my friend in Puerto Rico  Puerto
1         0  1598814115664994307  2022-12-02 22:59:24+00:00  Say hello to my friend in Puerto Rico    Rico

期望输出为每条推特匹配完整的国家/地区名称,无重复:

sentiment                   id                       date                                   text          word
0          0  1598814115664994307  2022-12-02 22:59:24+00:00            I think it is in South Asia   South Asia
1          0  1598814115664994307  2022-12-02 22:59:24+00:00  Say hello to my friend in Puerto Rico  Puerto Rico
解决方案

核心思路是直接用完整推特文本与df_globe中的国家/地区名称做模糊匹配,同时避免重复记录。

代码实现

import pandas as pd
from fuzzywuzzy import fuzz

# 读取数据
df_globe = pd.read_csv('fuzzyCountriesAndRegions.csv')
# 清理Country列的多余标点(比如末尾逗号)
df_globe['Country'] = df_globe['Country'].str.strip().str.rstrip(',')
df_tweets = pd.read_csv('tweets.csv')

df_result = []

for idx, row in df_tweets.iterrows():
    tweet_text = row['text']
    matched_countries = set()  # 用集合避免重复匹配同一名称
    for country in df_globe['Country']:
        # 用token_set_ratio做模糊匹配,阈值设为90
        ratio = fuzz.token_set_ratio(tweet_text, country)
        if ratio >= 90:
            matched_countries.add(country)
    # 为每个匹配到的名称生成结果行
    for country in matched_countries:
        df_result.append({
            'sentiment': row['sentiment'],
            'id': row['id'],
            'date': row['date'],
            'text': row['text'],
            'word': country
        })

# 转换为DataFrame
df_result = pd.DataFrame(df_result)
print(df_result)

代码说明

  1. 清理数据:先去除df_globe['Country']末尾的冗余标点,避免干扰匹配。
  2. 去重机制:用集合存储匹配到的名称,防止同一推特多次匹配到相同条目。
  3. 完整匹配:直接用推特全文与完整的国家/地区名称做模糊匹配,确保捕获多词名称(如South Asia)。
  4. 结果构建:遍历每条推特收集匹配项,再统一生成结果DataFrame,避免原代码中重复追加的问题。

如果需要更精准的子串匹配,可以将匹配逻辑替换为partial_ratio,它会检查国家名是否是推特文本的子串后再做模糊匹配:

# 替换匹配逻辑
ratio = fuzz.partial_ratio(tweet_text.lower(), country.lower())
if ratio >= 90:
    matched_countries.add(country)

内容的提问来源于stack exchange,提问作者Rashid Abramov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:01:00