如何不忽略空格遍历子串,匹配DataFrame中的多词条目?
问题描述
我有两个DataFrame:
df_globe(国家/地区对照表):
Country Region 0 Andorra Europe 1 Andorran Europe 2 Andorrano Europe 3 United Arab Emirates MENA 4 Saudi Arabia MENA 5 u.a.e MENA 6 Democratic People's Republic of Korea, East Asia 6 Puerto Rico, Americas .. ... ... 539 Americas Americas 540 MENA MENA 541 South Asia South Asia 542 Sub-Saharan Africa Sub-Saharan Africa 543 Pacific Pacific
df_tweets(推特数据):
sentiment id date text 0 0 1598814115664994307 2022-12-02 22:59:24+00:00 I think it is in South Asia 1 0 1598814115664994307 2022-12-02 22:59:24+00:00 Say hello to my friend in Puerto Rico
原代码将推特文本拆分为单个token逐个匹配,导致输出重复且仅匹配到单个词,结果如下:
sentiment id date text word 0 0 1598814115664994307 2022-12-02 22:59:24+00:00 I think it is in South Asia South 0 0 1598814115664994307 2022-12-02 22:59:24+00:00 I think it is in South Asia South ... 0 0 1598814115664994307 2022-12-02 22:59:24+00:00 I think it is in South Asia Asia 1 0 1598814115664994307 2022-12-02 22:59:24+00:00 Say hello to my friend in Puerto Rico Puerto 1 0 1598814115664994307 2022-12-02 22:59:24+00:00 Say hello to my friend in Puerto Rico Rico
期望输出为每条推特匹配完整的国家/地区名称,无重复:
sentiment id date text word 0 0 1598814115664994307 2022-12-02 22:59:24+00:00 I think it is in South Asia South Asia 1 0 1598814115664994307 2022-12-02 22:59:24+00:00 Say hello to my friend in Puerto Rico Puerto Rico
解决方案
核心思路是直接用完整推特文本与df_globe中的国家/地区名称做模糊匹配,同时避免重复记录。
代码实现
import pandas as pd from fuzzywuzzy import fuzz # 读取数据 df_globe = pd.read_csv('fuzzyCountriesAndRegions.csv') # 清理Country列的多余标点(比如末尾逗号) df_globe['Country'] = df_globe['Country'].str.strip().str.rstrip(',') df_tweets = pd.read_csv('tweets.csv') df_result = [] for idx, row in df_tweets.iterrows(): tweet_text = row['text'] matched_countries = set() # 用集合避免重复匹配同一名称 for country in df_globe['Country']: # 用token_set_ratio做模糊匹配,阈值设为90 ratio = fuzz.token_set_ratio(tweet_text, country) if ratio >= 90: matched_countries.add(country) # 为每个匹配到的名称生成结果行 for country in matched_countries: df_result.append({ 'sentiment': row['sentiment'], 'id': row['id'], 'date': row['date'], 'text': row['text'], 'word': country }) # 转换为DataFrame df_result = pd.DataFrame(df_result) print(df_result)
代码说明
- 清理数据:先去除
df_globe['Country']末尾的冗余标点,避免干扰匹配。 - 去重机制:用集合存储匹配到的名称,防止同一推特多次匹配到相同条目。
- 完整匹配:直接用推特全文与完整的国家/地区名称做模糊匹配,确保捕获多词名称(如
South Asia)。 - 结果构建:遍历每条推特收集匹配项,再统一生成结果DataFrame,避免原代码中重复追加的问题。
如果需要更精准的子串匹配,可以将匹配逻辑替换为partial_ratio,它会检查国家名是否是推特文本的子串后再做模糊匹配:
# 替换匹配逻辑 ratio = fuzz.partial_ratio(tweet_text.lower(), country.lower()) if ratio >= 90: matched_countries.add(country)
内容的提问来源于stack exchange,提问作者Rashid Abramov
相关产品推荐
相关产品推荐

