Python Pandas:验证DataFrame A子串是否存在于DataFrame B字符串中
解决Pandas字符串匹配循环的TypeError问题
我有两个DataFrame:
import pandas as pd countries_list_a = pd.DataFrame({'Country' : ['Australia', 'United Kingdom', 'United States'], 'Code' : ['A', 'B', 'C']}) countries_list_b = pd.DataFrame ({'Name' : ['Jack', 'Maria', 'David'], 'Geo' : ['New York, United States', 'Sydney, Australia', 'London, United Kingdom']})
需要迭代检查countries_list_a的Country列每个字符串是否存在于countries_list_b的Geo列任意单元格中,同时跟踪匹配的索引并存入数组。
手动验证的代码可以正常运行:
tmp_country_a = countries_list_a["Country"][0] tmp_country_b = countries_list_b["Geo"][1] if (tmp_country_a in tmp_country_b): print ("Correct")
但编写循环时触发了错误:
import numpy as np index_a = np.size(countries_list_a["Country"]) index_b = np.size(countries_list_b["Geo"]) for i in range(index_a): tmp_country_a = countries_list_a["Country"][i] for j in range(index_b): tmp_country_b = countries_list_b["Geo"][j] if (tmp_country_a in tmp_country_b): print("Correct")
错误信息:
TypeError: argument of type 'int' is not iterable
错误原因
问题出在np.size()的使用上:np.size()返回的是列的元素总数,但如果DataFrame的列索引不是连续的整数,用range(index_a)去取countries_list_a["Country"][i]时,会取到错误的索引值,甚至返回整数类型的索引而非字符串内容,导致in操作触发类型错误。
另外,嵌套循环处理Pandas数据效率极低,更推荐用Pandas原生的向量化操作。
解决方案
方式1:修复循环逻辑(适合理解基础逻辑)
把np.size()换成len()获取列长度,同时用.iloc[i]按位置取值(避免索引不连续的问题),还能收集匹配的索引对:
index_a = len(countries_list_a["Country"]) index_b = len(countries_list_b["Geo"]) # 存储匹配的索引对(a行索引, b行索引) matches = [] for i in range(index_a): tmp_country_a = countries_list_a["Country"].iloc[i] for j in range(index_b): tmp_country_b = countries_list_b["Geo"].iloc[j] if tmp_country_a in tmp_country_b: print("Correct") matches.append((i, j)) print("匹配的索引对:", matches)
方式2:Pandas向量化操作(推荐,高效简洁)
利用Pandas的字符串处理和合并功能,更高效地完成匹配并获取索引:
# 从countries_list_b的Geo列提取国家信息,新增Country列 countries_list_b["Country"] = countries_list_b["Geo"].str.split(", ").str[1] # 按Country列合并两个DataFrame,自动匹配相同国家 merged_df = pd.merge(countries_list_a, countries_list_b, on="Country") # 提取原DataFrame的索引对 match_indices = list(zip(merged_df.index_x, merged_df.index_y)) print("匹配结果:") print(merged_df) print("\n匹配的索引对:", match_indices)
或者用字典映射的方式快速匹配:
# 建立国家到countries_list_b索引的映射 country_to_b_idx = {geo.split(", ")[1]: idx for idx, geo in countries_list_b["Geo"].items()} # 遍历countries_list_a收集匹配索引 matches = [] for a_idx, country in countries_list_a["Country"].items(): if country in country_to_b_idx: b_idx = country_to_b_idx[country] matches.append((a_idx, b_idx)) print(f"Correct: 国家{country}在countries_list_b的索引{b_idx}处匹配") print("\n匹配的索引对:", matches)
内容的提问来源于stack exchange,提问作者JohnRambo
相关产品推荐
相关产品推荐

