You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:验证DataFrame A子串是否存在于DataFrame B字符串中

解决Pandas字符串匹配循环的TypeError问题

我有两个DataFrame:

import pandas as pd

countries_list_a = pd.DataFrame({'Country' : ['Australia', 'United Kingdom', 'United States'], 'Code' : ['A', 'B', 'C']})

countries_list_b = pd.DataFrame ({'Name' : ['Jack', 'Maria', 'David'], 'Geo' : ['New York, United States', 'Sydney, Australia', 'London, United Kingdom']})

需要迭代检查countries_list_a的Country列每个字符串是否存在于countries_list_b的Geo列任意单元格中,同时跟踪匹配的索引并存入数组。

手动验证的代码可以正常运行:

tmp_country_a = countries_list_a["Country"][0]
tmp_country_b = countries_list_b["Geo"][1]

if (tmp_country_a in tmp_country_b):
 print ("Correct")

但编写循环时触发了错误:

import numpy as np
index_a = np.size(countries_list_a["Country"])
index_b = np.size(countries_list_b["Geo"])

for i in range(index_a):
 tmp_country_a = countries_list_a["Country"][i]
 for j in range(index_b):
  tmp_country_b = countries_list_b["Geo"][j]
  if (tmp_country_a in tmp_country_b):
   print("Correct")

错误信息:

TypeError: argument of type 'int' is not iterable

错误原因

问题出在np.size()的使用上:np.size()返回的是列的元素总数,但如果DataFrame的列索引不是连续的整数,用range(index_a)去取countries_list_a["Country"][i]时,会取到错误的索引值,甚至返回整数类型的索引而非字符串内容,导致in操作触发类型错误。

另外,嵌套循环处理Pandas数据效率极低,更推荐用Pandas原生的向量化操作。


解决方案

方式1:修复循环逻辑(适合理解基础逻辑)

把np.size()换成len()获取列长度,同时用.iloc[i]按位置取值(避免索引不连续的问题),还能收集匹配的索引对:

index_a = len(countries_list_a["Country"])
index_b = len(countries_list_b["Geo"])

# 存储匹配的索引对(a行索引, b行索引)
matches = []

for i in range(index_a):
    tmp_country_a = countries_list_a["Country"].iloc[i]
    for j in range(index_b):
        tmp_country_b = countries_list_b["Geo"].iloc[j]
        if tmp_country_a in tmp_country_b:
            print("Correct")
            matches.append((i, j))

print("匹配的索引对:", matches)

方式2:Pandas向量化操作(推荐,高效简洁)

利用Pandas的字符串处理和合并功能,更高效地完成匹配并获取索引:

# 从countries_list_b的Geo列提取国家信息,新增Country列
countries_list_b["Country"] = countries_list_b["Geo"].str.split(", ").str[1]

# 按Country列合并两个DataFrame,自动匹配相同国家
merged_df = pd.merge(countries_list_a, countries_list_b, on="Country")

# 提取原DataFrame的索引对
match_indices = list(zip(merged_df.index_x, merged_df.index_y))

print("匹配结果:")
print(merged_df)
print("\n匹配的索引对:", match_indices)

或者用字典映射的方式快速匹配:

# 建立国家到countries_list_b索引的映射
country_to_b_idx = {geo.split(", ")[1]: idx for idx, geo in countries_list_b["Geo"].items()}

# 遍历countries_list_a收集匹配索引
matches = []
for a_idx, country in countries_list_a["Country"].items():
    if country in country_to_b_idx:
        b_idx = country_to_b_idx[country]
        matches.append((a_idx, b_idx))
        print(f"Correct: 国家{country}在countries_list_b的索引{b_idx}处匹配")

print("\n匹配的索引对:", matches)

内容的提问来源于stack exchange,提问作者JohnRambo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 11:01:20