Pytrends遍历DataFrame报cannot convert series to int错误
报错根因
- 遍历逻辑完全失效:
for id in df_full_1仅会遍历DataFrame的列名,不会逐行读取行数据,且循环内部完全没有使用遍历变量取值,循环写了等于没写 - 参数类型不匹配:
pytrend.get_historical_interest要求year_start/month_start/day_start等时间参数必须传入单个整数,第一个搜索词参数要求传入单个字符串或字符串列表,但你直接传入了df_full_1["year"]这类整列Series对象。你虽然确认了列的存储类型是int64,但传入的是包含上百个值的多值序列,pandas无法将其转换为方法要求的单个int值,因此抛出类型错误。
修正后代码
import pandas as pd from pytrends.request import TrendReq pytrend = TrendReq() result_list = [] # 正确逐行遍历DataFrame for idx, row in df_full_1.iterrows(): # 逐行取出单值参数 keyword = row["ICO_Name"] y = row["year"] m = row["month"] d = row["day"] # 拉取单条关键词的对应时间趋势数据 single_trend = pytrend.get_historical_interest( keyword, year_start=y, month_start=m, day_start=d, year_end=y, month_end=m, day_end=d, sleep=1 # 不要设为0,加1秒间隔避免触发反爬 ).drop(columns='isPartial') # 补充ICO标识字段,方便后续区分数据归属 single_trend["ICO_Name"] = keyword result_list.append(single_trend) # 拼接所有结果为完整DataFrame g_trends = pd.concat(result_list).reset_index(drop=True)
补充注意事项
- 不要将
sleep参数设为0,高频请求很容易触发谷歌趋势的反爬机制,导致IP被临时封禁,建议设置1-2秒的请求间隔 - 如果需要拉取ICO上线前后一段时间的趋势数据(比如上线前后1个月),只需要对应调整结束时间的年/月/日取值即可,禁止直接传入DataFrame整列作为参数
- 如果单次需要查询多个关键词,可以将关键词组装为Python原生列表传入,不要直接传入pandas Series对象
内容的提问来源于stack exchange,提问作者Marc Ericsson
相关产品推荐
相关产品推荐

