多进程池结果追加至DataFrame报错:TypeError类型拼接异常求助
问题解决:TypeError 无法将整数存入DataFrame
你的报错原因很明确:df.append()仅支持拼接Series或DataFrame对象,你直接传入整数类型的price,同时也没把对应的URL关联存入,导致类型不匹配。
修正方案1:循环中构建Series追加
把每个url和price包装成一个Series,指定列名后再追加:
import json from bs4 import BeautifulSoup import requests from multiprocessing import Pool import pandas as pd df = pd.DataFrame(columns=['url', 'price']) def get_price(url): soup = BeautifulSoup(requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}).content, "html.parser") data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text()) return url, int(data['offers']['price']) if __name__ == '__main__': urls = [ 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-hardwear-graduated-link-necklace-63008966/', 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-t-smile-pendant-35189459/', ] with Pool(processes=4) as pool: for url, price in pool.imap_unordered(get_price, urls): # 构建Series,指定列名后追加 df = df.append(pd.Series({'url': url, 'price': price}), ignore_index=True) # 导出到Excel/CSV df.to_excel('tiffany_prices.xlsx', index=False) df.to_csv('tiffany_prices.csv', index=False)
修正方案2:先收集所有结果再创建DataFrame(推荐)
循环追加DataFrame效率较低,更优的方式是先把多进程返回的所有(url, price)元组收集成列表,再一次性生成DataFrame:
import json from bs4 import BeautifulSoup import requests from multiprocessing import Pool import pandas as pd def get_price(url): soup = BeautifulSoup(requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}).content, "html.parser") data = json.loads(soup.find_all('script', {'type': 'application/ld+json'})[-1].get_text()) return url, int(data['offers']['price']) if __name__ == '__main__': urls = [ 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-hardwear-graduated-link-necklace-63008966/', 'https://www.tiffany.co.uk/jewelry/necklaces-pendants/tiffany-t-smile-pendant-35189459/', ] with Pool(processes=4) as pool: # 收集所有结果为列表 results = list(pool.imap_unordered(get_price, urls)) # 从列表直接创建DataFrame,指定列名 df = pd.DataFrame(results, columns=['url', 'price']) # 导出到Excel/CSV df.to_excel('tiffany_prices.xlsx', index=False) df.to_csv('tiffany_prices.csv', index=False)
额外提示
- pandas 2.0+版本中
append()方法已被标记为废弃,推荐使用pd.concat(),不过方案2的方式本身就更高效且符合新版本规范。 - 多进程爬取时建议增加异常处理,避免单个URL请求失败导致整个程序崩溃,比如在
get_price函数中添加try-except块。
内容的提问来源于stack exchange,提问作者Seedizens
相关产品推荐
相关产品推荐

