如何让concurrent.futures.ProcessPoolExecutor适配字典输入正常运行
代码问题排查
- 问题1:Windows系统下使用
concurrent.futures.ProcessPoolExecutor必须添加if __name__ == '__main__'入口判断,否则子进程启动会失败,出现无输出、卡住的情况。 - 问题2:
executor.map(函数, 字典)默认遍历字典的key,你传入getCurrency_data的参数是"USD-IDR"这类货币对字符串,但函数内部还在循环整个links字典,导致每个子进程都会重复爬取全部10个链接,完全没有实现多进程拆分任务的效果,反而会重复执行10次全量爬取逻辑。 - 问题3:你没有接收
executor.map返回的结果迭代器,也没有对函数返回的DataFrame做打印、存储等处理,自然看不到爬取结果输出。 - 小笔误:打印耗时的代码里把
seconds拼写为secondes,不影响运行但属于格式错误。
修正后代码
修正思路:将任务拆分为单个货币对的爬取,每个子进程只处理一个链接的爬取任务,最后合并所有子进程返回的结果生成最终DataFrame。
import requests from bs4 import BeautifulSoup import pandas as pd import time import concurrent.futures def getCurrency_data(currency_item): # 每次只处理单个货币对的爬取任务 key, value = currency_item user_agent = "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.88 Safari/537.37" data = requests.get(value, headers={'User-Agent': user_agent}) soup = BeautifulSoup(data.content, 'html.parser') span_tag = [] tags1 = soup.find_all('div', {'class':'top bold inlineblock'}) for div in tags1: spans = div.find_all('span') for span in spans: x = span.text span_tag.append(x) current_tmp = span_tag[0] change_tmp = span_tag[1] cur = [] tags2 = soup.find('div', {'class':'clear overviewDataTable overviewDataTableWithTooltip'}) for a in tags2.findAll('div', {'class':'first inlineblock'}): for b in a.findAll('span', {'class':'float_lang_base_2 bold'}): cur.append(b.text) prevclose_tmp = cur[0] open_tmp = cur[1] oneyearchange_tmp = cur[2] # 返回单条数据字典,方便后续合并 return { "Currency": key, "Current": current_tmp, "Change": change_tmp, "Prev. Close": prevclose_tmp, "Open": open_tmp, "1 Year Change": oneyearchange_tmp } if __name__ == '__main__': t1 = time.perf_counter() links = {"USD-IDR":"https://www.investing.com/currencies/usd-idr", "USD-JPY":"https://www.investing.com/currencies/usd-jpy", "USD-CNY":"https://www.investing.com/currencies/usd-cny", "USD-EUR":"https://www.investing.com/currencies/usd-eur", "USD-SGD":"https://www.investing.com/currencies/usd-sgd", "USD-THB":"https://www.investing.com/currencies/usd-thb", "USD-MXN":"https://www.investing.com/currencies/usd-mxn", "USD-MYR":"https://www.investing.com/currencies/usd-myr", "USD-KRW":"https://www.investing.com/currencies/usd-krw", "USD-INR":"https://www.investing.com/currencies/usd-inr"} with concurrent.futures.ProcessPoolExecutor() as executor: # 传入字典items,每个迭代元素为(货币对, 链接)元组 results = executor.map(getCurrency_data, links.items()) # 合并所有子进程返回结果为DataFrame df_currency = pd.DataFrame(results) # 输出爬取结果 print(df_currency) t2 = time.perf_counter() print(f'Finished in {t2-t1} seconds')
内容的提问来源于stack exchange,提问作者ohai
相关产品推荐
相关产品推荐

