Python中如何将Pool.map返回的含None列表转换为有效字典
解决多进程Pool.map返回结果的字典整理问题
先给你捋清楚问题根源:一是多进程环境下用全局字典的坑,二是返回结果的格式需要清洗。咱们一步步来搞定:
1. 重构你的get_accomodation_wifi_rating函数
原来的函数里操作全局字典hotels_with_low_wifi_quality在多进程里完全行不通——每个子进程都会创建自己的字典副本,所以返回的结果要么是None,要么是某个子进程里的零散小字典。咱们改成让函数返回单个酒店的有效数据(或者None),逻辑更清晰也更安全:
def get_accomodation_wifi_rating(hotel): page = requests.get(hotel) soup = BeautifulSoup(page.text, 'lxml') # 清洗酒店名称:去掉前后换行和多余空白 hotel_name = soup.find(id="hp_hotel_name").text.strip() reviews = soup.find(class_="review_list_score_breakdown_right") # 层层判断,不符合条件直接返回None if reviews is None: return None wifi_tag = reviews.find(attrs={"data-question": "hotel_wifi"}) if wifi_tag is None: return None wifi_rating = wifi_tag.find(class_="review_score_value") if wifi_rating is None: return None wifi_score = wifi_rating.text.strip() # 把评分转换成float类型,替换逗号为点 wifi_score_num = float(wifi_score.replace(",", ".")) # 只返回评分低于7的酒店数据,用元组更简洁 if wifi_score_num < 7: return (hotel_name, wifi_score_num) else: return None
2. 处理Pool.map返回的结果列表
现在map()会返回一个包含元组或None的列表,咱们只需要过滤掉无效的None,再把有效元组转换成干净的字典:
from multiprocessing import Pool # 假设你的hotels_url_list已经定义好 with Pool() as pool: results = pool.map(get_accomodation_wifi_rating, hotels_url_list) # 用字典推导式快速生成最终字典,简洁高效 hotels_with_low_wifi_quality = { name: score for name, score in results if name is not None # 过滤掉None项 }
额外说明:如果不想改原函数怎么办?
要是你坚持保留原函数返回小字典的逻辑,也可以这么处理结果:
results = pool.map(get_accomodation_wifi_rating, hotels_url_list) hotels_with_low_wifi_quality = {} for d in results: if d is not None: # 逐个清洗小字典里的键值对 for raw_name, raw_score in d.items(): clean_name = raw_name.strip() clean_score = float(raw_score.replace(",", ".")) hotels_with_low_wifi_quality[clean_name] = clean_score
但还是更推荐第一种重构函数的方式——多进程环境下尽量避免用全局变量,主进程统一处理结果才是更可靠的做法,逻辑也更容易维护。
内容的提问来源于stack exchange,提问作者Ian Spitz
相关产品推荐
相关产品推荐

