如何简化for循环中BeautifulSoup操作的多try/except异常捕获逻辑
问题根因
你自定义的异常处理函数不生效,核心原因是调用函数时,括号内的表达式会先执行完成再传入函数,异常在进入异常处理逻辑前就已经抛出,自然无法被捕获。
解决方案
将需要执行的取值逻辑封装成无参数可调用对象(比如lambda表达式)传入处理函数,在函数内部执行表达式再做异常捕获,即可正常捕捉对应错误。
第一步:修改异常处理函数
新增KeyError到捕获列表,适配你获取标签属性的字典取值场景:
def ex_handler(func): try: return func() except (AttributeError, IndexError, KeyError): return None
第二步:调整变量取值写法
所有需要做异常捕获的取值操作,都包裹在lambda中传入ex_handler即可,示例:
# 原来的写法(异常提前抛出) last_name = ex_handler(soup.find('span', class_='fn').text.strip().split()[1]) # 改后写法(延迟执行,异常可被捕获) last_name = ex_handler(lambda: soup.find('span', class_='fn').text.strip().split()[1])
改后完整代码
def ex_handler(func): try: return func() except (AttributeError, IndexError, KeyError): return None def get_fighter_meta(fighter_urls): """Scrapes meta from fighters page""" fighter_data = [] for counter, fighter_url in enumerate(fighter_urls, start=1): soup = get_soup(fighter_url) first_name = ex_handler(lambda: soup.find('span', class_='fn').text.strip().split()[0]) last_name = ex_handler(lambda: soup.find('span', class_='fn').text.strip().split()[1]) full_name = ex_handler(lambda: f'{first_name} {last_name}' if first_name and last_name else None) nickname = ex_handler(lambda: soup.find('span', class_='nickname').text.strip()) image_url = ex_handler(lambda: f"https://www.xxxxxx.com/{soup.find('img', attrs={'itemprop': 'image'})['src']}") dob = ex_handler(lambda: soup.find('span', attrs={'itemprop': 'birthDate'}).text.strip()) location = ex_handler(lambda: soup.find('span', class_='locality').text.strip()) nationality = ex_handler(lambda: soup.find('strong', attrs={'itemprop': 'nationality'}).text.strip()) association = ex_handler(lambda: soup.find('span', attrs={'itemprop': 'name'}).text.strip()) height = ex_handler(lambda: soup.find('span', class_='item height').text.strip()[-9:-2]) weight = ex_handler(lambda: soup.find('span', class_='item weight').text.strip()[-9:-2].strip()) weight_class = ex_handler(lambda: soup.find('strong', class_='title').text.strip()) win_loss_loop = ex_handler(lambda: [i.text.strip() for i in soup.find_all('span', class_='counter')]) wins = ex_handler(lambda: win_loss_loop[0]) losses = ex_handler(lambda: win_loss_loop[1]) graph_tag_loop = ex_handler(lambda: [i.text.strip() for i in soup.find_all('span', class_='graph_tag')]) win_ko = ex_handler(lambda: graph_tag_loop[0][:2].strip()) win_submission = ex_handler(lambda: graph_tag_loop[1][:2].strip()) win_decisions = ex_handler(lambda: graph_tag_loop[2][:2].strip()) loss_ko = ex_handler(lambda: graph_tag_loop[3][0][:2].strip()) loss_submission = ex_handler(lambda: graph_tag_loop[4][:2].strip()) loss_decisions = ex_handler(lambda: graph_tag_loop[5][:2].strip()) fighter_meta = { 'First_name': first_name, 'Last_name': last_name, 'Full name': full_name, 'Nickname': nickname, 'Image_url': image_url, 'Date_of_birth': dob, 'Location': location, 'Nationality': nationality, 'Association': association, 'Height': height, 'Weight': weight, 'Weight_class': weight_class, 'Wins': wins, 'Losses': losses, 'Win_by_ko': win_ko, 'Win_by_submission': win_submission, 'Win_decision': win_decisions, 'Loss_by_ko': loss_ko, 'Loss_by_submission': loss_submission, 'Loss_by_desision': loss_decisions } fighter_data.append(fighter_meta) print(f'Saving: {full_name} - {counter} of {len(fighter_urls)}') return fighter_data
补充说明
如果需要排查异常原因,可以在ex_handler的except分支中加打印逻辑,输出异常详情方便定位问题。
内容的提问来源于stack exchange,提问作者DECROMAX
相关产品推荐
相关产品推荐

