使用Selenium爬取Trustpilot时无明显错误触发except的原因排查
select_cat_and_or_subcat方法触发异常的原因及修复 问题背景
我正在用Selenium爬取https://uk.trustpilot.com/,已完成Cookie授权、点击「查看全部」获取分类及子分类列表,get_catogories方法返回字典结构如下:
- 键:分类名称
- 值:(分类URL, 子分类字典)
- 子分类字典:键为子分类名称,值为子分类URL
实现的select_cat_and_or_subcat方法代码如下:
def select_cat_and_or_subcat(self): """ Allows the user to select which catorgy and sub catogory where they would like to begin extracting data from """ #gets the data of catogories and sub_cat catogories_and_sub_cat = self.get_catogories() keys = catogories_and_sub_cat.keys() #presents the catogory and subcat to the user for key in keys: print('---------------------------------------') print(f'Catogory: {key} ') _, sub_cat = catogories_and_sub_cat[f'{key}'] for sub_key in sub_cat.keys(): print(f"SubCatorgy: {sub_key}") try: user_catogory = input('Please select which catogory you would like to explore: ') #uses the user input as a key and gets out the cat url and sub_cat dictionary cat_url, cat_sub_cat = catogories_and_sub_cat[f'{user_catogory}'] #may like to explore and extract data from main catorgory or sub_cat user_sub_catorgory = input('Please select the sub-catorgory or 0 for no sub-catorgory: ') except: print('exiting....due to invalid catorgory being selected') try: if int(user_sub_catorgory) == 0: #if zero navigate to main cat url self.driver.get(cat_url) print('done') #else if sub_cat check if its within the cat_sub keys elif user_sub_catorgory in cat_sub_cat.keys(): sub_url_for_cat = cat_sub_cat[f'{user_sub_catorgory}'] #load the webpage for the sub_cat self.driver.get(sub_url_for_cat) except: print('exiting....due to invalid response')
运行时明明输入的分类/子分类看起来正确,但程序总是进入except块,请问原因是什么?
触发异常的核心原因
分类名称显示与字典键不匹配
从你提供的分类列表输出可见,分类名称显示为Animals & Pets(HTML转义后的&),但get_catogories返回的字典键是网页原始文本Animals & Pets。当你按照显示的&输入分类名称时,字典中找不到对应键,触发KeyError进入第一个except块。异常捕获范围过广+错误后未终止流程
第一个try块使用裸except,会捕获所有类型异常(比如输入中断、变量错误等),无法定位具体问题。且即使第一个try块出错,程序仍会执行第二个try块,此时user_sub_catorgory、cat_url等变量可能未定义,触发NameError进入第二个except块。子分类输入的匹配问题
若用户输入的子分类名称大小写、空格与字典键不一致(比如输入animal health而非Animal Health),会触发KeyError;若输入的非0值无法转为整数(比如输入不匹配的字符串),会触发ValueError,均会进入第二个except块。
修复方案
统一分类名称的显示与字典键
打印分类时直接使用字典原始键,避免HTML转义问题:print(f'Catogory: {key} ') # 确保key是未转义的原始文本,比如"Animals & Pets"精确捕获异常并终止错误流程
只捕获特定异常,出错后直接返回,避免后续代码访问未定义变量:try: user_catogory = input('Please select which catogory you would like to explore: ').strip() cat_url, cat_sub_cat = catogories_and_sub_cat[user_catogory] user_sub_catorgory = input('Please select the sub-catorgory or 0 for no sub-catorgory: ').strip() except KeyError: print('exiting....due to invalid category being selected') return # 出错后终止方法,避免进入后续try块 except Exception as e: print(f'exiting....unexpected error: {str(e)}') return优化子分类输入的校验逻辑
直接用字符串判断是否为0,避免整数转换异常,同时明确校验子分类是否存在:try: if user_sub_catorgory == '0': self.driver.get(cat_url) print('done') elif user_sub_catorgory in cat_sub_cat: sub_url_for_cat = cat_sub_cat[user_sub_catorgory] self.driver.get(sub_url_for_cat) print('done') else: print('exiting....invalid sub-category selected') except Exception as e: print(f'exiting....unexpected error: {str(e)}')可选:用数字选择替代字符串输入
为分类/子分类分配序号,让用户输入数字选择,彻底避免字符串匹配问题:# 打印分类时添加序号 category_list = list(catogories_and_sub_cat.keys()) for idx, key in enumerate(category_list, 1): print(f'{idx}. Catogory: {key} ') _, sub_cat = catogories_and_sub_cat[key] sub_list = list(sub_cat.keys()) for sub_idx, sub_key in enumerate(sub_list, 1): print(f' {sub_idx}. SubCatorgy: {sub_key}') # 用户输入序号选择分类 user_cat_idx = int(input('Please select category number: ')) user_catogory = category_list[user_cat_idx-1]
内容的提问来源于stack exchange,提问作者Darman

